# mmv-voice v0.3: Long-form formatting repair evaluation

## Decision

**HOLD.** Version 0.3 substantially improved long-form formatting, but it missed the frozen held-out threshold: 5 of 11 held-out chunks were delivered with formatting (`45.45%`), below the required `50%`. One additional held-out chunk was valid but unchanged, so total validator acceptance was 6 of 11 (`54.55%`); the decision rule explicitly required *formatted* chunks, and that definition was not changed after observing the result.

The other adoption gates passed. English and Japanese FLEURS acceptance and punctuation scores were unchanged, delivered WER equaled Whisper WER in every meeting condition, and all 13 existing-plus-new tests passed.

## What changed

The candidate is a minimal two-part change to the v0.2.4 copy. The lexical validator itself is unchanged and remains case-sensitive.

**F1 — deterministic known-scaffold removal.** Before validation, v0.3 repeatedly removes only a leading, single-line route-transformer echo matching one of three anchors:

- `premise embedded in your question`
- `re-anchor against what I actually know`
- `as of my training data`

The match is the same narrow regular expression used by the existing GUI rewrite path. It does not remove arbitrary disclaimers or suffixes. Both the raw model output and the stripped candidate are retained in the audit record. A known scaffold followed by a changed word or changed letter case still fails the unchanged validator.

**F2 — sentence-aware chunking.** At the 1,000-character limit, v0.3 chooses the last sentence end (`. ? ! 。 ？ ！`), otherwise the last complete whitespace run, otherwise the last complete source atom. It never splits an ASCII word or numeric atom, preserves exact reconstruction, and handles Japanese text without spaces. An indivisible atom longer than the limit remains whole and is rejected by the existing unsent-source guard.

## Frozen evaluation

The plan and decision rule were frozen before inference with SHA-256 `d279c5438a2e8cec11fb584d964ff5b60609f18758c133ac21af4b8d0bbe90e4`.

The development meeting was AMI `ES2004a`, the meeting that motivated the repair. Its previously saved Whisper transcript was reused. The held-out meeting was AMI `IS1009a`, from a different scenario series. Its Mix-Headset signal was transcribed once with Whisper `large-v3-turbo`, automatic English detection, and temperature 0. Both formatter conditions used the same E2 prompt, `gemma4:12b-it-qat` at digest `38044be4f923…`, temperature 0, speaker attribution off, and physical GPU0 through an isolated Ollama server.

Held-out source provenance:

- Audio: [AMI IS1009a Mix-Headset WAV](https://groups.inf.ed.ac.uk/ami/AMICorpusMirror/amicorpus/IS1009a/audio/IS1009a.Mix-Headset.wav), 838.833 seconds, SHA-256 `a1bd1c4488221173961be84f35ba924b9bdaf0c0dacc9d7b1dbd54db40f918d6`.
- Manual annotations: [AMI manual annotations v1.6.2](https://groups.inf.ed.ac.uk/ami/AMICorpusAnnotations/ami_public_manual_1.6.2.zip), SHA-256 `b56e5babb2496b8795deeeda7e71178d7fbc9963f94276cf2a3f4b56ebbc9f9d`.
- License: [Creative Commons Attribution 4.0](https://groups.inf.ed.ac.uk/ami/corpus/license.shtml).

For regression, the evaluation reused saved ASR for the same 12 `en_us` FLEURS clips from the v0.2.4 evaluation and the same 12 `ja_jp` FLEURS clips from the prior Japanese formatter evaluation. No FLEURS audio was transcribed again.

## Long-form results

| Dataset | Condition | Chunks | Formatted | Unchanged | Source retained | Accepted total | Wall time |
|---|---:|---:|---:|---:|---:|---:|---:|
| AMI ES2004a (development) | v0.2.4 | 12 | 1 | 0 | 11 | 1 | 51.76 s |
| AMI ES2004a (development) | v0.3 | 13 | 9 | 0 | 4 | 9 | 49.29 s |
| AMI IS1009a (held-out) | v0.2.4 | 10 | 1 | 0 | 9 | 1 | 40.04 s |
| AMI IS1009a (held-out) | v0.3 | 11 | 5 | 1 | 5 | 6 | 40.72 s |

The different chunk counts are expected: sentence-aware boundaries shorten some chunks and therefore change the partition while preserving exact source reconstruction.

On development data, v0.3 formatted `69.23%` of chunks, versus `8.33%` for v0.2.4. On held-out data, it formatted `45.45%`, versus `10.00%` for v0.2.4. This is a large practical improvement on these two meetings, but the held-out value remains below the frozen adoption floor.

### Which fix helped

Both fixes contributed, but the evaluation does not provide a clean causal ablation of F2.

- Applying F1 alone to the already-saved v0.2.4 candidates matched six development scaffolds and five held-out scaffolds. It newly rescued one chunk in each meeting; the remaining matched candidates still contained a case or lexical change and correctly failed.
- In the full v0.3 run, F1 stripped six candidates in each meeting. Four stripped development candidates and three stripped held-out candidates passed the unchanged validator.
- F2 changed the partitions from 12 to 13 development chunks and from 10 to 11 held-out chunks. Among full-v0.3 outputs that did not require F1, five development chunks and three held-out chunks were accepted (the held-out count includes one valid unchanged chunk). This is consistent with sentence-aligned starts reducing boundary-case failures, but model inference was rerun on different chunks, so this is not an isolated F2 counterfactual.

The remaining five held-out chunks were source-retained because their candidates still failed with `content_changed`. The guard therefore continued to fail safely.

## WER preservation

WER used Unicode NFKC normalization, lowercase alphanumeric word tokens with internal apostrophes retained, and a deterministic speaker merge ordered by word start time, end time, speaker ID, and source order. Punctuation and non-lexical annotation events were excluded.

| Dataset | Reference words | Whisper words | S / D / I | Whisper WER | v0.2.4 delivered WER | v0.3 delivered WER |
|---|---:|---:|---:|---:|---:|---:|
| ES2004a | 2,653 | 2,250 | 162 / 481 / 78 | 0.271768 | 0.271768 | 0.271768 |
| IS1009a | 2,018 | 1,758 | 185 / 362 / 102 | 0.321606 | 0.321606 | 0.321606 |

For all four meeting-condition pairs, the delivered normalized token sequence exactly equaled the raw Whisper token sequence. This proves lexical fail-safe behavior under this scorer; it does not prove that every accepted punctuation decision was natural or correct.

## FLEURS regression

| Language | Condition | Accepted clips | Position F1 | Position+type F1 |
|---|---:|---:|---:|---:|
| English (`en_us`) | v0.2.4 | 11 / 12 | 0.888889 | 0.864198 |
| English (`en_us`) | v0.3 | 11 / 12 | 0.888889 | 0.864198 |
| Japanese (`ja_jp`) | v0.2.4 | 11 / 12 | 0.767442 | 0.720930 |
| Japanese (`ja_jp`) | v0.3 | 11 / 12 | 0.767442 | 0.720930 |

There was no observed regression in clip acceptance, punctuation position F1, or punctuation position+type F1. These are single deterministic formatter runs over saved ASR, not independent statistical replicates.

## Frozen gate verdicts

| Gate | Result | Evidence |
|---|---|---|
| Held-out formatted chunks at least 50% | **FAIL** | `5 / 11 = 45.45%` |
| English acceptance loss no worse than 1 clip | PASS | `11 / 12` versus `11 / 12` |
| Japanese acceptance loss no worse than 1 clip | PASS | `11 / 12` versus `11 / 12` |
| English position+type F1 loss no worse than 0.02 | PASS | `0.864198` versus `0.864198` |
| Japanese position+type F1 loss no worse than 0.02 | PASS | `0.720930` versus `0.720930` |
| Delivered WER equals Whisper everywhere | PASS | Exact in all four meeting-condition pairs |
| Existing and new tests pass | PASS | 13 passed, 0 failed |

Because adoption required every gate, the single held-out coverage failure determines **HOLD**.

## Verification and limitations

The unified implementation-and-tests diff applies cleanly to the v0.2.4 working tree. Tests cover exact known-scaffold removal, rejection after a scaffold plus a changed word or case, non-removal of unknown text, repeated known echoes, exact chunk reconstruction, limit compliance, no ASCII mid-word split, sentence preference, Japanese without spaces, and preservation of oversized atoms. The baseline checkout was not edited, committed, pushed, or uploaded.

This is exploratory, unreviewed engineering evidence. AMI and FLEURS training contamination is unresolved. F1 and F2 were designed after inspecting ES2004a, and there is only one held-out meeting. There was no human naturalness or readability rating. The fixed baseline-first run order does not rule out cache or order effects. AMI Mix-Headset contains overlapping speakers, while the WER scorer imposes one deterministic stream and is not overlap-aware. Consequently, the results support a narrow claim—v0.3 improves safe long-form formatting coverage on these meetings without changing delivered lexical tokens—but not a corpus-wide quality or usability claim.
