# mmv-voice v0.3b: repetition suppression evaluation

**Decision: HOLD.** The candidate met the held-out formatting target at **4/5 chunks (80.00%)**, but it failed the frozen ASR non-inferiority gate on the new held-out meeting. W1 increased TS3010a WER from 43.26% to 85.01% (+41.75 pp); the allowed increase was 1.00 point. All other frozen gates passed.

This is exploratory, descriptive, unreviewed local engineering evidence. It supports the release decision below, not a general performance or human-quality claim.

## Change evaluated

v0.3b consists of released v0.2.4 plus the previously frozen v0.3 formatter changes and one ASR change:

- **F1:** strip only the three exact known governance-scaffold families before the unchanged lexical/numeric validator.
- **F2:** prefer sentence endings, then whitespace, then a complete source atom when making exact-reconstruction 1,000-character chunks.
- **W1:** pass `condition_on_previous_text=False` to Whisper. `word_timestamps` and `hallucination_silence_threshold` were not enabled, and compression-ratio, log-probability, and no-speech thresholds remained at library defaults.

The W1 choice and all adoption thresholds were frozen and SHA-256 hashed before inference. The shipped v0.2.4 checkout was not modified.

## Data and protocol

The new held-out item was AMI **TS3010a**, a 1,041.024-second Mix-Headset scenario meeting from the TS3010 series. This series had not been used for ES2004a/IS1009a development. The audio was downloaded from [the AMI corpus mirror](https://groups.inf.ed.ac.uk/ami/AMICorpusMirror/amicorpus/TS3010a/audio/TS3010a.Mix-Headset.wav) (SHA-256 `d3a7d934b73fa5b9cfcff54ce3909ba59b544ef43c248c931082aa62056c1625`). The manual word reference came from [AMI public manual annotations v1.6.2](https://groups.inf.ed.ac.uk/ami/AMICorpusAnnotations/ami_public_manual_1.6.2.zip) (SHA-256 `b56e5babb2496b8795deeeda7e71178d7fbc9963f94276cf2a3f4b56ebbc9f9d`). AMI releases these signals and annotations under [CC BY 4.0](https://groups.inf.ed.ac.uk/ami/corpus/license.shtml).

Whisper `large-v3-turbo` ran once per ASR configuration on physical GPU0. AMI used the application's automatic language selection; the English and Japanese FLEURS clips used their frozen `en` and `ja` language arguments and temperature 0. The formatter was `gemma4:12b-it-qat`, digest `38044be4f923e5a55264ed7df4eaac2676651a905f735197c504045140c02bd3`, temperature 0, on an isolated loopback Ollama server. Each saved Whisper transcript was reused for all corresponding formatter conditions.

AMI WER lowercases NFKC word tokens, removes punctuation, retains lexical apostrophes, and orders manual speaker words by start time, end time, speaker, then source order. FLEURS reports aggregate English WER and Japanese CER. Punctuation scoring reports boundary-position F1 and boundary-position-plus-type F1. The unchanged validator guarantees lexical/numeric preservation; the evaluation independently rechecked token signatures and ASR error rates after delivery.

## ASR results

| Dataset | Baseline | W1 | Change | Frozen allowance | Gate |
|---|---:|---:|---:|---:|---|
| TS3010a held-out WER | 43.26% | 85.01% | +41.75 pp | ≤ +1.00 pp | **FAIL** |
| ES2004a dev WER | 27.18% | 27.21% | +0.04 pp | ≤ +1.00 pp | PASS |
| IS1009a dev WER | 32.16% | 28.69% | -3.47 pp | ≤ +1.00 pp | PASS |
| FLEURS en_us WER, 12 clips | 6.75% | 6.75% | +0.00 pp | ≤ +0.50 pp | PASS |
| FLEURS ja_jp CER, 12 clips | 6.66% | 6.66% | +0.00 pp | ≤ +0.50 pp | PASS |

W1 removed the targeted repetition pattern on all three AMI meetings. Tokens in identical runs of length at least three changed from **20 to 0** on TS3010a, **3 to 0** on ES2004a, and **129 to 0** on IS1009a. The longest IS1009a run fell from 88 consecutive `okay` tokens to 2.

The held-out failure is not hidden by that success. Automatic language selection returned `nl` for both TS3010a runs. Baseline previous-text conditioning allowed much of the English meeting to remain English, whereas W1 produced substantially more Dutch rendering. Against the English manual reference, the resulting substitutions/deletions outweighed the removed repetition. Because automatic language selection was frozen as the application path, the evaluation did not rescue the result by pinning English after observation.

## Formatter results and fix contributions

| Held-out condition | ASR input | Formatted chunks | Accepted chunks | Known scaffold stripped |
|---|---|---:|---:|---:|
| v0.2.4 | baseline | 0/6 | 0/6 | 0 |
| v0.3 (F1+F2) | same baseline | 2/6 | 2/6 | 4 |
| v0.3b (F1+F2+W1) | W1 | 4/5 | 4/5 | 0 |

On identical baseline ASR, F1+F2 added **2 formatted chunks** (0/6 to 2/6). The complete W1 pipeline added another **2 formatted chunks** relative to v0.3 (2/6 to 4/5), and its 80% rate passed the 50% formatting gate. That second delta is a pipeline contribution, not a same-input formatter ablation: W1 shortened and changed the transcript, so its denominator is five chunks rather than six.

## FLEURS regression

| Language | Acceptance v0.2.4 → v0.3b | Position F1 v0.2.4 → v0.3b | Position+type F1 v0.2.4 → v0.3b | Gate |
|---|---:|---:|---:|---|
| English | 11/12 → 11/12 | 0.8889 → 0.8889 | 0.8642 → 0.8642 | PASS |
| Japanese | 11/12 → 11/12 | 0.7674 → 0.7674 | 0.7209 → 0.7209 | PASS |

The 24 short FLEURS transcripts were byte-for-byte unchanged by W1 at the transcript level, so ASR, acceptance, and punctuation scores were unchanged. This is a regression result, not evidence that W1 is inert on longer or low-energy audio.

## Frozen gate verdicts

| Gate | Verdict |
|---|---|
| New held-out v0.3b formatted chunks ≥ 50% | PASS — 4/5 (80%) |
| W1 AMI WER ≤ baseline +1.0 point, each meeting | **FAIL** — TS3010a +41.75 points; both dev meetings pass |
| W1 FLEURS en WER / ja CER ≤ baseline +0.5 point | PASS — both unchanged |
| FLEURS acceptance and position+type F1 regression limits | PASS — both languages unchanged |
| Delivered tokens equal corresponding Whisper tokens everywhere | PASS |
| Existing and new tests | PASS — 14/14 |

The conjunction therefore yields **HOLD**, despite the formatting-rate and repetition-suppression gains.

## Limitations

- **Development contamination:** W1 was selected in response to failures on ES2004a and IS1009a. Their gains are descriptive and not independent confirmation.
- **Model-data contamination:** AMI and FLEURS may occur in model training data. No training-corpus cutoff or contamination-controlled subset was available, so broad benchmark claims are blocked.
- **One new meeting:** TS3010a is one held-out meeting from one site/series. It cannot estimate performance across meetings, accents, languages, acoustic conditions, or durations.
- **Automatic-language interaction:** both held-out runs selected Dutch. The large WER loss is specific to the frozen automatic-language application path and may not extrapolate to language-pinned transcription; pinning after seeing the result would be a new experiment.
- **No human rating:** no one rated punctuation naturalness, readability, or semantic appropriateness. The validator proves a bounded lexical/numeric property, not overall quality.
- **Order effects:** formatter conditions ran in the fixed order v0.2.4, v0.3, v0.3b on a warmed local service. Runtime comparisons are therefore not causal.
- **Single ASR run:** each configuration ran once. The release gate is deterministic and conservative, but this study does not estimate run-to-run variance.

## Reproducibility and operational checks

`v03b.diff` applies cleanly to v0.2.4 and includes `voice_mmv.py`, `whisper_gui.py`, the existing formatter tests, and the new transcribe-option test. Fourteen tests passed both in the candidate copy and after applying the diff to a fresh v0.2.4 archive. All generated ASR records identify physical GPU0; no GPU1 workload was launched. The isolated Ollama process was stopped by its recorded PID, port 11440 was verified closed, and the existing service on port 11434 remained listening. No commit, push, upload, or shipped-checkout edit was made.
