# Bounded-edit guard on a new meeting, and English marker prompt E2 vs E3 (2026-09-26)

Exploratory, descriptive, unreviewed. Closes the two items left open by v0.2.11: the guard thresholds had been set on the same held-out outputs they were evaluated on, and the English readable prompt placed no `（※要確認）` markers on read speech. Code under test: tag v0.2.11 unmodified; the E3 prompt was injected per call for two extra conditions. Plan frozen (sha256 `3e896dbddb3a…`) before any download or inference.

## Inputs

- AMI `ES2008a` (scenario series ES2008, not used before), Mix-Headset, 17.4 min, CC BY 4.0; manual word reference (2506 words). Whisper large-v3-turbo with the application's default automatic language detection (detected: `en`), temperature 0, one run.
- FLEURS test split (revision `70bb2e84b976…`), 12 new clips per language (English, Japanese, Mandarin) after excluding every clip used in any earlier experiment, language pinned.

## Conditions

| Code | Engine | English prompt |
|---|---|---|
| V12 | default punctuation-only, 12B | — |
| R12 / R26 | readable draft with the shipped v0.2.11 guard, 12B / 26B-A4B | E2 (shipped) |
| R12E3 / R26E3 | same | E3: E2 plus three worked examples of marking an uncertain name, an unclear word and an ambiguous number |

3 repeats each at temperature 0, alternating order, one isolated loopback Ollama on one GPU.

## English FLEURS, 12 clips

| Metric | V12 | R12 | R26 | R12E3 | R26E3 |
|---|---:|---:|---:|---:|---:|
| Raw Whisper error rate | 4.38% | 4.38% | 4.38% | 4.38% | 4.38% |
| Delivered error rate | 4.38% | 4.38% | 4.38% | 4.38% | 4.38% |
| Invented tokens / 100 | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 |
| Omitted tokens / 100 | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 |
| Guard kept source (chunks) | 0/36 | 3/36 content_added 3 | 3/36 content_added 3 | 3/36 content_added 3 | 3/36 content_added 3 |
| Chunks formatted / unchanged / retained | 3/33/0 | 3/30/3 | 3/30/3 | 6/27/3 | 6/27/3 |
| Model markers placed / on real errors / on correct text | 0 / 0 / 0 | 0 / 0 / 0 | 0 / 0 / 0 | 3 / 3 / 0 | 3 / 3 / 0 |
| Error regions left unmarked / total | 24/24 | 24/24 | 24/24 | 21/24 | 21/24 |
| Punctuation position F1 / +type | 0.912 / 0.912 | 0.912 / 0.912 | 0.912 / 0.912 | 0.912 / 0.912 | 0.912 / 0.912 |
| Clips with 3 identical repeats | 12/12 | 12/12 | 12/12 | 12/12 | 12/12 |
| Median seconds per chunk (order-confounded) | 4.85 | 2.90 | 6.66 | 2.48 | 4.73 |

## Japanese FLEURS, 12 clips

| Metric | V12 | R12 | R26 |
|---|---:|---:|---:|
| Raw Whisper error rate | 3.49% | 3.49% | 3.49% |
| Delivered error rate | 3.49% | 3.82% | 3.63% |
| Invented tokens / 100 | 0.00 | 0.35 | 0.33 |
| Omitted tokens / 100 | 0.00 | 0.47 | 0.52 |
| Guard kept source (chunks) | 0/36 | 0/36 | 9/36 content_added 9 |
| Chunks formatted / unchanged / retained | 29/7/0 | 30/6/0 | 21/6/9 |
| Model markers placed / on real errors / on correct text | 0 / 0 / 0 | 13 / 6 / 7 | 3 / 0 / 3 |
| Error regions left unmarked / total | 48/48 | 40/48 | 48/48 |
| Punctuation position F1 / +type | 0.854 / 0.841 | 0.852 / 0.852 | 0.623 / 0.623 |
| Clips with 3 identical repeats | 11/12 | 8/12 | 12/12 |
| Median seconds per chunk (order-confounded) | 3.27 | 2.84 | 6.22 |

## Mandarin FLEURS, 12 clips

| Metric | V12 | R12 | R26 |
|---|---:|---:|---:|
| Raw Whisper error rate | 7.66% | 7.66% | 7.66% |
| Delivered error rate | 7.66% | 7.91% | 7.91% |
| Invented tokens / 100 | 0.00 | 0.24 | 0.24 |
| Omitted tokens / 100 | 0.00 | 0.00 | 0.00 |
| Guard kept source (chunks) | 0/36 | 0/36 | 3/36 omission_over_limit 3, content_added 3 |
| Chunks formatted / unchanged / retained | 33/3/0 | 33/3/0 | 30/3/3 |
| Model markers placed / on real errors / on correct text | 0 / 0 / 0 | 0 / 0 / 0 | 6 / 6 / 0 |
| Error regions left unmarked / total | 36/36 | 36/36 | 30/36 |
| Punctuation position F1 / +type | 0.983 / 0.963 | 0.928 / 0.879 | 0.861 / 0.812 |
| Clips with 3 identical repeats | 12/12 | 12/12 | 12/12 |
| Median seconds per chunk (order-confounded) | 3.14 | 2.72 | 6.24 |

## AMI ES2008a meeting

| Metric | V12 | R12 | R26 | R12E3 | R26E3 |
|---|---:|---:|---:|---:|---:|
| Raw Whisper error rate | 31.13% | 31.13% | 31.13% | 31.13% | 31.13% |
| Delivered error rate | 31.13% | 31.21% | 31.09% | 30.96% | 30.83% |
| Invented tokens / 100 | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 |
| Omitted tokens / 100 | 0.00 | 0.21 | 0.09 | 0.17 | 0.09 |
| Guard kept source (chunks) | 0/39 | 18/42 content_added 6, omission_over_limit 15, negation_count_changed 3 | 30/42 omission_over_limit 27, content_added 21, negation_count_changed 6 | 23/42 content_added 9, omission_over_limit 19, negation_count_changed 3 | 34/42 omission_over_limit 34, content_added 22, negation_count_changed 6 |
| Chunks formatted / unchanged / retained | 6/0/33 | 24/0/18 | 12/0/30 | 19/0/23 | 8/0/34 |
| Model markers placed / on real errors / on correct text | 0 / 0 / 0 | 0 / 0 / 0 | 0 / 0 / 0 | 0 / 0 / 0 | 0 / 0 / 0 |
| Error regions left unmarked / total | 849/849 | 849/849 | 849/849 | 849/849 | 849/849 |
| Clips with 3 identical repeats | 1/1 | 0/1 | 1/1 | 0/1 | 0/1 |
| Median seconds per chunk (order-confounded) | 3.86 | 3.38 | 2.00 | 3.40 | 1.99 |

## Frozen verdicts

| Rule | Result | Detail |
|---|---|---|
| `guard:R12:en_us retained<=2/36` | FAIL | 3 |
| `guard:R12:ja_jp retained<=2/36` | PASS | 0 |
| `guard:R12:cmn_hans_cn retained<=2/36` | PASS | 0 |
| `guard:R12:meeting omitted<=1.0` | PASS | 0.21 |
| `guard:R12:meeting invented==0` | PASS | 0.0 |
| `guard:R12:meeting err<=raw+0.5pt` | PASS | 0.08 |
| `guard:R26:en_us retained<=2/36` | FAIL | 3 |
| `guard:R26:ja_jp retained<=2/36` | FAIL | 9 |
| `guard:R26:cmn_hans_cn retained<=2/36` | FAIL | 3 |
| `guard:R26:meeting omitted<=1.0` | PASS | 0.09 |
| `guard:R26:meeting invented==0` | PASS | 0.0 |
| `guard:R26:meeting err<=raw+0.5pt` | PASS | -0.04 |
| `E3:(a) 12B hits +>=3` | PASS | 0->3 |
| `E3:(b) 12B precision>=0.5` | PASS | 3/3 |
| `E3:(c) 12B not worse` | PASS | {'hits': [0, 3], 'markers': [0, 3], 'err_f': [0.0438, 0.0438], 'err_m': [0.3121, 0.3096], 'inv_f': [0.0, 0.0], 'guard_f': [3, 3]} |
| `E3:(c) 26B not worse` | PASS | {'hits': [0, 3], 'markers': [0, 3], 'err_f': [0.0438, 0.0438], 'err_m': [0.3109, 0.3083], 'inv_f': [0.0, 0.0], 'guard_f': [3, 3]} |
| `E3:ADOPT` | PASS |  |

Guard rules: read-speech retention ≤ 2/36 outputs per language and condition; meeting after-guard omission ≤ 1.0 per 100, invented = 0, error rate ≤ raw + 0.5 pt. E3 adoption: 12B marker hits +≥3 over E2 (clips + meeting), hits/placed ≥ 0.5, no worsening (error rate +0.5 pt, invented 0 on clips, guard retention ≤ E2 + 2) for 12B and 26B.

## Outcome and follow-up changes (v0.2.12)

- **Guard on the new meeting: all declared expectations met** for both models (after-guard omission 0.21 and 0.09 per 100, invented 0, error rate within 0.1 pt of raw Whisper). On this native-English meeting the guard returned 18/42 (12B) and 30/42 (26B) readable chunks to the source; the default engine formatted 6/39.
- **Guard on read speech: the declared bound (≤ 2/36 outputs) was missed** by 12B English (3/36, a singular/plural change each time) and by 26B in all three languages (English 3, Japanese 9, Mandarin 3). Per the frozen plan the thresholds were **not** changed. Inspection showed one false-positive class in the Japanese added-content rule — a source name wrapped in new particles (`彼はウェールズは`) was treated as a new kana run — which v0.2.12 fixes by allowing kana runs already present in the source as segmentation fragments; applied to these saved candidates the fix lowers the 26B Japanese retentions from 9 to 6/36 and changes nothing else. The remaining retentions are paraphrases (`参照とされます` → `参照されるとされています`) and singular/plural or particle-level rewrites — conservative by design, documented as the cost of the bound.
- **Invented tokens on Japanese read speech (0.35 per 100 for 12B) are particles** (`や`, `へ`) added by grammar repair; the guard allows particle-level kana by design, and the metric counts them because they are in neither the source nor the reference.
- **E3 adopted** by the frozen rule: 12B marker hits rose from 0 to 3 with 3/3 on real errors and no metric worsened for either model. The gain is small and concentrated: the three hits are the same misheard name (`Dr. Malar Balasumbramian（※要確認）`) in three repeats of one clip, and **no English condition placed a single marker on the meeting** (849 error regions unmarked). E3 ships as the English prompt from v0.2.12; the marker behaviour on English meetings remains an open limitation.

## Limitations

- One new meeting and 12 clips per language; exploratory; single automatic reference; no human rating; FLEURS/AMI may be in training data; repeats are temperature-0 re-runs.
- "Markers on real errors" is a token-overlap heuristic (preceding 3 words / 12 characters); it cannot judge whether the marked span was in fact unclear to a listener.
- Token metrics do not see meaning changes that keep the same tokens.

Raw outputs, the frozen plan and hash, prompts_en.json (E2/E3 texts) and the input manifest are kept in the measurement archive; the plan hash is recorded in RESULTS.json.
