# 12B bounded readable editing — 2026-09-27

Exploratory engineering comparison; unreviewed. Unchanged gemma4:12b-it-qat weights and model digest. Baseline: v0.2.12 (b9c55b9). No fine-tuning, additional judge or generation pass. Default punctuation-only mode remains unchanged.

## Changes
- Reject detected uncertainty/condition changes instead of delivering the changed claim with a warning marker.
- Check numeral order, not only the multiset; preserve relative order of surviving content tokens.
- A model-specific prompt narrows edits to avoid conflicts with the guard, selected separately per language using the frozen rule. Original source, draft, differences and reason stay available.

## Evaluation
Two rounds of 36 new FLEURS clips (12 each English/Japanese/Mandarin), 72 unique audio clips total, each condition repeated three times at temperature 0. Whisper large-v3-turbo, pinned language. The two rounds produced 540 formatter outputs: round1 baseline/proposal, round2 baseline/proposal/guard-only ablation. Same Whisper output goes to all conditions. References never go to the model. Input manifests, audio hashes, code/prompt hashes and raw rows are retained in the local measurement archive. These are repeated runs, not 108 independent samples per language.

Round1 compact prompt was rejected: Mandarin omission exceeded the +0.2/100 bound; Japanese punctuation and useful markers deteriorated. Round2 retained the original annotation guidance and narrowed the allowed edits. Old evaluation inputs, including AMI ES2008a, were development data, not new validation.

## Round2 results
WER/CER is mean clip error percentage (English words; Japanese/Chinese characters). Added/omitted are token metrics per 100 source tokens, not semantic ratings. Punctuation is position F1. Marker hits use the previous 3-word/12-character overlap heuristic, not a calibrated human judgment.

| Language | Condition | WER/CER % | Added/100 | Omitted/100 | Punctuation F1 | Retained/36 | Marker hits/placed | Median s |
|---|---|---:|---:|---:|---:|---:|---:|---:|
| en_us | old | 3.754 | 0.000 | 0.000 | 0.979 | 6/36 | 3/3 | 0.81 |
| en_us | new | 4.171 | 0.000 | 0.000 | 0.924 | 3/36 | 3/3 | 0.91 |
| en_us | guard | 3.754 | 0.000 | 0.000 | 0.979 | 6/36 | 3/3 | 0.84 |
| ja_jp | old | 5.116 | 0.000 | 0.086 | 0.627 | 9/36 | 6/9 | 0.89 |
| ja_jp | new | 4.771 | 0.000 | 0.000 | 0.680 | 6/36 | 6/6 | 0.98 |
| ja_jp | guard | 5.116 | 0.000 | 0.086 | 0.627 | 9/36 | 6/9 | 0.95 |
| cmn_hans_cn | old | 5.564 | 0.000 | 0.000 | 0.900 | 3/36 | 9/9 | 0.84 |
| cmn_hans_cn | new | 5.564 | 0.000 | 0.000 | 0.946 | 0/36 | 18/21 | 0.92 |
| cmn_hans_cn | guard | 5.564 | 0.000 | 0.000 | 0.900 | 3/36 | 9/9 | 0.90 |

old = published prompt/guard; new = proposed prompt/new guard; guard = published prompt/new guard.

## Selection
- en_us: guard; checks: {"error": true, "added": true, "omitted": true, "punctuation": false, "marker_hits": true}.
- ja_jp: new; checks: {"error": true, "added": true, "omitted": true, "punctuation": true, "marker_hits": true}.
- cmn_hans_cn: new; checks: {"error": true, "added": true, "omitted": true, "punctuation": true, "marker_hits": true}.

Any language failing a proposed-prompt criterion keeps the original prompt with the improved guard. Selection uses this panel; another independent panel is required before claiming a general quality gain.

## Limits and counterexamples
- No human fluency rating or calibrated semantic judge. This does not establish a percentage for overall text quality (`文章品質`) or an improvement in ASR correction ability. Training-set overlap cannot be ruled out; new local IDs/references do not establish training cleanliness.
- Guards are lexical heuristics: changing grammatical roles, scope or meaning using the same tokens can still escape. They intentionally retain some legitimate rewrites. Kana-only Japanese content is only partly covered.
- Annotation is not reliable error detection; an unmarked phrase remains unverified. Added-token zero is not proof of no invented meaning.
- The development meeting contains one repeated-filler generation hitting the output cap; source is retained. Meeting development performance must not be presented as held-out evidence. No new multilingual meeting test was run.
- Three temperature-zero repeats assess repeatability; CIs cluster by clip. There is no powered hypothesis test or claim of statistical significance. G1: frozen before formatter inference; G2: exploratory/unreviewed; G3: descriptive and cluster CIs only; G4: no calibrated human-quality inference; G5: training contamination unresolved; G6: counterexamples/unit negative controls, no research novelty claim.

## Local reproduction
Archive: `/home/happy/.codex/work/mmv_12b_polish_20260927/`. `round2/FREEZE.json`, `round2/samples.json`, `round2/INPUT_AUDIT.json`, `round2/runs/`, `round2/RESULTS.json`, `SELECTION.json`. Frozen measured code is `candidate_round1/` and `candidate_round2/`; `candidate/` contains the selected local patch. `round2/run_holdout.py` uses loopback 11447 and the existing MMV adapter; its snapshot symlink keeps the round2 code immutable. No GH/HF publication in this task.

## Runtime and regression verification

46 regression tests passed under Xvfb with CUDA_VISIBLE_DEVICES=0. Real GUI postprocessing used the selected model/prompt configuration for English, Japanese and Mandarin and checked displayed text and audit rows. The 72 new clips were actually transcribed with Whisper; the GUI smoke reused those transcripts rather than rerecording them.

11 constructed adverse edits (uncertainty/condition loss, swapped numerals and reordered people) were rejected 5/11 by the baseline and 11/11 by the new guard. These targeted regression examples are not an accuracy estimate. Ordinary punctuation/filler controls pass.

An initial GUI startup probe encountered this host's CUDA device enumeration failure. GPU status probes now show the failed device and continue, and device selection skips failed probes instead of aborting. The healthy/broken-device branches were tested with controlled failures. GPU repair itself is outside this change; the actual model/audio runs used GPU0. No system-wide GPU setting was changed.

The final English prompt is unchanged. The development meeting's 7-to-4 retention improvement belonged to the rejected English prompt and is NOT a gain of the selected final configuration. Do not quote that improvement as shipped behavior.
