MMV Voice

Linux desktop · Local by default

A clearer transcript.
Still your words.

Turn recordings into readable text on your own computer. Whisper listens. MMV adds punctuation. A preservation check keeps your original words intact.

No account or API key for local use. Free software, AGPL-3.0.

WORKFLOW ILLUSTRATIONON YOUR COMPUTER
ORIGINAL WORDS

We met today the budget is 320 dollars

PUNCTUATION, NOT A REWRITE

We met today, the budget is 320 dollars.

A changed word or number? The original is retained.
Keep it on your machine.

Local speech and language models. Cloud processing is a separate, explicit choice.

See what changed.

Original, candidate and preservation result stay available for review.

Save what you need.

Copy the transcript, save a text file or keep the full review digest.

One recording. One clear workspace.

Open a recording, follow its progress, then review the original beside the result. Notes and extra checks stay out of the way until you need them.

MMV Voice desktop workspace: recording controls on the left, progress at the top, original and transcript side by side, and text export below
Desktop interface preview. This page is a guide; audio processing runs in the installed desktop app.
Where your data goes: everything runs inside your computer; the network is used once at install time; the cloud engine is opt-in and off by default
Everything runs inside your computer. The network is used once, by the installer; the cloud engine is opt-in and asks before sending.

Local-first, concretely

Ready when you are.

Check your setup first. The installer asks before downloading anything.

git clone https://github.com/mobius-style/mmv-voice && cd mmv-voice
bash install.sh --check   # shows what is missing; changes nothing
bash install.sh           # installs locally, asks before every download
bash launch.sh            # opens the app

Linux desktop, Python 3.10+, ffmpeg, python3-tk, and Ollama running. An NVIDIA GPU with 8–16 GB is recommended. The installer never uses sudo and never starts or stops Ollama.

Using it

  1. Open an audio file — transcription starts at once, language auto-detected.
  2. Formatting follows: chunks of up to 1,000 characters go to the local model for punctuation and paragraph breaks.
  3. Read the result; any chunk that failed the check is shown exactly as Whisper wrote it.
  4. The "Review" tab shows source, candidate and reason for every chunk.
  5. Save the text (optional: minutes or a digest file).

Optional readable draft (off by default). A checkbox in "Optional steps" lets the local model rewrite for readability instead of only adding punctuation. In that mode the word-and-number guarantee does not apply; unclear names and phrases are kept as heard and marked (※要確認) ("please verify"), the original and a diff stay in the Review tab, and every result needs human review. Measured on 38 known inputs (report) and on 36 new clips plus one meeting with the shipped code (held-out comparison): on read speech it invents nothing and repairs nothing (error rate unchanged); on a meeting it dropped 3.5–9.3 content words per 100, so v0.2.10 adds a code-level guard that keeps the original chunk when numerals change, content disappears or content is added (on those meeting outputs: 14–15 of 18 chunks kept; on a second, native-English meeting 18–30 of 42, with content loss ≤ 0.2 per 100 and nothing invented). Since v0.2.13 numerals must keep their order, a changed "might"/"if" retains the original, and words cannot swap places. Review stays mandatory.

How your words are protected

A candidate that only adds punctuation is used; a candidate that changes a number or a word is discarded and Whisper's text is kept
Only punctuation and whitespace may differ from Whisper's text; anything else sends the chunk back to Whisper's version. Enforced in code and tests, not a formal proof; it cannot fix words Whisper misheard.

Measured, with limits.

Docs, UI, prompts and tests are written for English. v0.2.4 removes the one sentence that made the model add stray semicolons. On 12 new English FLEURS clips, punctuation now roughly matches Whisper's own (type-aware F1 0.854 vs 0.821, intervals overlap) instead of clearly worse (v0.2: 0.659); v0.2.4 leaves most clips' punctuation alone (7 of 12) and adds punctuation on 4; 11 of 12 clips were accepted, and no word or number was ever changed — the one clip whose candidate changed a word fell back to Whisper's text.

Small read-speech panels; the 95% intervals overlap; FLEURS training overlap unresolved. Details: v0.2.4 evidence.

Long recordings: what to expect today

On a real 17.5-minute, four-person meeting (AMI ES2004a), transcription took 32 s including model load (27 s transcribing) on one RTX 5070 Ti (WER 27% on overlapping, informal speech). Formatting stepped aside on 11 of 12 chunks, returning Whisper's text unchanged — your words stay safe, but on meetings you mostly get Whisper's own punctuation. Observed causes: the MMV governance layer prepended a boilerplate sentence to 6 conversational chunks; in the other 5 the model changed capitalisation, mostly where the 1,000-character split cut a sentence. Two fixes were measured and put on hold: stripping the boilerplate and splitting at sentence ends (v0.3) helped but missed its gate; adding Whisper's condition_on_previous_text=False (v0.3b) removed repetition loops and reached 4 of 5 formatted chunks, but Whisper's automatic language detection then labelled an English meeting as Dutch and WER rose from 43% to 85%. Reports: v0.3, v0.3b.

Measured results

PanelCandidate acceptedSource retainedPunctuation-position F1
English, 12 new clips × 3 (2026-09-26), v0.2.411/12 clips1/120.846 → 0.878
English meeting, 17.5 min, 12 chunks (AMI), v0.2.41/12 chunks11/12—
English, 8 held-out clips × 3 (2026-09-24), v0.2 prompt8/8 clips0/80.766 → 0.769
Mandarin (reference), 8 clips × 3 (2026-09-26)21/24 (7/8 clips)3/240.278 → 0.786
Japanese, 12 new clips × 3, v2 prompt91.7% (33/36)3/360.857
Japanese, same clips, previous v1 prompt83.3% (30/36)6/360.857
English, same 8 clips, regression (pre-release build)8/80/8—

Acceptance means the candidate passed the preservation check, not that its punctuation is right. The position-accuracy metric is identical for both prompts (paired delta +0.001, 95% CI [−0.089, +0.088]). Small panels, read speech, no population guarantee. Details and failure cases are in the status report.

What it is not

It does not correct transcription errors — a candidate that fixes a misheard word is rejected because the words changed. Punctuation can still alter meaning; the check does not prove semantic fidelity. Mandarin has a small reference measurement; other languages are unevaluated.