Initial commit of current state
This commit is contained in:
+125
@@ -0,0 +1,125 @@
|
||||
# Transcription accuracy
|
||||
|
||||
What was measured, what shipped, and what turned out to be wrong along the way.
|
||||
|
||||
Benchmark throughout: eight sentences of technical speech (85 words) containing
|
||||
the vocabulary this project actually uses — pipecat, Kokoro, Metamate,
|
||||
Phabricator, fbsource, SFSpeechRecognizer. The audio is synthesised, so it is
|
||||
cleaner and more evenly paced than a person at a microphone. Trust the ordering,
|
||||
not the absolute rates, and prefer `compare_engines.py` on your own voice.
|
||||
|
||||
Scoring follows the [Inscribe
|
||||
benchmark](https://get-inscribe.com/blog/apple-speech-api-benchmark.html):
|
||||
lowercased, depunctuated, contractions expanded and **digits written out**. That
|
||||
last part matters — see the correction below.
|
||||
|
||||
## Where it ended up
|
||||
|
||||
| Configuration | WER |
|
||||
| --- | --- |
|
||||
| old `SFSpeechRecognizer`, no help | 20.0% |
|
||||
| `SpeechTranscriber`, no help | 16.5% |
|
||||
| old + vocabulary + repair rules | 10.6% |
|
||||
| `DictationTranscriber` + vocabulary + repair | 9.4% |
|
||||
| **`SpeechTranscriber` + repair — the default** | **9.4%** |
|
||||
|
||||
The top two tie, but not at equal cost. `SpeechTranscriber` reaches 9.4% with
|
||||
**five repair rules and no vocabulary**, where the old model needs 39 curated
|
||||
terms *and* seven rules to match it. It is also better with no configuration at
|
||||
all, which is the state anyone starts in.
|
||||
|
||||
## The two mistakes I made measuring this
|
||||
|
||||
**Scoring punished the new model for formatting.** My first normaliser
|
||||
lowercased and stripped punctuation but left digits alone, so
|
||||
`SpeechTranscriber` was charged errors for writing "400 milliseconds" instead of
|
||||
"four hundred milliseconds", and "SF speech recognizer" instead of
|
||||
"SFSpeechRecognizer". Neither is a mishearing. The Inscribe benchmark names this
|
||||
exact trap. Rescored properly, a 12.9%-vs-14.1% loss for the new model became a
|
||||
9.4% tie that it wins on effort.
|
||||
|
||||
**"Dilution is cheap" was wrong.** Padding the term list with a thousand random
|
||||
dictionary words cost only 1.1 points, so auto-generating 100 terms from project
|
||||
filenames looked safe. It erased the entire benefit — back to no-biasing levels
|
||||
— because filenames like `bot`, `plan` and `hack` are ordinary English, and
|
||||
boosting ordinary words drags correct speech onto them. Obscure words are inert
|
||||
by comparison. Discovered terms are now filtered against the system dictionary,
|
||||
capped at 12, and ranked behind the hand-written list.
|
||||
|
||||
## What each mechanism is worth
|
||||
|
||||
**Vocabulary biasing** (`vocabulary.txt` → `contextualStrings`) removed 30% of
|
||||
errors on the old model and costs nothing at runtime. `SpeechTranscriber`
|
||||
**ignores it** — output is byte-identical with and without terms — so on the
|
||||
default engine the file now serves only Claude's system prompt and the fallback
|
||||
engines. Keep it short and specific.
|
||||
|
||||
**Repair rules** (`corrections.txt`) are deterministic substitutions for
|
||||
mistakes that recur identically. Worth more on the new engine than the old,
|
||||
because its errors land phonetically close to the target — "Kakoro" for Kokoro,
|
||||
where the old model produced "coral". Five rules do what 39 terms plus seven
|
||||
rules do elsewhere.
|
||||
|
||||
**LLM post-correction** took 16.5% to 7.1% on the old engine, and 5-best beat
|
||||
1-best (7.1% against 8.2%). Not worth a separate pass on the conversation path,
|
||||
since Claude already reads every transcript — the vocabulary goes into his
|
||||
system prompt instead, which is free. Worth a real pass for meeting notes.
|
||||
|
||||
**Punctuation and task hint** on the old API measurably do nothing for accuracy.
|
||||
The new engine punctuates and capitalises automatically.
|
||||
|
||||
**Noise beats all of it.** At 10 dB SNR, WER nearly tripled to 69% and neither
|
||||
vocabulary, high-pass filtering nor gain normalisation recovered any of it. The
|
||||
microphone is the highest-leverage component in the pipeline.
|
||||
|
||||
## Which model is which
|
||||
|
||||
| | old `SFSpeechRecognizer` | `DictationTranscriber` | `SpeechTranscriber` |
|
||||
| --- | --- | --- | --- |
|
||||
| Model | old | **byte-identical to old** | **new** |
|
||||
| Honours vocabulary | yes | yes | no |
|
||||
| Long audio | needs stitching | native | native |
|
||||
| Punctuation | no | no | automatic |
|
||||
| Best measured | 10.6% | 9.4% | **9.4%, far less tuning** |
|
||||
|
||||
`DictationTranscriber` is not a different model: it produces character-for-
|
||||
character identical output to `SFSpeechRecognizer` on the same audio, verified
|
||||
sentence by sentence rather than inferred from the docs. It is the old
|
||||
recogniser reached through the new API, and its only advantage is vocabulary
|
||||
support.
|
||||
|
||||
Both new-API modules handle long speech natively — 40 of 40 sentences recovered
|
||||
from 83 seconds — so the transcript stitching in `apple_stt.py` is only needed
|
||||
on the legacy fallback path.
|
||||
|
||||
For scale, the published benchmark measured 5,559 LibriSpeech utterances on an
|
||||
M2 Pro: SpeechAnalyzer 2.12% clean and 4.56% noisy, against SFSpeechRecognizer's
|
||||
9.02% and 16.25% — the legacy engine scoring worse than Whisper Tiny. It
|
||||
validates its harness by reproducing OpenAI's published Whisper numbers within
|
||||
+0.11 to +0.42. Caveats: English read speech, one machine.
|
||||
|
||||
## Ruled out
|
||||
|
||||
**Denoising.** *When De-noising Hurts* ([arXiv:2512.17562](https://arxiv.org/abs/2512.17562))
|
||||
applied MetricGAN+ across 500 recordings and 10 conditions: **all 40
|
||||
configurations got worse**, Whisper degrading 8.82% → 25.83% at 10 dB SNR, and
|
||||
even clean audio losing 1.3–3.2 points. So Pipecat's `RNNoiseFilter` is not
|
||||
worth the dependency surgery it needs, and the commercial filters (Koala, Krisp,
|
||||
AIC) are the same bet with a licence. Worth checking macOS Voice Isolation is
|
||||
off, since it is the same class of processing applied for free.
|
||||
|
||||
**Custom language models.** `SFSpeechLanguageModel` trains on example sentences
|
||||
and supports X-SAMPA pronunciations, so it should beat `contextualStrings` — but
|
||||
its data builder is Swift-only, and it cannot be combined with
|
||||
`SpeechTranscriber` anyway.
|
||||
|
||||
## Still open
|
||||
|
||||
- **Real speech.** Every number here is synthesised audio with config files
|
||||
tuned around it. `compare_engines.py` records you and runs every engine.
|
||||
- **The external microphone**, versus the built-in one. Probably the largest
|
||||
real-world effect available and free to test.
|
||||
- **5-best into Claude.** `--alternatives` exists in the helper and is unused;
|
||||
measured 8.2% → 7.1% on the old engine.
|
||||
- **Confidence-gated clarification** — ask again rather than answer a
|
||||
low-confidence guess.
|
||||
Reference in New Issue
Block a user