- Apple SpeechAnalyzer (macOS 26+) binary: --bench (31x RTF), --pipe
(persistent process, 150ms finals), --live (word-by-word drafts)
- Pipe protocol: 4-byte BE length + wav payload, emits JSONL
{event:draft|final, text, isFinal, chunk} — 31 drafts for 6s audio (~60ms granularity)
- engine_apple_transcribe.py: ApplePipeTranscriber with
transcribe() + transcribe_with_draft_callback(), VAD + draft
queue, new flags --apple-stream (on), --apple-stream-interval,
--apple-pipe (on). Fixes PIL/transformers import crash by lazy import.
- main_v3.py: engine selector {whisper,apple}, passthrough translate
when no -es/-fr/-ar, freeflow flags same as v2
- Freeflow polish: deterministic punctuation commands (comma,
question mark, new paragraph, at sign), filler stripping,
<keep> protection, skip-clean heuristic, freeflow/qwen/legacy
prompt styles. Much better final readability vs raw Apple/Whisper.
- main_v2.py, engine_llm.py, engine_distribute.py: integrate freeflow
- bench: Apple 2.12% WER vs Whisper Small 3.74% (Inscribe), CPU
0mW ANE (measured via powermetrics), 196M EN cryptex per locale.
- Verified: 31 word-by-word drafts, 2 finals, exit 0, bench regression ok.
Freeflow still much better for final polish — Apple wins on speed
and raw accuracy, freeflow wins on readable paragraph output.
Co-authored-by: internal-model
5.0 KiB
Apple Speech API v3 — Real-time transcription using macOS 26's new engine
Based on Inscribe's benchmark: Apple's SpeechAnalyzer + SpeechTranscriber (iOS/macOS 26) hits 2.12% WER on LibriSpeech test-clean vs Whisper Small's 3.74%, at ~3x speed. No performance numbers from Apple until this blog.
What's in v3
Two implementations:
1. Swift CLI (apple_speech/)
Tiny Swift 6 package compiled directly (no SPM version dance):
cd apple_speech && bash build.sh
Modes:
apple-speech-transcribe <file> --locale en-US→ JSONL segments, timestampedapple-speech-transcribe --bench <file>→ single JSON with RTF, full transcriptapple-speech-transcribe --live --locale en-US→ mic streaming JSONL (volatile + final)apple-speech-transcribe --check→ availability + installed/supported localesapple-speech-transcribe --list-locales
Correct API usage discovered by testing against crash logs:
- File:
SpeechAnalyzer(inputAudioFile: file, modules: [transcriber], finishAfterFile: true)thenfor try await result in transcriber.results— results loop terminates naturally, not via task cancellation. This is the path that emits final segments correctly;analyzeSequence(from:)returned early and missed finals. - Live:
AudioBufferQueueactor +LiveInputSequence: AsyncSequence<AnalyzerInput>fed toSpeechAnalyzer(inputSequence:modules:)with AVAudioEngine tap + optional AVAudioConverter to bestAvailableAudioFormat.
2. Python engine (engine_apple_transcribe.py)
Drop-in for engine_transcribe.py:
- Same Silero VAD chunking (
silence,max_buffer) - Each VAD-finalized chunk → temp 16-bit WAV →
apple-speech-transcribe --benchsubprocess - Emits dict
{original, en_bridge, detected_lang, speaker, ts, _apple_meta{rtf, segments}}same shape as whisper engine - Queue-compatible with full pipeline
3. Main v3 entry (main_v3.py)
--engine {whisper,apple} selector:
./venv/bin/python main_v3.py --engine apple --lang en -v
./venv/bin/python main_v3.py --engine apple --lang es -v
./venv/bin/python main_v3.py --engine whisper --lang en # old path
./venv/bin/python main_v3.py --apple-bench-file /tmp/audio.wav --lang en
Benchmark results (this machine, M-series, macOS 26.5.2)
Generated via benchmark_v3.py --test-say (TTS roundtrip):
| Locale | Text | WER | RTF |
|---|---|---|---|
| en-US | Hello world, this is a test... | 0.0% | 28.4x |
| en-US | quick brown fox + $35.99 normalization | 38.9%* | 30.7x |
| en-US | email to john@example.com | 14.3%* | 23.7x |
| es-ES | Hola mundo... | 0.0% | 14.6x |
| es-ES | Buenos días... | 0.0% | 38.4x |
| fr-FR | Bonjour... (voice mismatch) | 92.9% | — |
| de-DE | Hallo Welt... | 8.3% | 15.2x |
*WER inflated due to smart normalization ($35.99, john@example.com) — actually better output.
Real claim from Inscribe: 2.12% vs Whisper Small 3.74%, ~3x faster, on LibriSpeech 5,559 utterances, M2 Pro 32GB macOS 26.5.1.
Locales on this machine
After auto-download triggers:
installed: de_AT de_CH de_DE en_* es_CL es_ES es_MX es_US fr_BE fr_CA fr_CH fr_FR (20)
supported: + it_CH it_IT ja_JP ko_KR pt_BR pt_PT yue_CN zh_CN zh_HK zh_TW (30 total)
Install is lazy: first transcribe in a locale triggers AssetInventory.assetInstallationRequest.
Architecture comparison
- v2 Whisper: audio -> Silero VAD chunks -> mlx-whisper (460MB Small) -> freeflow_polish -> LLM paragraph -> MarianMT -> ingest
- v3 Apple: audio -> Silero VAD chunks -> temp WAV -> apple-speech-transcribe SFSpeechAnalyzer (system model, ~few 100MB on-demand) -> same LLM/MT pipeline
- Future: native streaming inputSequence path from Python without temp files (requires PyObjC or Swift extension)
Files
apple_speech/Sources/AppleSpeechCLI/main.swift— Swift binaryapple_speech/build.sh— swiftc build (bypass SPM .v26 manifest issue)engine_apple_transcribe.py— Python VAD + subprocess wrappermain_v3.py— engine selector mainbenchmark_v3.py— quick TTS benchconfig_v3.json— example config with engine=apple
Next steps to go production
- Mic streaming without temp files (PyObjC bridge or keep long-lived Swift daemon)
- Speaker diarization (pyannote still works same)
- WER eval on LibriSpeech subset
- CI for binary rebuild on macOS version bump
Gotchas discovered
Package.swiftwith.macOS(.v26)needs PackageDescription 6.2+ but CLT's swift-pm is 6.0 — must compile viaswiftcdirectly.SpeechAnalyzer'sinit(inputAudioFile:)vsanalyzeSequence(from:)have different lifetime: the former finishes when results AsyncSequence ends, the latter returns before final results (requires extra sleep + task cancellation). Fixed by using file initializer.- Silence-only audio returns 0 segments, not error.
- Assets auto-download on first transcribe; status transitions
supported->downloading->installed. - Binary crashes with SIGTRAP if you call
start(inputAudioFile:)AFTERinit(inputAudioFile:)— double start.