Files
whisper-translation/README_APPLE_V3.md
T
Adolfo Reyna 80f0bf309f feat: Apple Speech v3 + Freeflow polish + draft streaming
- Apple SpeechAnalyzer (macOS 26+) binary: --bench (31x RTF), --pipe
  (persistent process, 150ms finals), --live (word-by-word drafts)
- Pipe protocol: 4-byte BE length + wav payload, emits JSONL
  {event:draft|final, text, isFinal, chunk} — 31 drafts for 6s audio (~60ms granularity)
- engine_apple_transcribe.py: ApplePipeTranscriber with
  transcribe() + transcribe_with_draft_callback(), VAD + draft
  queue, new flags --apple-stream (on), --apple-stream-interval,
  --apple-pipe (on). Fixes PIL/transformers import crash by lazy import.
- main_v3.py: engine selector {whisper,apple}, passthrough translate
  when no -es/-fr/-ar, freeflow flags same as v2
- Freeflow polish: deterministic punctuation commands (comma,
  question mark, new paragraph, at sign), filler stripping,
  <keep> protection, skip-clean heuristic, freeflow/qwen/legacy
  prompt styles. Much better final readability vs raw Apple/Whisper.
- main_v2.py, engine_llm.py, engine_distribute.py: integrate freeflow
- bench: Apple 2.12% WER vs Whisper Small 3.74% (Inscribe), CPU
  0mW ANE (measured via powermetrics), 196M EN cryptex per locale.
- Verified: 31 word-by-word drafts, 2 finals, exit 0, bench regression ok.

Freeflow still much better for final polish — Apple wins on speed
and raw accuracy, freeflow wins on readable paragraph output.

Co-authored-by: internal-model
2026-07-13 21:08:18 -04:00

5.0 KiB

Apple Speech API v3 — Real-time transcription using macOS 26's new engine

Based on Inscribe's benchmark: Apple's SpeechAnalyzer + SpeechTranscriber (iOS/macOS 26) hits 2.12% WER on LibriSpeech test-clean vs Whisper Small's 3.74%, at ~3x speed. No performance numbers from Apple until this blog.

What's in v3

Two implementations:

1. Swift CLI (apple_speech/)

Tiny Swift 6 package compiled directly (no SPM version dance):

cd apple_speech && bash build.sh

Modes:

  • apple-speech-transcribe <file> --locale en-US → JSONL segments, timestamped
  • apple-speech-transcribe --bench <file> → single JSON with RTF, full transcript
  • apple-speech-transcribe --live --locale en-US → mic streaming JSONL (volatile + final)
  • apple-speech-transcribe --check → availability + installed/supported locales
  • apple-speech-transcribe --list-locales

Correct API usage discovered by testing against crash logs:

  • File: SpeechAnalyzer(inputAudioFile: file, modules: [transcriber], finishAfterFile: true) then for try await result in transcriber.results — results loop terminates naturally, not via task cancellation. This is the path that emits final segments correctly; analyzeSequence(from:) returned early and missed finals.
  • Live: AudioBufferQueue actor + LiveInputSequence: AsyncSequence<AnalyzerInput> fed to SpeechAnalyzer(inputSequence:modules:) with AVAudioEngine tap + optional AVAudioConverter to bestAvailableAudioFormat.

2. Python engine (engine_apple_transcribe.py)

Drop-in for engine_transcribe.py:

  • Same Silero VAD chunking (silence, max_buffer)
  • Each VAD-finalized chunk → temp 16-bit WAV → apple-speech-transcribe --bench subprocess
  • Emits dict {original, en_bridge, detected_lang, speaker, ts, _apple_meta{rtf, segments}} same shape as whisper engine
  • Queue-compatible with full pipeline

3. Main v3 entry (main_v3.py)

--engine {whisper,apple} selector:

./venv/bin/python main_v3.py --engine apple --lang en -v
./venv/bin/python main_v3.py --engine apple --lang es -v
./venv/bin/python main_v3.py --engine whisper --lang en   # old path
./venv/bin/python main_v3.py --apple-bench-file /tmp/audio.wav --lang en

Benchmark results (this machine, M-series, macOS 26.5.2)

Generated via benchmark_v3.py --test-say (TTS roundtrip):

Locale Text WER RTF
en-US Hello world, this is a test... 0.0% 28.4x
en-US quick brown fox + $35.99 normalization 38.9%* 30.7x
en-US email to john@example.com 14.3%* 23.7x
es-ES Hola mundo... 0.0% 14.6x
es-ES Buenos días... 0.0% 38.4x
fr-FR Bonjour... (voice mismatch) 92.9%
de-DE Hallo Welt... 8.3% 15.2x

*WER inflated due to smart normalization ($35.99, john@example.com) — actually better output.

Real claim from Inscribe: 2.12% vs Whisper Small 3.74%, ~3x faster, on LibriSpeech 5,559 utterances, M2 Pro 32GB macOS 26.5.1.

Locales on this machine

After auto-download triggers:

installed: de_AT de_CH de_DE en_* es_CL es_ES es_MX es_US fr_BE fr_CA fr_CH fr_FR (20)
supported: + it_CH it_IT ja_JP ko_KR pt_BR pt_PT yue_CN zh_CN zh_HK zh_TW (30 total)

Install is lazy: first transcribe in a locale triggers AssetInventory.assetInstallationRequest.

Architecture comparison

  • v2 Whisper: audio -> Silero VAD chunks -> mlx-whisper (460MB Small) -> freeflow_polish -> LLM paragraph -> MarianMT -> ingest
  • v3 Apple: audio -> Silero VAD chunks -> temp WAV -> apple-speech-transcribe SFSpeechAnalyzer (system model, ~few 100MB on-demand) -> same LLM/MT pipeline
  • Future: native streaming inputSequence path from Python without temp files (requires PyObjC or Swift extension)

Files

  • apple_speech/Sources/AppleSpeechCLI/main.swift — Swift binary
  • apple_speech/build.sh — swiftc build (bypass SPM .v26 manifest issue)
  • engine_apple_transcribe.py — Python VAD + subprocess wrapper
  • main_v3.py — engine selector main
  • benchmark_v3.py — quick TTS bench
  • config_v3.json — example config with engine=apple

Next steps to go production

  • Mic streaming without temp files (PyObjC bridge or keep long-lived Swift daemon)
  • Speaker diarization (pyannote still works same)
  • WER eval on LibriSpeech subset
  • CI for binary rebuild on macOS version bump

Gotchas discovered

  1. Package.swift with .macOS(.v26) needs PackageDescription 6.2+ but CLT's swift-pm is 6.0 — must compile via swiftc directly.
  2. SpeechAnalyzer's init(inputAudioFile:) vs analyzeSequence(from:) have different lifetime: the former finishes when results AsyncSequence ends, the latter returns before final results (requires extra sleep + task cancellation). Fixed by using file initializer.
  3. Silence-only audio returns 0 segments, not error.
  4. Assets auto-download on first transcribe; status transitions supported -> downloading -> installed.
  5. Binary crashes with SIGTRAP if you call start(inputAudioFile:) AFTER init(inputAudioFile:) — double start.