13 KiB
Pipecat Lag + STT Accuracy — Tuning Plan (Mac mini M4)
For Hermes: Use
subagent-driven-developmentto implement task-by-task. Verify against real trace.
Goal: Cut end-to-end voice turn from ~7–9s to <2.5s for Grace (child voice) and fix Whisper base mis-transcriptions, while keeping the stack local (no new cloud bill) with an optional Apple STT fast-path.
Current trace (from proc_221ead32955d, session 10:01–10:04, Grace):
greeting -> speaking: 0.37sOK. Kokoro synthesize0.235s— not the bottleneck.- Per turn (e.g. "Teach it to be kind"):
- VAD
stop_secs 0.65s+ Silero confidence0.62— OK Whisper baseprocessing 0.21–0.29s, TTFB 0.78–0.94s (speech-end → text) — high for M4- SmartTurn V3 hold: +2.24–2.40s
_idle_watcher: Disabling receiver after 2.3s idle / Discarding old frames+stop_ttfb 0.78s— dominates lag. Second +2–3sappend_audio: End of Turn complete due to stop_secs. Silence in ms: 3000.0on short turns ("Yes.", "Keep going.") — SmartTurn is waiting 3s extra silence before firingUserTurnStop. LLM muse-spark-1.2 via 127.0.0.1:8642:0.93–6.1s(median ~3.8s). Grows with context — context had duplicate greeting, 6 turns, 20k tokens prompt.- TTS chunking OK but sequential per sentence, no streaming.
- VAD
- STT accuracy (Whisper base):
- Good:
and my daughter is here...,My name is Grace, nice to meet you too.✓ - Bad child voice:
I like good stuff(likely "I like princess stuff"?),Princess, tough.(fragmented),go on and do it, inventors, together.(should be "go on an adventure together" — the LLM had to guess),subscribe and watch again Enjoy.(nonsense split, 2 turns merged as one with 4s gap). RNNoise 16000↔48000 resampling,TurnAnalyzer VAD stop_secs differs from recommended 0.2swarning, duplicate greeting in context.
- Good:
Architecture: Keep Pipecat pipeline. Fix in layers: transport/VAD/SmartTurn → STT model/prompt → LLM context/streaming → TTS streaming → observability. All files under ~/Projects/pipecat-voice-gateway/server/.
Tech Stack: Pipecat 1.7, faster-whisper 1.2.1, WhisperSTTService / WhisperMLXSTTService, SileroVADAnalyzer, LocalSmartTurnAnalyzerV3, OpenAILLMService @ :8642, Kokoro :7332 warm + MLX.
Task 1 — Baseline observability (metrics you can read)
Objective: Make lag measurable without grepping DEBUG.
Files:
- Modify:
server/bot.py:1-40— logging + PipelineParams - Create:
server/metrics.py(optional, tiny helper)
Step 1: Enable structured per-turn timing
In run_bot() after PipelineParams, ensure enable_metrics=True, enable_usage_metrics=True already set (is). Add a MetricsLogger or on_metrics handler:
from pipecat.processors.metrics import MetricsProcessor # if exists, else use logger filter
# Add after worker creation:
@worker.event_handler("on_metrics")
async def on_metrics(metrics): # check Pipecat API: worker.on_metrics or PipelineWorker metrics callback
logger.info(f"[metrics] {metrics}")
If API differs, at minimum lower VAD/STT/TTS logger to INFO and add a one-line turn summary in on_client_ready helper that computes speech_stopped -> transcript -> llm_first_token -> tts_first_audio.
Step 2: Persist logs sanely
The current bg proc writes to stdout only; tee /tmp/pipecat-7860.log mask from earlier run hid logs. Ensure server/bot.py launch (both manual and launchctl) writes to StandardOutPath /tmp/pipecat-7860.log and also logger.add("/tmp/pipecat-7860.log") if needed.
Step 3: Verify
Run: tail -f /tmp/pipecat-7860.log | grep -E "TTFB|processing time|EndOfTurn|prompt tokens|kokoro" shows per-turn numbers. Same via curl to public URL.
Commit: feat: per-turn metrics logging
Task 2 — SmartTurn / VAD tuning (biggest win: −2–4s)
Objective: Remove the 2.3s idle + 3s silence hold.
Files:
- Modify:
server/bot.py:145-180(VADProcessor,LocalSmartTurnAnalyzerV3,LLMUserAggregatorParams)
Step 1: Read current Pipecat docs for SmartTurn config Search in venv:
grep -rn "stop_secs\|idle_watcher\|LLMUserAggregatorParams\|TurnAnalyzer" .venv/lib/python*/site-packages/pipecat --include="*.py" | head -n 80
Note: SileroVADAnalyzer(params=VADParams(confidence=0.62, start_secs=0.12, stop_secs=0.65)) vs the warning recommended 0.2s. And LLMUserAggregatorParams(vad_analyzer=SileroVADAnalyzer(... stop_secs=0.8)) is duplicated.
Step 2: Reduce hold
- Set VAD
stop_secs 0.35–0.45s(child speech has longer pauses but 0.65 is too slow):vad = VADProcessor( vad_analyzer=SileroVADAnalyzer(params=VADParams(confidence=0.58, start_secs=0.10, stop_secs=0.40)), audio_idle_timeout=1.0) # was 1.5 - If using
LocalSmartTurnAnalyzerV3, passstop_secsconsistently — do not create two different Silero instances with 0.65 and 0.8. Use one set of params shared with the aggregator, or disable SmartTurn for the fast path:# Option A: disable SmartTurn, rely on VAD only for snappy turn smart_turn = None # Option B: keep it but lower thresholds smart_turn = LocalSmartTurnAnalyzerV3(stop_secs=0.40) # check constructor — if no arg, wrap via LLMUserAggregatorParams(smart_turn_params=...) - Set
ttfs_p99_latencyon STT to match measured value so Pipecat doesn't pad:(measure after tuning; target 0.4–0.6s).stt = WhisperSTTService(settings=WhisperSTTService.Settings(model="base", language="en"), ttfs_p99_latency=0.45)
Step 3: Fix duplicate context turn
In bot.py:192-194 both user_aggregator and assistant_aggregator may double-add greeting. The trace shows assistant message duplicated twice. Deduplicate:
if llm and context is not None:
greeting = "Hey — Hermes on your Mac mini..."
context.add_message({"role": "assistant", "content": greeting}) # once
await worker.queue_frame(TTSSpeakFrame(text=greeting))
Step 4: Verify
Run: speak "Yes." — should trigger LLM within ~0.6s of VAD stop, not 2.8s. Log should show no Disabling receiver after 2.3s idle before turn.
Task 3 — STT accuracy: Whisper base → small / MLX + prompting
Objective: Fix child-voice errors (I like good stuff, go on and do it, inventors).
Files:
- Modify:
server/bot.py:147(STT instantiation) - Modify:
server/kokoro_bridge.py/server/apple_stt_bridge.py(optional Apple path)
Why base is suboptimal: Pipecat Model.BASE = 74M params, trained on adult speech; WER higher on high-pitched child voice, especially at base + CPU int8. The gateway already installs pipecat-ai[mlx-whisper].
Step 1: Try Whisper small (244M, ~3× base) — easiest drop-in
# Cost: ~0.35s vs 0.25s on M4, but accuracy jump for child voice
stt = WhisperSTTService(settings=WhisperSTTService.Settings(model="small", language="en"), ttfs_p99_latency=0.55)
Benchmark: faster_whisper.WhisperModel("small", device="cpu", compute_type="int8") or float16 if available — test offline on a recorded Grace wav before wiring.
Step 2 (preferred on M4): Whisper MLX small or large-v3-turbo
Apple Silicon MLX is faster than faster-whisper CPU:
from pipecat.services.whisper.stt import WhisperSTTServiceMLX
stt = WhisperSTTServiceMLX(settings=WhisperSTTServiceMLX.Settings(model="small", language="en"))
# or: model="mlx-community/whisper-large-v3-turbo" # need to check Model enum supports it
Quantized variant large-v3-turbo-q4 may trade tiny WER for speed — test.
Step 3: Prompt/initial_prompt for domain Whisper respects an initial prompt. Pass kid-domain hints:
stt = WhisperMLXSTTServiceMLX(settings=..., extra={"prompt": "Grace, princess, inventors, adventure, dragon, castle"})
# Check BaseWhisperSTTSettings prompt field vs STTSettings extra
For faster-whisper, verify prompt actually forwards to model — read base_stt.py:148 and stt.py transcribe call.
Step 4: Compare to Apple on-device
server/apple_stt_bridge.py already wraps speech_transcribe_file (SpeechAnalyzer, 6.3× realtime, 2.12% WER) — test it:
# Record 5s Grace utterance to /tmp/grace.wav, then:
python -c "from apple_stt_bridge import transcribe_apple_stt; print(transcribe_apple_stt('/tmp/grace.wav', locale='en-US'))"
If WER < faster-whisper small, implement a hybrid: try Apple first (single MCP roundtrip ~120ms + transcribe), fall back to MLX Whisper. Requires a custom SegmentedSTTService subclass that buffers VAD segment to temp file then calls transcribe_apple_stt — latency ~0.6–1.0s but accuracy wins for kids.
Step 5: Filter non-speech
Set no_speech_prob 0.45 or push_empty_transcripts=False already — but check WhisperSTTSettings(no_speech_prob=...) to suppress the subscribe and watch again hallucination on near-silence. Lower min_volume 0.55 in VADParams.
Verify:
Record 5 utterances from Grace, compare transcriptions across base / small / MLX small / Apple. Pick lowest WER; log latency per model.
Task 4 — LLM latency: context pruning + streaming TTS
Objective: Cut 0.9–6.1s LLM tail.
Files:
- Modify:
server/bot.py:166-180, 198-210(LLM + context + TTS)
Step 1: Prune context
Current context grew to prompt 20k → 20.7k tokens over 6 turns (includes full brain hint). Symptoms: 6s on "subscribe and watch again". Fix:
- Set
LLMContextmax turns (e.g. keep last 8 messages + system, drop older):context = LLMContext(messages=[{"role":"system","content": SYSTEM_PROMPT}]) # On each turn, after add_message, trim if len(context.messages) > 10 - Or use
LLMContextAggregatorPairwithcontext_aggregatorlimits — check Pipecat LLMContext API formax_messages.
Step 2: Stream LLM → TTS (don't wait for full completion)
Current: OpenAILLMService generates full completion (processing 4.7s) then _push_tts_frames synthesizes all sentences sequentially. Switch to incremental:
- Ensure
OpenAILLMServicestreams deltas (LLMTextFrameper token). ThenWarmKokoroTTSServiceshould useTTSServicesentence-based chunking (already viapush_framein Pipecat 1.7 — verifyWarmKokoroTTSService.run_ttsyields per sentence, not per full response). If not streaming, wrap as:# In WarmKokoroTTSService: split on sentence boundaries inside run_tts, yield TTSStartedFrame per sentence - Enable
PipelineParams(enable_metrics=True)already does streaming; confirmon_client_readydoesn't block.
Step 3: Hermes gateway model routing
muse-spark-1.2 via :8642 is correct; check if a faster local model (e.g. qwen2.5:3b via Ollama :11434) could be used for kid chat with sub-500ms first token — compare curl http://127.0.0.1:8642/v1/models vs Ollama latency, but keep Hermes for tool use.
Step 4: Deduplicate greeting token cost Fix Task 2 duplicate greeting — saves ~300 tokens/turn.
Verify:
processing time drops from 5.3s to <1.5s for short turns; TTS first audio appears within 200ms of LLM TTFB.
Task 5 — Audio input hygiene (RNNoise + transport)
Objective: Stop resampling thrash and timeouts.
Files:
- Modify:
server/bot.py:230-245(_make_rnnoise,TransportParams)
Step 1: Lock sample rate
Log shows RNNoise enabling resampling: 16000 <-> 48000. SmallWebRTC negotiates 16k or 48k depending on browser. Set transport explicitly:
TransportParams(audio_in_enabled=True, audio_out_enabled=True, audio_in_sample_rate=16000, audio_out_sample_rate=24000)
So RNNoise knows input rate up front and avoids SOXR per-frame resample jitter (adds ~20ms).
Step 2: Confirm RNNoise not double-instantiated
_make_rnnoise() is called per-transport lambda — good. Don't also create in run_bot.
Step 3: Suppress spurious timeout spam
Timeout: No audio frame received within the specified time. every 2–3s is normal idle but clutters logs. Lower log level for smallwebrtc.transport to INFO, or increase audio_idle_timeout handling in Pipecat — just filter it from metrics.
Task 6 — E2E validation script (repeatable)
Files:
- Create:
server/scripts/bench_stt.py— records 5s, runsbasevssmallvsMLXvsApple, prints WER + latency - Create:
server/scripts/bench_turn.py— drives a WebRTC client and measuresVAD stop -> LLM first token -> TTS first audio
Run:
cd ~/Projects/pipecat-voice-gateway/server
uv run python scripts/bench_stt.py --wav /tmp/grace.wav
uv run python scripts/bench_turn.py --host https://voice.reynafamily.com
# Targets: VAD->STT <0.6s, STT->LLM <0.4s, LLM->TTS <0.3s, E2E <2.0s p50, WER <5% on kid voice
Risks
WhisperMLX smallmodel download ~500MB via HF — cache on first run, may needHF_TOKEN.- Apple STT via MCP is outside Pipecat's segmented pipeline — adds file I/O + HTTP hop; if used, must be benchmarked against local MLX for p99.
- Cutting
VAD stop_secstoo low (0.2s) splits Grace's pauses into multiple turns — test child-specific 0.35–0.45s sweet spot.
Open questions
- Is
muse-spark-1.2the right LLM for kid chat, or should we route kid sessions to a faster 3B with constrained toolset (voiceprofile already is)? - Should kid STT get a separate model (MLX
small) vs adult default (base)? Could switch via diarizer speaker hint.