- Add detailed perf baseline Pi5 (live draft 800ms/8KB, quick ack >30 chars, instant audio Alex<Kokoro, macmini-live-final 753ms vs 11s fallback) - Rank pain points P0-P3: Mac mini SPOF :7331 down now, patch fragility 1675 lines, ground-truth lock boy->voice, disk IO storm, no sentence TTS queue, single VAD vs adaptive FreeFlow effectiveSilenceThreshold, empty trigger - Deep dive mrinalwadhwa/freeflow 91★: PolishPipeline 43 rules, <keep> tags, collapseAdjacent, cleanSpuriousCommas, stripNoise/fillers, knownTerms, guard_against_truncation 7f96ccd, skip 83% 0.55s median, warm backup 91% zero handshake, adaptive silence far=0.001 near=ambient*1.2 capped, DictationPipeline actor detached AVAudioEngine + BT SCO 3s timeout, chunk buffer sentence boundary, SessionTiming - OSS V2V 2025-26: Pipecat 6.5k★ TurnTracker/aggregator/barge-in/e2e tester, LiveKit/TEN, Moshi 7B Mimi codec 12.5Hz 160ms full-duplex dual-stream, Qwen2.5-Omni 2.8-4s — borrow ideas + why stay pipeline text middle = tool-use - 3-layer ground-truth fix: L1 deterministic 0ms, L2 semantic ANE 3B with precedingText + KNOWN_TERMS + alternatives + overlap 0.45 truncate 0.25, L3 Hermes disclaimer + paraphrase confirmation builder build_voice_user_message() - Native M2 target WS /api/voice/live single endpoint spec: binary PCM s16le 16k / empty <1600B trigger / JSON start/stop/ping, events ready→listening→draft volatile 60ms 400ms→transcript final polished+raw+alts→instant_response contextual>30→instant_audio Kokoro 200ms→response_text streaming sentence buffer timed 300ms tail + crossfade→done background, VAD dual 1250ms adaptive, barge-in level>0.12, multi-partial while tool runs - Power truth M2 Air 14,2 macOS 26.5.2: SpeechAnalyzer 0mW ANE 7.3W spike 33% XPC 27-38x RTF 196M EN, Foundation 3B ANE 258-498ms 2GB quant - Phased plan P0→P3 quick wins: disclaimer wrapper + knownTerms + alt transcriptions + guard tune + in-mem ByteArray + adaptive silence + empty trigger + sentence queue + warm daemons + desktop hook ChatGPT parity - Refs: patches, hermes_plugin/__init__.py, pythonwhisper ApplePipeTranscriber pipe volatile 31 drafts/6.3s, apple-speech-api skill ANE traces, HTML diagrams hermes-live-voice-native + oss-comparison + imperfect-fix
34 KiB
Future Optimizations — ESP32 Voice Gateway (M2 native path + FreeFlow learnings)
Goal: WS
/api/voice/livesingle websocket endpoint (compat with current/api/esp32/voice/wsframing binary PCM s16le 16k mono) — instant captioning feedback, fast local LLM initial voice response custom based on partial captioning, empty audio / stream stopped trigger, multi-partials timed to audio tail while big LLM works. Native M2 Air implementation, no Mac mini hop. Keep ESP32 compat by endpoint shape, but design doc is device-agnostic.
This doc captures all learnings from July 2026 audit of Pi5 api_server.py custom gateway (6949 vs 5124 stock, 1675-line patch), Mac mini ANE accelerator (:7331/mcp speech_live_transcribe ApplePipeTranscriber volatile 60ms, apple_llm_quick_reply Foundation 3B ANE 258-498ms, speech_synthesize_base64 Alex <500ms / voicebox_quick_reply Aiden Qwen 1.7B), desktop hooks use-voice-conversation idle→listening→transcribing→thinking→speaking VAD 0.075/1250ms, 60s cap, 3-strike halt, and FreeFlow/macOS SpeechAnalyzer patterns.
0. Current perf baseline (Pi5) — reference
- Live draft: snapshot
pcm_buffer→ in-mem WAV → base64 → HTTP POST to Mac minispeech_live_transcribe— 800ms intervalESP32_LIVE_DRAFT_INTERVAL_MS=800+MIN_BYTES=8000, dedup vialast_draft_text. Eventsready → listening → draft/interim_transcript → thinking{transcribing} → transcript/final → instant_response/quick_reply → instant_audio → response_text → done{background} + thinking{waiting} - Quick ACK: draft >30 chars →
apple_llm_quick_reply {draft, context, instructions}else use finaltranscript[:400]— guard generic "Perfect." triggers regenerate from final - Instant audio:
speech_synthesize_base64 {text, voice:"Alex"}boy voice <500ms preferred, fallbackvoicebox_quick_reply profile:"Aiden"Qwen3-TTS 1.7B 2-4s, pushed via Screen MCPplay_audio_base64 wav_base64 volume80tohttp://<source_ip>/api/mcp(resolved from headerX-Hermes-Screen-Url> query >X-Device-ID=iphone*lookuplocal_devices.yaml> peer_ip) - Final STT: same Mac mini pipe full buffer
macmini-live-final753ms vs local faster-whisper base 11-13s cold fallback - Full turn: persistent
esp32:{device_id}/voice:{id}conversation viaresponse_store, tool-use capable (calendar, memory, kanban, ESP32draw_textetc.) - TTS chain: Edge AriaNeural (Pi default, free 4.3 MOS) mp3→ffmpeg
16k mono s16le WAV - Audit:
~/.hermes/cache/audio/esp32/ + voice_requests.jsonl - Failure mode: raw STT is ground truth — hears "boy" instead of "voice" for voice domain, model keeps wrong transcript, answers about boy mode. Mac mini down (it is now — :7331 refused) kills all live.
- Patch fragility:
api_server.py6949 LOC patch breaks onhermes update— this plugin repo fixes via~/.hermes/plugins/auto-restore.
1. Pain Points — ranked, with origin
| # | Pain | Impact | Where | OSS ref / inspiration |
|---|---|---|---|---|
| P0 | Mac mini MCP single point failure :7331/mcp down now |
No live drafts, no quick ack, no instant audio — full gateway degradation to slow base | Pi5 gateway uses RPC instead of local Swift pipe | FreeFlow alternative: on-device SpeechAnalyzer 26.5 + FoundationModels 3B direct ANE no hop |
| P0 | api_server.py patch 1675 lines wiped by hermes update |
Must restore, flakey 3-strike | Pi5 git stash --include-untracked + pull |
This repo hermes_plugin/__init__.py auto-restore — TODO true plugin injection without file touch |
| P1 | Ground-truth lock: boy → voice kept as truth |
Hermes tool calls wrong thing | f'[Voice input...]\n{transcript}' implies perfect |
FreeFlow never treats raw as final — 3-stage polish + skip heuristic + anti-truncation guard |
| P1 | Disk IO storm: raw_path .s16le temp file per chunk |
60 chunks/min file writes unneeded | Pi5 WS handler writes every frame | FreeFlow + new spec: ByteArray only, one WAV encode at stop — engine_apple_transcribe.py ApplePipeTranscriber |
| P1 | No streaming LLM→TTS sentence queue | Whole answer TTS at end, perceived latency +2s even after ACK | Desktop + ESP32 background | Pipecat sentence aggregator + Cartesia streaming; voice-playback.ts already has takeSpeechChunk /.+?[.!?]/ + 15s stall guard |
| P2 | Single VAD (RMS 0.075) tuned once, not adaptive | near-field AirPods vs far-field built-in built-in speech 0.002-0.005 RMS same threshold → silent presses rejected as speech or vice versa | Pi5 SILENCE_RMS_THRESHOLD 200 vs desktop 0.075 |
FreeFlow effectiveSilenceThreshold() far-field 0.001, near-field max(ambient*1.2,0.0005) capped 0.01; Silero VAD 1MB unified |
| P2 | Empty audio trigger missing as intentional signal | Tap no speech should = "yes?" prompt not dropped | Pi5 len < threshold error path |
User spec: <1600B = 50ms guard treat as immediate final / clarification trigger |
| P3 | PIL UDP spam for screen animations | Network spam per draft | _send_screen_animation per draft |
Skip diff, throttle 200ms + only on state change ready/listening/thinking/speaking |
2. FreeFlow (mrinalwadhwa/freeflow) deep dive — what to steal
Source: https://github.com/mrinalwadhwa/freeflow 91★ Apache-2.0 macOS 14+, dictation-anywhere hotkey app. Not voice assistant, but polish pipeline is gold reusable.
Why it matters for voice gateway:
FreeFlow problem is same: Apple SpeechAnalyzer raw is 2.12% WER 196M EN / 174M ES but robotic, homophones, spurious commas. It never trusts STT. Median keyRelease→injection 0.55s. Warm backup WS pre-opened — 91% zero handshake. 83% skip LLM polish entirely via local heuristic. When polish needed, gpt-4.1-nano 320-780ms (original) but macOS 26 path uses on-device SpeechAnalyzerDictationProvider + FoundationModelChatClient local — same pattern as our ANE 3B.
Code audited July 2026:
FreeFlowKit/Sources/FreeFlowKit/Services/PolishPipeline.swiftpolishModel = "gpt-5.4-nano"— 43 punctuation rules,<keep>tags,collapseAdjacentPunctuation,cleanSpuriousCommas,stripNoisePhrases,stripFillerSounds,buildCloudSystemPrompt(context),knownTerms(Redis, Postgres, macOS… domain terms capitalized),guard_against_truncation(commit7f96ccd Preload models on startup and guard against aggressive truncation2mo ago),word_overlap_ratioconcept,stripKeepTags→\n\n,sanitizeContextField80 charsfocusedFieldContent,toneLabel(for: bundleID) casual- Prompts:
PolishPromptEnglish.swifthuge tested 80-line system prompt: removes throat-clearing (let me think, hold on, how do I put this, let me recall→ keep actual content), self-correctionno wait, actually, sorry, I mean, let me rephrase, never mind, or rather, make that → keep final, abandon restartno, forget it, nah, that won't work → drop before, stutter doublethe the → thebutwait wait wait → Wait, wait, wait,comma keep same sentence,2-item inline never list,3+ items always vertical list with - or 1. 2. 3. if ordered, recapping meeting format- Person: update, numbers always digitsone → 1,minus ten → -10, phone dashes, emailjohn at example dot com → john@example.com, preservekinda/gonna/wannacontractions,<keep>must preserve exactly PolishPromptLocal.swiftsystemPromptLocal = "Clean up this dictated text. Return only the cleaned text."— for FoundationModels 3B ANE short window 4096 tokens guardrailsDictationPipeline.swiftDictationPipeline: PipelineProvidingactor:activate() → recordingimmediate HUD return, audio setup in Detached task (AVAudioEngine 500-900ms + BT SCO 3s timeoutdetachedWithTimeout),pendingContextparallel, adaptiveeffectiveSilenceThreshold()far-field 0.001 vs near-field ambient*1.2 capped 0.01 floor 0.0005 multiplier 1.2, early silence short-circuit peakRMS <= threshold after 200ms = 4 ticks 50ms, streaming session reuse adoptionfreshConnection vs adoptedBackup vs adoptedStaleBackup maxBackupAge 180s, chunk bufferrawChunkBufferuntilendsAtSentenceBoundarythenpolishChunk()+ inject, rolling commit chunkerTimeAndSilenceChunkingStrategyspeechDebounce 10s (only genuine extended silence triggers commit not thinking pauses),ChunkInjectedFlagthread-safe,RecoveryBoxaudio saved for retry,minimumAudioDuration 0.1s,SessionTimingdiagnostic
Borrow list for this gateway:
- Deterministic pipeline (0 ms, no key):
substituteDictatedPunctuation+cleanSpuriousCommas+stripFillerSounds+collapseAdjacentPunctuation— already ported partially in~/Projects/pythonwhisper/freeflow_polish.py(deterministic_polish(),substitute_dictated_punctuation,strip_keep_tags,normalize_formatting,capitalize_known_terms,collapse_adjacent_punctuation,strip_noise_phrases,strip_filler_sounds,clean_spurious_commas) — KNOWN_TERMS currently DB/infra only, add Hermes domain - Skip heuristic 83% rate:
is_clean_enough_to_skip_llm()— if ends sentence + no fillers regex + no repeat(\w+)\s+\1→ skip expensive LLM — medians 0.55s — implement in WS before ANE LLM call - Guards:
guard_against_truncation len(polished)/len(pre) <0.25 → fallback,word_overlap_ratio <0.45 → fallback,match_input_casingleading lowercase preserve when continuing sentence (precedingTextmid-sentenceendsAtSentenceBoundaryfalse) - Chunk buffer: accumulate volatile intermediates
rawChunkBufferuntil sentence boundary, then polish batch — don't polish volatilespeecheech → speechword-level jitter per draft, only on sentence boundary or final — reduces ANE calls 31 drafts → ~5 polishes per 6.3s audio - Adaptive silence: replace fixed 200 / 0.075 with FreeFlow's proximity-aware adaptive — use same for client AudioWorklet level + server VAD — far-field built-in needs 0.001 floor, near-field AirPods ambient measured
- Warm backup: pre-open Swift pipe daemons (
apple-speech-transcribe --pipe volatile+apple-llm-polishJSONL pipe keepingLanguageModelSessionKV warm) — already validatedpythonwhisperApplePipeTranscriberkeep-alivestdin 4-byte BE len + wav→stdout JSONL draft/final 60ms word-level 31 drafts, fixexit(0)hang by return naturally not Process exit,readExactnil on EOF,PYTHONPATH=""isolate PIL_imagingcollision - Context injection:
AppContext80 chars lastfocusedFieldContent+ preceding lines + tone casual → use for polish; for voice gateway = last 2 finalsrefined_context_history deque maxlen2+ conversation summaryesp32:{id}
3. OSS V2V Landscape 2025-26 — where we sit
Family 1 — Pipeline orchestrators (STT → LLM → TTS, text middle = tool-use, modular):
- Pipecat (Daily pipecat-ai/pipecat) 6.5k★: Python frame DAG, Daily/WebRTC transport
VAD → STT → LLM → TTS, context aggregator keeps partials + tool results, interruption via VAD confidence + cancel + preserve context, TurnTrackeruser_start → vad_silence → user_stop → thinking → bot_startenum withheardSpeechRef, token streaming + sentence aggregator + Cartesia/Eleven. Lat 800-1200ms e2e $0.10/min Deepgram+Cartesia. Steal: TurnTracker + aggregator + barge-in cancel TTS keep partial LLM context + parallel tool calls without dropping audio + e2e tester audio file → asserts caption+timeline - LiveKit Agents + TEN (Agora): same frame pipeline, WebRTC room, graph editor Rust core extensions vision. Steal: visual flow state JSON for debugging
ready/listening/captioning/thinking/speaking, SFU recording infra mature but cloud
Family 2 — End-to-end speech models (no text middle, full-duplex UX magic, but poor tool-use, GPU-hungry):
- Moshi — Kyutai Apache 2.0: Mimi codec 12.5Hz, 7B full-duplex inner monologue + 2 audio streams simultaneous (user+bot), 160ms theoretical, 2.2GB quant, A100/CUDA needed, weak function calling hallucinates. Steal: dual-stream idea — keep STT open while TTS playing via
echoCancellation:true+ second AnalyserNode RMS >0.12 while speaking → half→full duplex for free - Qwen2.5-Omni / Mini-Omni2 / Ultravox / Freeze-Omni: Whisper encoder + LLM + soundstorm codec audio in → thought text → audio out, raw audio encoder brief → no STT errors emotion speaker id, but 3-5s TTFT, large VRAM, needs fine-tune for tools. Note: text middle needed for Hermes because calendar/memory/kanban tool calling via Claude Spark > omni accuracy; skip STT path would need audio→LLM direct that can't yet call tools reliably
Our position — hybrid unique 2026:
- Only project with Apple Speech 26 pipe volatile 60ms 2.12% WER 196M EN + Foundation 3B on-device ANE 258-498ms both local — FreeFlow validated 27-38x RTF 0mW ANE 7W spike single-core XPC
localspeechrecognition - Fast voice from partials contextual not generic — no OSS does this; Moshi inner thought generic "yeah", Pipecat none by default, we have draft >30 chars → Foundation quick ack
"On it — checking your calendar for tomorrow 3 meetings…"already measured - Timed multi-partials while tool runs — Pipecat usually waits all tools then speaks; we want instant → progress "found 3 events" → final timed to audio tail queue sentence buffer
/.+?[.!?]/+ start next in last 300ms prev tail + crossfade 150ms - In-mem only, no ffmpeg, no temp wav per chunk — most OSS still writes file per chunk; we already have
ByteArrayonly path in spec - WS compat
WS /api/voice/livesame binary PCM s16le 16k framing as Pi5 ESP32 WS — new design says "keep endpoint with websocket" compatible but diagrams show no ESP32/other devices focus on MacBook only — so web can send WebM/opus or PCM flag, backward compat for future hardware - Offline $0 7W vs cloud $0.10/min + net jitter
- Latency budget native: VAD 1250ms + fast LLM 300ms + Kokoro 200ms = 1750ms first audio, caption 60ms instant feels <1s streaming, big LLM token from 200ms streaming sentence queue. vs Pipecat ~2200ms cloud, Moshi ~600ms magic but no tools, Qwen-Omni 2.8-4s
4. Ground-truth lock fix — 3-layer defense (key feature request)
User worry: Pi hears "boy instead of voice" and model keep it as ground truth. FreeFlow fix is not trust acoustic as truth but semantic context.
Before (current):
user_message = f'[Voice input from ESP32 device "{device_id}"]\n{transcript}'
=> Hermes thinks typed ground truth = boy mode.
After — 3 layers:
L1 deterministic 0ms: deterministic_polish() from freeflow_polish.py — strips fillers/noise, collapses ,,,, Apple spurious commas — boy still boy here
L2 semantic ANE 3B with context (fixes boy→voice):
precedingText = last 2 finals + voice:{id} summary(FreeFlowrefined_context_history deque 2)- Prompt includes domain
KNOWN_TERMS += voice, gateway, hotkey, dictate, Kokoro, VAD, barge-in, ANE, SpeechAnalyzer, DictationPipeline, FreeFlowKit - FreeFlow
systemPromptEnglishfull (80 lines tested) — self-correctionI mean…→ keep final, etc. - ANE 3B
polish_line(substituted, prev1, prev2)→ if context = "voice gateway setup", "boy" acoustically close → semantic mismatch → corrects "voice" - Guards: truncate 0.25 / overlap 0.45 / casing / keep tags
- Alternatives: Apple
SpeechTranscriberattribute options[.alternativeTranscriptions]returns list — already parsed in pipe--pipeJSONLalternatives— pass to LLM:Alternatives: boy mode (0.4), voice mode (0.6)— model picks semantic fit
L3 Hermes LLM disclaimer (your ask) + multi-partial ask confirm:
Build user_message with explicit uncertainty + polished + raw + alternatives + instruction:
[Voice input — speech-to-text, may contain errors]
This was spoken, not typed. Homophones like
boy/voice, four/for, new/knew, dictate/dictation
are common STT errors. Don't treat transcription
as ground truth.
Polished transcript (best guess, 1 edit, overlap 0.83):
"Set voice mode"
Raw STT alternatives (low confidence word jitter):
• "set boy mode" (raw final)
• "set voice mode" (draft volatile alternative, 3x seen)
Draft volatile instability is normal (speecheech →
speech final) — don't treat as rewrite.
Context: user was talking about voice gateway setup
for MacBook native, previous turn: "add optimization
doc with FreeFlow ref".
If request unclear or semantically off in voice gateway
context, paraphrase back what you understood and ask for
confirmation before acting. Don't hallucinate tool call
on shaky transcript. Example: "Just to confirm — you mean
voice mode?"
User originally said (as heard / best):
"Set voice mode"
Ephemeral system add for whole voice:{id} session: "Voice session: STT imperfect 2.12% WER even polished ANE. If confidence low or term out-of-domain, clarify." + Hermes domain bias list.
This makes big LLM say "Just to confirm — you mean voice mode?" if still ambiguous instead of silently wrong. Fast ACK uses same polished transcript so <1.7s audio also correct.
5. Native M2 Air target — single WS endpoint spec (user final ask)
Endpoint: WS /api/voice/live (keep shape compat with Pi5 /api/esp32/voice/ws for future ESP32, but current design MacBook only focus)
Transport:
- Client → Server: binary
PCM s16le 16k monochunks 20ms AudioWorklet (replacesScriptProcessor), OR empty binary (<1600B = 50ms guard = trigger), JSON control{start:{format:"pcm_s16le"|"webm", sample_rate:16000, locale:"en-US"}, stop, ping} - Server → Client:
{ready} → listening{threshold} → draft volatile{draft,isFinal:false,is_volatile, alternatives, level} every ~400ms → interim_transcript → thinking{transcribing} → transcript final{polished,true, overlap, raw, alts, source:"apple" } → instant_response{quick, draft_count} contextual fast ack → instant_audio{b64/wavBase64, voice, quick_ms,speak_ms} → response_text streaming token via sentence buffer → audio_chunk{audio_b64, text, seq} timed to audio tail (start N in last 300ms of N-1 + 150ms crossfade) → done{background, instant_ok, keep_instant} + thinking{stage:waiting}
Instant caption feedback:
- Apple
SpeechAnalyzer --pipe volatilekeep-alive single process, stdin 4-byte BE len + wav bytes len 0 = EOF, stdout JSONL draft/final word-level per chunk (~60ms granularity, 31 drafts/6.3s validated, EN 196M ES 174M),bestAvailableAudioFormat+AVAudioConverter+AudioBufferQueueactor, report[.volatileResults, .alternativeTranscriptions]attribute[.audioTimeRange, .transcriptionConfidence]. Volatile jitter normal. - In-mem only
ByteArrayaccumulator — zero disk, deduplast_draft_textsame as Pi5_do_live_draftbut no file. Composer desktop shows italic .55 live while listening, solid on finaltranscript
Trigger:
- VAD 1250ms dual
effectiveSilenceThresholdFreeFlow adaptive (far-field built-in 0.001, near-field max(ambient*1.2,0.0005) capped 0.01) — same threshold client AudioWorklet RMS AnalyserNode + server —heardSpeechRefpattern fromuse-mic-recorder+ Pi5 3-strikeCONTINUOUS_NO_SPEECH_LIMIT=3halt - Empty audio trigger: binary <1600B or
{start}then immediate{stop}→ freeze draft → final transcript immediately → if still <2 words → prompt "yes?" clarification via fast ACK, not error — user spec
Fast local LLM initial voice response custom based on partial captioning:
- Foundation 3B ANE (reuse
apple_speech/.build/release/apple-llm-polishpipe JSONL{"id","mode":"line"|"paragraph","text","prev1","prev2","context","prevSource","language"}→{"id","ok","text","ms"}KV cache warm, 258-498ms line polish, temp 0.1 line / 0.2 parapermissiveContentTransformationsguardrails, assetcom_apple_MobileAsset_UAF_FM_GenerativeModels+ Overrides 2GB quantized, ANE+GPU) — contextual: draft+final+context → "On it — checking your calendar for tomorrow, 3 meetings…" not generic "one sec" — must be >30 chars else use final[:400] for context else regenerate if generic "Perfect." - Spoken instantly: Kokoro ONNX 82M standalone
kokoro==0.9.4voice=param required, model cache~/.cache/huggingface/hub/hexgrad_Kokoro-82M326MB 4-5x RTF 0.36-0.71s inference load once 4-6s daemon resident, MPS 19x SLOWER LSTM must CPU, Edge AriaNeural cached 4.3 MOS fallback,voicebox_quick_replyAiden Qwen 1.7B 4.4 MOS if kokoro down,say -v Eddynot Siri 30MB MOS 3.8 don't use. Timed: play fast ACK <500ms Kokoro while big LLM already started parallel — don't wait fast finish to start big
Big LLM while fast response playing, multiple partials timed to playback:
voice:{device_id}persistent conversation,response_store, ephemeral system disclaimer- Streaming:
POST /api/sessions/{id}/chat/streamSSEdeltaexisting — after final transcript, submit polished transcript + disclaimer + alternatives + preceding context to big LLM (Claude Spark etc) streaming tokens → sentence buffer/.+?[.!?。!?]/soft boundary, ; :after 180 chars buf>220 (same as Pi5takeSpeechChunk) →playSpeechText(chunk, {source:'voice-conversation', messageId})queuePLAYBACK_STALL_MS=15000sequence guardstopVoicePlayback() - Multi-partial behavior: intent ack already spoken (fast), big LLM can send progressive updates while task runs:
"On it — checking your calendar…"(fast) →"Found 3 events tomorrow"(progress while tool runs) → full final. Timed to tail: schedule next audio start in last 300ms of previous Audio element tail — requires audio duration from WAV header or estimate, crossfade 150ms, no overlap - Tools: calendar fetch iCloud, memory, kanban etc — Pipecat usually waits all tools then speaks; we stream partials as tool results arrive via
tool_progress_callback("tool.started"/completed)→ inject as contextual voice update "found 5 events" etc
Barge-in:
- Client
AnalyserNodeRMSlevel>0.12whilespeaking→stopVoicePlayback()+startListening()→ server cancel queue viaping cancelor{barge_in}event — Pi5 solved via_tts_playingEvent + 300ms gap + cancel recorder beforeplay_audio_file, desktop pause analyser during speaking to avoid self-trigger but keep second node for barge - Server also
AnalyserNodeechoCancellation true to allow dual-stream half→full duplex for free (steal from Moshi idea)
Power/perf truth (measured M2 Air Mac14,2):
- SpeechAnalyzer: 0mW ANE in 50 powermetrics samples, 33.5% CPU single-core
localspeechrecognition XPC442mW idle → 7.3W spike E+P, GPU 67-193mW, 27-38x RTF file bench 18.9s→0.6s RC 31x pipe, RTF 31 drafts +2 finals, asset per-locale 174-196M EN 196M total 1.4G 4 locales preinstalled path/System/Library/AssetsV2/com_apple_MobileAsset_UAF_...Whisper Base 140M single multi-lang but venv Torch 800M — system footprint win no venv but heavier per-locale - Foundation 3B: ANE+GPU ~258-498ms measured LM, 2GB quant, 4K context offline no key
- LLM log
logs/llm_requests.jsonl→provider:apple model:apple-fm apple_ms 351-498vs Ollama 4-8s - Missing log bug cause Apple LLM not wired fixed
2e6d6d2
6. Implementation plan — low-hanging wins ordered
Phase 0 — this plugin repo already done:
- Pi custom gateway persisted as plugin
~/.hermes/plugins/esp32-voice-gatewayauto-restore bundled patch, verifies 400 not 404
Phase 1 — polish layer (today, 80 lines reuse existing code):
hermes_plugin/future/freeflow_polish_port.py— copyfreeflow_polish.pydeterministic + skip + guards +KNOWN_TERMS += Hermes domain voice/gateway/hotkey/Kokoro/VAD/ANE/SpeechAnalyzer, unit test boy→voice case with precedingContextvoice gateway setup- Swift pipe daemons local — already built:
apple_speech/.build/release/apple-speech-transcribe --pipe+apple-llm-polish— M2 Air native, daemon pattern model resident 0.36s inference after load - Alternatives in prompt — enable
.alternativeTranscriptionsreport option, passalternativesarray in draft event{"event":"draft","text","alternatives":["boy mode","voice mode"]} - Guard tuning — overlap 0.45 not 0.7 to allow boy→voice 1-word edit, truncate 0.25 fallback
- Composer opacity .55 italic live draft — no changes to Pi5 — future desktop
- Skip PIL UDP spam — throttle screen animation to state change only
- In-mem pipe — remove
raw_paths16le temp files, one WAV at STOP_s16le_to_wav_bytessame as Pi5 but no file per chunk, bytearray copy every 800mssnap = buffer.copy()to avoid race
Phase 2 — native /api/voice/live plugin proper (replace patch):
Create hermes_plugin/plugin_new/ true standalone plugin injecting routes at runtime without touching file — hook APIServerAdapter.connect — mount POST /api/voice + WS /api/voice/live reusing helpers _s16le_to_wav_bytes, _write_pcm_s16le_wav, _resolve_esp32_screen_url verbatim + audit logger, but using native Apple pipes instead of Mac mini MCP URLs from MACMINI_MCP_URL/TOKEN env / config.yaml mcp_servers.macmini
Phase 3 — streaming sentence TTS timed queue:
Port takeSpeechChunk regex + 15s stall guard from voice-playback.ts to server audio_chunk events timed 300ms before prev tail end, sentence buffer while LLM token streams, stopVoicePlayback() barge-in level>0.12
Phase 4 — Hermes disclaimer + confirmation behavior:
Add build_voice_user_message() wrapper that injects STT imperfect disclaimer + raw + polished + alts + domain context + instruction paraphrase back & ask if ambiguous — apply to both ESP32 existing POST /api/esp32/voice + new /api/voice/live — minimal change reduces hallucination risk immediate win
Phase 5 — desktop use-voice-conversation-live hook ChatGPT Live parity:
use-mic-recorder.ts expose PCM chunks via AudioWorklet 20ms + AnalyserNode RMS, WS client connect ws://localhost:8642/api/voice/live, live draft composer, streaming SSE response_text → sentence TTS queue per sentence, barge-in via stopVoicePlayback() + startListening(), continuous loop 3-strike silence halt, waveform pill level*100% width, Press Space to send, header dot green listening yellow transcribing blue thinking purple speaking
Phase 6 — testing harness:
Pipecat/LiveKit e2e tester pattern: audio file → WS → asserts caption + audio timeline regression, plus boy→voice specific test case with context
7. Latency target vs OSS
- Hermes M2 Native target (this spec): caption 60ms volatile ANE instant confidence, first sound 1750ms (1250 VAD + 300 ANE ack + 200 Kokoro) but feels <1s because caption at 60ms, streaming token from 200ms sentence N speaks while N+1 gen, barge-in 12ms RMS detect
- Pipecat Cloud Deepgram+Cartesia: ~800-1200ms e2e streaming but no contextual fast ack, $0.10/min, net jitter, 350ms partials
- Moshi 7B full-duplex: 160ms codec theoretical + 400ms tokens → 600ms magic UX but needs A100/CUDA 30W GPU, weak tool-use, no caption (needs separate ASR), not offline M2
- Qwen2.5-Omni: 800ms encoder + 1.2s prefill → 2.8-4s first audio, good quality not streaming sentence TTS, 12GB VRAM
Winner for Hermes tool-use: native pipeline with text middle auditable transcript + freeflow deterministic readable, fast contextual voice unique.
8. References — in this repo + external
Inside this repo:
patch/api_server_esp32_live_draft_with_llm.patch— 1675-line current running code Pi5 verified 2026-07-14, origin~/.hermes/hermes-agent/gateway/platforms/api_server.py+1500 customhermes_plugin/__init__.py— auto-restore plugin logicdocs/*.patch— evolution reference
External open source:
- mrinalwadhwa/freeflow https://github.com/mrinalwadhwa/freeflow — Apache-2.0 91★ macOS dictation anywhere, 288 commits, last 7f96ccd guard truncation.
FreeFlowKit/Sources/FreeFlowKit/Services/PolishPipeline.swift+DictationPipeline.swift+ PromptsPolishPromptEnglish.swiftlocal/casual/qwen/hindi/kannada/tamil, ProposalsOpenAIRealtimeProvider.swiftWS WSS streaming backup warm standby 180sSessionTiming, adaptive silenceeffectiveSilenceThreshold(). BENCHMARK.md median 0.55s keyRelease→inject 83% skip LLM. mic compatibility issue #2. - Pipecat https://github.com/pipecat-ai/pipecat — 6.5k★ Python frame processors, frame DAG, context aggregator, TurnTracker, barge-in cancel preserve context, Daily transport — steal TurnTracker enum + aggregator + testing harness
- LiveKit Agents + TEN — WebRTC SFU agents graph editor — steal flow state JSON visualization
- Moshi Kyutai https://github.com/kyutai-labs/moshi — 7B full-duplex Mimi codec 12.5Hz 2.2GB quant 160ms — steal dual-stream while TTS playing keep STT open echoCancellation true half→full duplex free
- Qwen2.5-Omni / Mini-Omni2 / Ultavox — audio LLM raw encoder → LLM → codec, no STT errors but high TTFT weak tools
Local prior art:
~/Projects/pythonwhisper/—freeflow_polish.pydeterministic_polish()port,engine_apple_transcribe.pyApplePipeTranscriberpipe volatile 31 drafts/6.3s 60ms granularity,engine_apple_llm.pyAppleLLMreader thread + pending-id queues ANE 258-498ms,engine_llm.pyrolling paragraph refinedeque maxlen2post_correct_line()Apple path primary Ollama fallback +logs/llm_requests.jsonl,engine_kokoro_tts.py, SQLitetranscribe.py61K~/Projects/pythonwhisper/apple_speech/— Swift CLIapple-speech-transcribe --pipe,apple-llm-polish120KB FoundationModels 3B pipe JSONL KV warm, build workaroundswiftc -O -parse-as-library -target arm64-apple-macosx26.0 -sdk .../MacOSX26.5.sdk Sources/... -framework Speech -framework AVFoundation -framework Foundation(Package.swift v26 needs PD 6.2 but CLTools SPM 6.0), ANE power traceane_power 0mW 7.3W CPU 33% XPC, assets EN 196M, benchmarks Inscribe 2.12% clean vs Whisper Small 3.74% — see~/.hermes/skills/apple-speech-api/SKILL.md~/.hermes/skills/hermes-live-voice/— full audit refspi5-voice-gateway-audit.mdws-protocol.mddesktop-live-voice-plan.mdnative-m2-websocket-only.mdoss-v2v-comparison-2026.md- HTML diagrams:
~/hermes-live-voice-native.html(MacBook-only WS-only target) +~/hermes-voice-oss-comparison.html+~/hermes-voice-gateway-design.html(Pi+Mac mini pain points) +~/hermes-voice-imperfect-fix.html(boy→voice 3-layer)
Protocol docs:
- Binary
pcm_s16le 16k mono, WS controlstart{format,sample_rate}/stop/ping, Events orderready → listening → draft/interim_transcript → thinking{transcribing} → transcript/final → instant_response/quick_reply → instant_audio → response_text → done{background} + thinking{waiting}, HeadersX-Hermes-Screen-Url,X-Device-ID=iphone*→local_devices.yaml, peer_iphttp://{ip}/api/mcpaliaseslittle32/esp32_screen, MCP toolsclear_screen,draw_text,play_audio_base64,play_mp3_base64, animationshermes_voice_animations.py ack/captioning/waitingESP32_VOICE_ANIMATION_PATH - Audio tools:
tools/voice_mode.pyrecorder RMS 200/3s hallucination filter,tools/tts_tool.py300+ voices Edge/Eleven,hermes_cli/voice.pyCtrl+B push-to-talk continuous VAD 3-strike
9. Quick wins TODO checklist
- P0 Add
docs/future-optimizations.md(this file) — done - P0 Add
build_voice_user_message()wrapper with disclaimer + alternatives + confirmation instruction — minimal code, biggest reliability gain - P0 Add Hermes domain
KNOWN_TERMSvoice/gateway/hotkey/Kokoro/VAD/ANE - P1 Enable
.alternativeTranscriptionsin Apple pipe + emit in WS draft/final so LLM sees homophone candidates - P1 Tune guard
overlap 0.45 truncate 0.25to allow boy→voice single-word fix - P1 In-mem only ByteArray no file per chunk + snap copy every 800ms race fix
- P1 Adaptive silence
effectiveSilenceThreshold()far-field 0.001 near-field ambient*1.2 cap 0.01 - P1 Empty audio <1600B trigger as clarification not error
- P2 Sentence TTS queue timed to tail 300ms start before tail + 150ms crossfade stall 15s
- P2 Warm backup Swift pipe daemons model resident 0.36s inference
- P2 Barge-in
level>0.12 && speakingstop TTS + startListening + server cancel - P2 Multi-partial while tool runs — instant ack → progress found N → final, stream via tool_progress_callback
- P3 Skip PIL UDP spam throttle to state change
- P3 Desktop hook
use-voice-conversation-liveChatGPT parity — AudioWorklet 20ms + live draft italic .55 + waveform + header dots - P3 True standalone plugin without file touch (eliminate patch reapply) + integration test curl 400 not 404 + boy→voice e2e test
10. Open questions / decisions
- TTS choice M2 native: Kokoro 82M standalone
kokoro==0.9.4imperative 200ms CPU daemon vs Edge AriaNeural 4.3 MOS cached vs Eleven fallback? Start Kokoro daemon (copy 325MB models from Pi5~/kokoro-models/kokoro-v1.0.onnx+ voices.bin) butmlx-audiowrapper needsmisaki[en]+spacy+pydanticconflicts — use standalonekokoropip notmlx_audio. MPS is 19x SLOWER LSTM keep CPU. - Foundation 3B prompt for fast ack: should we reuse
systemPromptLocal = "Clean up… Return only…"for polish but new quick reply promptdraft + final + "quick ack under 20 words contextual on intent"with low temp 0.1/0.2? Already measured Foundation general + contentTagging UseCases ~2GB quantized ANE+GPU. - WS compat long-term: keep
/api/esp32/voice/wsalias for ESP32 firmware already flashed or only/api/voice/live+ adapter route? Current plugin runs both by patch; future true plugin could mount both aliases. - FreeFlow license reuse: Apache-2.0 clean, prompts can be reused with attribution — ideal for private repo (Reyna family private) not selling.
- Ephemeral vs stored transcript: Should voice gateway keep both raw + polished in audit for metrics, but only polished + disclaimer in
voice_requests.jsonl? Recommended: log raw, polished, alts, overlap, guard decisions for future tuning.
Generated 2026-07-15. Sources: Pi5 live gateway audit, Mac mini MCP 37 tools, pythonwhisper ANE traces 0mW ANE 7W CPU, FreeFlow 7f96ccd, Pipecat/Moshi/Qwen OSS landscape.