Files
VoiceAgent/README.md
T

21 KiB
Raw Blame History

Local voice conversation with Claude Code

Talk to Claude out loud. Speech recognition and speech synthesis run entirely on this Mac; only Claude itself is remote, reached through the same claude CLI and the same auth a terminal session uses.

mic → hold fn → SpeechTranscriber (macOS 26) → Claude Agent SDK → Kokoro → speakers
Stage What runs Where
Turn taking hold the fn key, any app local
Speech to text SpeechTranscriber, macOS 26's on-device model local
The conversation Claude Code, with tools Meta AI Gateway
Text to speech Kokoro (kokoro-onnx) local
Audio I/O portaudio via sounddevice local

Usage

cd ~/voice-agent
./talk

Hold the 🌐 fn key and speak; release to send. This works whatever app has focus, so you can keep working while you talk. Speaking again cuts Claude off mid-sentence. Ctrl-C to stop.

If Input Monitoring isn't granted (see below), it falls back to SPACE to start talking, SPACE again to send — which also stays available as a backstop even when the hold key is working.

Useful flags:

./talk --llm-engine hermes            # Hermes LLM engine (default)
./talk --hermes-model hermes-3        # specify Hermes model (default hermes-3)
./talk --llm-engine apple             # local Apple Silicon MLX model
./talk --llm-engine claude            # Claude Code CLI engine
./talk --list-devices                # see microphones and speakers
./talk --list-voices                 # see macOS system voices
./talk --input-device "Adolfo i16"   # pick a device by index or name substring
./talk --tts apple --voice Moira     # Irish macOS voice instead of Kokoro
./talk --voice am_michael            # any Kokoro voice id
./talk --hold-key right-option       # a different hold key, or 'none'
./talk --voice-activity              # hands-free instead of push-to-talk
./talk --claude-model claude-opus-5  # trade latency for capability
./talk --allow-writes                # give Claude Edit, Write and Bash too
./talk --cwd ~/some/project          # work in a specific project directory
./talk --load-settings               # load your ~/.claude plugins and skills
./talk --log-level DEBUG             # watch the frames flow

If something misbehaves, run the self-test first — it isolates the broken stage instead of making you read pipeline logs:

.venv/bin/python selftest.py

First run

Kokoro downloads about 350 MB of voice models into ~/.cache/pipecat the first time. Apple's recognizer needs no download.

macOS will need microphone permission. The first ./talk should prompt. If it doesn't, or if the self-test reports "captured pure silence", add your terminal under System Settings → Privacy & Security → Microphone.

The speech model downloads itself the first time a locale is used.

macOS AirPods route switching checklist

When no audio device is pinned, VoiceAgent follows macOS's current default microphone and speaker while a conversation is running. Use this checklist when validating AirPods or another Bluetooth headset:

  1. Connect AirPods before starting ./talk; confirm that both microphone and speaker audio use them.
  2. Start a conversation on the built-in microphone and speakers, then connect AirPods. Confirm that each available direction moves independently to the new macOS default without ending the conversation.
  3. Disconnect AirPods during a conversation. Confirm that the available input and output return to the macOS defaults and that the conversation remains usable.
  4. In Control Center or Sound settings, explicitly switch back to the built-in microphone and speakers while AirPods remain connected. Confirm both routes follow those defaults.
  5. Repeat connect, disconnect, and manual default changes several times to catch delayed Bluetooth profile changes or a stale route listener.

Run this validation without --input-device or --output-device: omitted values intentionally follow macOS defaults. Supplying either flag pins only that direction (for example, a pinned USB microphone still allows an unpinned output to follow AirPods); pinning both directions disables automatic route following. Use ./talk --list-devices to identify a device by index or name substring.

This behavior depends on macOS Core Audio notifications and the device profile currently exposed by Bluetooth. A headset may briefly expose only output, or reject the negotiated sample rate while switching profiles; VoiceAgent keeps the previous working route when a replacement cannot open. The terminal/app still needs macOS Microphone permission, and the global hold-key needs Input Monitoring. A sandboxed packaged app also needs the appropriate microphone usage description and audio-input entitlement; those platform permissions cannot be granted by route switching code.

Turn taking, and why it's push-to-talk

Detecting the end of a turn by listening for silence is both slow and wrong here. Slow, because Silero waits 0.2 s of silence and the turn strategy waits another 0.6 s in case you resume — about 0.8 s of dead air on every turn, before Claude has even been asked. Wrong, because with no acoustic echo cancellation the microphone hears the speakers, so Kokoro's voice gets transcribed and Claude answers himself.

A key fixes both at once. The turn ends the instant you say it does, and audio outside a keypress is never transcribed, so self-hearing is structurally impossible rather than merely guarded against.

Holding fn is the default because a Quartz event tap sees it from any app, and — unlike a terminal, which only ever receives key presses — reports releases too, making real hold-to-talk possible. It needs Input Monitoring for your terminal, under System Settings → Privacy & Security → Input Monitoring. Restart the terminal afterwards.

Two things worth knowing about the fn key specifically. If it opens the emoji picker, set System Settings → Keyboard → "Press 🌐 to" → Do Nothing; the tap is listen-only by design, so it observes the key without stealing it from whatever else you have bound. And without Input Monitoring, macOS still returns a valid-looking event tap and simply never delivers events to it — a silent failure — which is why the SPACE toggle stays armed as a backstop even in hold mode. --hold-key none disables the tap entirely.

--voice-activity switches back to hands-free. That path keeps the echo guard (mic ignored while Claude speaks, plus --echo-tail, default 0.4 s), and --barge-in disables even that — headphones only, or Claude will interrupt himself.

Where the delay goes

Measured on this machine, from the end of your sentence to the first audio out:

Stage Push-to-talk Voice activity
Deciding your turn ended ~0 (keypress) ~0.8 s
Apple speech to text 0.05–0.10 s 0.05–0.10 s
Claude, first token (Sonnet) ~1.4 s ~1.4 s
Kokoro, first audio ~0.65 s ~0.65 s

Claude dominates what's left, which is why the default model is Sonnet rather than the largest available — for conversation, first-token latency beats raw capability. --claude-model trades back the other way.

Kokoro's share is the price of sentence-at-a-time synthesis: it waits for a sentence boundary before speaking so the prosody is right. Short first sentences therefore start talking sooner, which is part of why VOICE_STYLE asks for brevity.

Voices

--tts kokoro (default) is a local neural voice. --tts apple uses the macOS system voices, which are what say and Spoken Content offer.

For Irish, --tts apple --voice Moira works and is the only en-IE voice macOS ships. It's slightly faster than Kokoro here — 0.66 s to first audio versus 0.98 s — but only the "super-compact" variant is installed, which sounds noticeably synthetic. Downloading the Enhanced or Premium version of Moira from System Settings → Accessibility → Spoken Content → System Voice → English (Ireland) makes it much better, and requires no code change.

Siri's voices are not available, Irish or otherwise. Apple doesn't expose them to third-party apps, so neither say nor AVSpeechSynthesizer can see them — --list-voices shows everything that is reachable.

Which speech engine

The default is SpeechTranscriber, macOS 26's new on-device model, reached through the Swift helper in swift/. It needs that helper built and approved (see below); without it the agent falls back to the older dictation model automatically.

To hear the difference on your own voice rather than trusting a benchmark:

.venv/bin/python compare_engines.py          # records you, runs every engine

--stt-engine apple or --analyzer-module dictation switch to the older model, which is worse bare but honours the vocabulary file.

Building the helper

swift/build.sh

On this managed Mac the fresh binary is killed on launch (exit 137) until it has been through binary approval, and every rebuild needs approving again since the hash changes. build.sh says so when it happens, and the agent degrades to the older engine meanwhile rather than failing.

Getting the words right

Two files control how well unusual words are heard, and both are meant to be edited:

  • vocabulary.txt — terms the recognizer should expect. Measured at 23.5% → 16.5% word error rate on technical speech. Add the names, jargon and product names you actually say out loud.
  • corrections.txt — heard => replacement rules for mistakes that recur identically. Took it to 14.1%. Add a rule once you've seen the same wrong word twice.

A handful of terms is also discovered automatically from the project — unusual filenames, class names, git branches and authors — and from what Claude has been saying, since that predicts what you'll say next.

Keep vocabulary.txt short and relevant. Padding the list out is actively harmful, and not in the way you'd expect: a thousand random dictionary words cost only 1.1 points, but a hundred names harvested from this project erased the entire benefit. The problem is ordinary words — biasing towards "bot", "plan" or "hack" drags correct speech onto them. Discovered terms are therefore filtered against the system dictionary and capped at a dozen, and the total is capped at 40. Terms you write yourself are always kept, ordinary or not, on the assumption that you meant it.

Claude also gets the list in his system prompt, so he can resolve a mangled transcript while answering instead of needing a separate correction pass. That recovers most of the benefit of one at no latency cost.

--no-vocabulary turns all of it off; --vocabulary-file points somewhere else.

Personality, and where Claude works

By default Claude works in ~/Workspace (created if missing), not in this repo — the assistant is for everyday use, not for editing itself. --cwd points it elsewhere.

Three files there shape how it behaves, all git-tracked so you can see how they drift and roll back if they drift wrong:

File What it does
AGENTS.md Claude's personality, loaded on every launch
vocabulary.txt words the recognizer should expect
corrections.txt fixes for words it gets wrong the same way each time

They're created from the *.example templates in this repo the first time you point at a new workspace, then they're yours to grow.

Two things worth knowing:

  • HTML comments are stripped before Claude sees the file, so you can leave yourself notes in <!-- ... --> without them acting as instructions.
  • Keep it short. It's prepended to every turn, and a long file makes replies longer and more written-sounding. Describe character and habits; the formatting rules live in VOICE_STYLE in bot.py and are applied after the personality, so they win on "no markdown, keep it brief".

CLAUDE.md in the same directory also works — the CLI picks that one up by itself. AGENTS.md is handled here because the CLI ignores it.

Long-term memory

The agent reads your Metamate personal brain at startup — briefing.md, preferences.md, profile.md — so it already knows what you're working on. Ask "what am I most overdue on" and it answers from your actual pinned notes. Costs about 1.6 s at launch, cached locally so a VPN drop doesn't lose it.

The 31 project names under projects/ also become vocabulary, since "CIP-Unified-Cooldown" and "pSMSL" are exactly the words a recogniser fumbles.

It writes back too, but narrowly. The agent supplies three tools; when to use them is not in the agent — it's in ~/Workspace/.claude/skills/memory/, since the discipline is your setup rather than a property of the code. Edit that file to change the behaviour.

Tool Goes to For
remember_correction corrections.txt (local) a word the recognizer misheard
remember_preference brain preferences.md how you want to be worked with
remember_note brain notes.md a technique or lesson worth keeping

Corrections stay local because they're about this microphone and this recognizer — meaningless on another machine. The other two are true of you regardless, so they belong in the brain.

Entries land under a ## Learned in voice sessions heading so they never get tucked inside a section you wrote, and it stays obvious which lines came from the agent. The skill also tells it never to touch briefing.md, which the daily cron owns.

--no-brain turns all of it off.

Skills and the journal

Drop a skill in ~/Workspace/.claude/skills/<name>/SKILL.md and it's available in conversation — verified, not assumed: a test skill placed there was discovered and invoked with no extra configuration. What makes it work is loading the workspace as a project source (setting_sources=["project"]) plus skills="all"; without the first, a skills folder there is silently ignored.

There's a README in that folder covering the format and how to write skills that sound right when spoken rather than read.

Every turn is logged to ~/Workspace/journal.jsonl — what was heard, what was answered. The journal sits immediately after the transcript repair rather than at the end of the pipeline, because the user aggregator consumes TranscriptionFrames; anything downstream of it never sees what you said. The reply arrives separately, through the LLM's own callback. It's gitignored, being a verbatim record of everything said near the microphone. It exists so the next round of accuracy work can be measured on real conversations rather than synthesised audio.

What Claude is allowed to do

The shell is on by default, because most of what you'd ask about out loud lives behind meta — experiments, memory, tasks, calendar — and without it the agent can only apologise. Write, Edit and NotebookEdit are denied.

Be clear-eyed about what that means: a shell can write files perfectly well, so denying the edit tools is a speed bump against casual edits, not a security boundary. --read-only is the real boundary — it denies the shell too, leaving only reading and searching. --allow-writes denies nothing.

Worth knowing how that's enforced, because the obvious way doesn't work. Passing allowed_tools with permission_mode="bypassPermissions" restricts nothing — measured: Claude ran Bash while it was absent from that list, with no denial recorded. Only naming tools in disallowed_tools actually blocks them. If you change this, verify by asking it to run a shell command rather than trusting the config to mean what it looks like.

Formatting is stripped before anything is spoken, so a stray **bold** never becomes "asterisk asterisk". spoken_text.py wraps pipecat's MarkdownTextFilter — which runs after sentence aggregation, so markdown split across streaming chunks is already reassembled — and adds the cases it misses: strikethrough, bullet dashes, snake_case (read as words), and 3 * 4, which the base filter turns into "3 4" rather than "3 times 4".

Prompting alone wasn't enough for this. Claude is told not to emit markdown and mostly doesn't, but asking is probabilistic and hearing it once is enough.

Claude is also told to write for speech rather than for a screen: short answers, no markdown, no URLs, and a spoken heads-up before long tool calls. That prompt lives in VOICE_STYLE in bot.py and is the first thing to edit if the replies don't sound the way you want.

Files

  • bot.py — assembles the pipeline; all the tuning knobs are here
  • apple_stt.py — Apple dictation model as a Pipecat STT service
  • claude_llm.py — Pipecat processor that puts Claude Code in the LLM slot
  • speech_analyzer_stt.py — drives the Swift helper; swift/SpeechHelper.swift
  • compare_engines.py — record yourself, run every engine, compare
  • brain.py — reads and writes the Metamate personal brain
  • memory_tools.py — the remember_* tools Claude calls
  • journal.py — one JSON line per turn
  • vocabulary.py — term biasing and repair rules; templates in *.example.txt
  • transcript_repair.py — applies the repair rules to every transcription
  • push_to_talk.py — keyboard turn control, the default
  • global_hotkey.py — system-wide hold-key watcher via a Quartz event tap
  • apple_tts.py — macOS system voices as a Pipecat TTS service
  • echo_guard.py — self-hearing guard, used only in --voice-activity mode
  • sounddevice_transport.py — microphone and speaker transport
  • selftest.py — per-stage diagnostics
  • talk — launcher, sets the CA bundle and runs bot.py

Why this doesn't look like a stock Pipecat example

Five things forced changes worth knowing about before you edit anything.

The SDK's bundled CLI can't authenticate here. claude-agent-sdk ships its own copy of Claude Code inside the wheel and prefers it over the one on PATH. That copy is stock Anthropic: it knows nothing about this org's AI Gateway or the apiKeyHelper in /Library/Application Support/ClaudeCode/managed-settings.json, so every turn comes back "Invalid API key · Fix external API key". bot.py passes cli_path=shutil.which("claude") to force the installed CLI. This is easy to miss because it works fine when launched from inside a Claude Code session, which leaks the gateway environment variables to its children.

Apple's recognizer isn't in Pipecat. apple_stt.py adds it as a SegmentedSTTService. The one real trap: results arrive through the CoreFoundation runloop, so waiting on a threading.Event deadlocks — nothing pumps the runloop and the handler never fires. The service pumps it in 10 ms slices and yields to asyncio in between. Also, a Python exception escaping the result handler crosses back into Objective-C and aborts the process, so that handler catches everything.

No PyAudio. Pipecat's LocalAudioTransport needs PyAudio, which has no macOS wheel and must be compiled against a Homebrew portaudio. Homebrew can't write to /opt/homebrew here, and a locally compiled binary won't run until it's been through the approval process. sounddevice_transport.py is a port of that transport onto sounddevice, whose wheel ships a prebuilt portaudio. Same frame contract, so it drops into any Pipecat pipeline.

Python doesn't trust the TLS proxy. pip and any library that downloads models fail certificate verification until pointed at the system CA bundle. .venv/pip.conf handles pip; talk exports SSL_CERT_FILE for everything else. If you add a dependency, install it through .venv/bin/python -m pip so it picks up pip.conf.

Claude keeps its own history. ClaudeCodeLLM isn't a Pipecat LLMService subclass — the Agent SDK already owns conversation state and tool calling, so wrapping it in Pipecat's context machinery would mean two systems tracking the same conversation. It's a plain FrameProcessor that reads the newest user message off each LLMContextFrame and emits text frames. The Pipecat context still exists, but only so the aggregators can detect when a turn has ended.

Known rough edges

  • Launching from inside a Claude Code session doesn't work. The claude CLI tries to apply its own sandbox-exec profile, which macOS refuses when the parent is already sandboxed, and it exits with code 71. Use a normal terminal.
  • The SPACE fallback is a toggle, not hold-to-talk. A terminal only receives key presses, never releases. Only the event-tap path can do true hold-to-talk, which is why it's the default.
  • Only Moira is available in Irish, at the lowest quality tier. Download the Enhanced or Premium variant from System Settings to fix that.
  • Whisper hallucinates on silence when you fall back to it, emitting things like "Thank you." for breath noise. _NOISE_TRANSCRIPTS in claude_llm.py filters the common ones; add to that set if you hit others.