fix: iMac mic quiet - arecord raw + 12x software gain, strip WAV RIFF header, wakeword api 0.4.0, STT base.en tests, voiceID Adolfo 0.97 enrolled
This commit is contained in:
@@ -0,0 +1,87 @@
|
||||
# Jarvis v2 - Final Setup (iMac Intel i3 Ubuntu 26.04)
|
||||
|
||||
## What we built vs v1
|
||||
v1 (Jarvis repo main):
|
||||
- monolithic jarvis.py, macOS only, mlx-whisper (Apple Silicon), openWakeWord, gTTS, say command, Gemini CLI
|
||||
- No speaker ID, no vision, blocking pipeline
|
||||
|
||||
v2 (Jarvis repo -> branch v2):
|
||||
- Modular: audio/capture (PipeWire pw-record backend, no PortAudio needed), VAD (silero + Energy fallback), wakeword (openWakeWord TFLite hey_jarvis 0.4), STT (faster-whisper base INT8 CPU), SpeakerID (ECAPA-TDNN speechbrain + stats fallback), Vision (OpenCV Haar + face_recognition), LLM pluggable (ollama qwen2.5:3b / gemini_cli / hermes_api / echo), TTS (Piper ONNX + espeak + echo)
|
||||
- Pipeline state machine: IDLE -> WAKE -> RECORDING (with pre-roll 0.8s) -> TRANSCRIBING -> IDENTIFYING (voice+face fusion) -> THINKING -> SPEAKING -> FOLLOWUP window (12s no wake needed)
|
||||
- Tested on iMac: mic via pw-record 16k mono OK, VAD energy OK, faster-whisper base int8 loads, cam0 640x480 OK, SpeakerStore empty (ready to enroll), Echo LLM/TTS OK
|
||||
|
||||
## iMac paths
|
||||
~/Projects/jarvis-v2
|
||||
.venv/ (uv, python 3.14, torch CPU, faster-whisper, openwakeword, opencv-headless etc)
|
||||
jarvis.yaml (active config) + jarvis.yaml.example
|
||||
data/soul.md (persona)
|
||||
data/speakers/ (enroll .npy + .json)
|
||||
data/faces/ (enroll .npy + .json)
|
||||
data/transcriptions/YYYY-MM-DD.log
|
||||
models/ (Piper voices, gitignored)
|
||||
systemd/jarvis.service (user service)
|
||||
|
||||
## Commands on iMac
|
||||
export PATH=$HOME/.local/bin:$PATH
|
||||
cd ~/Projects/jarvis-v2
|
||||
source .venv/bin/activate
|
||||
|
||||
# List audio
|
||||
jarvis devices
|
||||
|
||||
# Mic level test
|
||||
jarvis test-mic
|
||||
# VAD test
|
||||
jarvis test-vad
|
||||
|
||||
# Low-level wav test
|
||||
python scripts/test_mic.py
|
||||
python scripts/test_wakeword.py
|
||||
python scripts/test_stt.py
|
||||
|
||||
# Enroll voices (do 3 samples per person for robustness, quiet room)
|
||||
jarvis enroll-speaker --name Adolfo --mic
|
||||
jarvis enroll-speaker --name Grace --mic
|
||||
jarvis enroll-speaker --name Elias --mic
|
||||
# or from file: jarvis enroll-speaker --name Adolfo --file /path/to/wav (16k mono)
|
||||
|
||||
# Enroll faces
|
||||
jarvis enroll-face --name Adolfo
|
||||
jarvis enroll-face --name Grace
|
||||
jarvis enroll-face --name Elias
|
||||
|
||||
# Run assistant (echo brain - placeholder)
|
||||
jarvis run
|
||||
# Later with ollama
|
||||
# install ollama: curl -fsSL https://ollama.com/install.sh | sh
|
||||
# ollama pull qwen2.5:3b
|
||||
# edit jarvis.yaml llm.provider=ollama
|
||||
# jarvis run
|
||||
|
||||
# Systemd always-on
|
||||
./scripts/install_service.sh
|
||||
systemctl --user start jarvis
|
||||
journalctl --user -u jarvis -f
|
||||
systemctl --user stop jarvis
|
||||
systemctl --user disable jarvis
|
||||
|
||||
## Configurable tiny-LLM gate (future)
|
||||
In jarvis.yaml wakeword.allow_followup_without_wake + conversation.followup_timeout
|
||||
Future: add local 0.5B Qwen to decide if utterance is addressed to Jarvis even without wakeword (use voice activity + speaker gaze)
|
||||
|
||||
## Piper TTS voices
|
||||
.venv/bin/python -m piper.download_voices en_US-lessac-medium
|
||||
# then edit jarvis.yaml tts.engine=piper tts.voice=en_US-lessac-medium
|
||||
|
||||
## Next steps priority
|
||||
1. Enroll speakers (you + kids)
|
||||
2. Enroll faces (webcam)
|
||||
3. Install Ollama + qwen2.5:3b for local LLM
|
||||
4. Test with wakeword: say "Hey Jarvis" + command
|
||||
5. Install Piper voice for natural TTS
|
||||
6. (Optional) Build PortAudio in ~/.local if you want sounddevice backend: ./scripts/build_portaudio.sh; export LD_LIBRARY_PATH=$HOME/.local/lib:$LD_LIBRARY_PATH
|
||||
7. Create repo Jarvis-v2 in Gitea web UI (private) then push main: git push origin main
|
||||
|
||||
## Gitea
|
||||
Current: Jarvis repo branch v2 contains full v2 code (pushed)
|
||||
Missing: Jarvis-v2 new repo creation blocked by server setting "Push to create is not enabled for users." Need to create via Gitea UI, then we can push main.
|
||||
Reference in New Issue
Block a user