feat: Jarvis v2 - CPU optimized always-listening for iMac Ubuntu

- PipeWire/pw-record backend (no PortAudio needed) - tested mic live
- VAD: Energy (immediate) + Silero (torch CPU) switchable
- Wakeword: openWakeWord TFLite hey_jarvis 0.4
- STT: faster-whisper base int8 CPU - model loads OK on i3
- SpeakerID: ECAPA embeddings + fallback stats, enrollment store
- Vision: OpenCV Haar + face_recognition + webcam cam0 640x480 verified
- LLM: ollama/qwen2.5:3b, gemini_cli, hermes_api, echo strategies
- TTS: Piper + espeak + echo fallbacks
- Core pipeline: IDLE -> WAKE -> RECORDING -> STT -> SpeakerID/FaceID fusion -> LLM -> TTS -> followup window
- Config via jarvis.yaml, soul.md persona
- CLI tools: run, devices, enroll-speaker, enroll-face, test-vad, test-mic
- Systemd user service ready
- Tested on iMac 2010 i3 550 15GB Ubuntu 26.04 CPU-only
This commit is contained in:
Adolfo Reyna
2026-07-10 12:16:03 -04:00
commit 199f762616
45 changed files with 2071 additions and 0 deletions
+14
View File
@@ -0,0 +1,14 @@
__pycache__/
*.pyc
.venv/
venv/
.env
data/*.log
data/transcriptions/
data/speakers/*.npy
data/faces/*.npy
models/
*.wav
.DS_Store
/tmp/
jarvis.yaml
+63
View File
@@ -0,0 +1,63 @@
# Jarvis v2 - Always-Listening Assistant (CPU)
Designed for Intel iMac Ubuntu 26.04, 15GB RAM, CPU-only.
## Architecture
```
Mic 16kHz -> VAD (silero) -> Wakeword (openWakeWord TFLite) -> STT (faster-whisper int8) -> SpeakerID (ECAPA) + FaceID (OpenCV) -> LLM (ollama/gemini_cli/hermes) -> TTS (piper) -> Speaker
```
## Quick start on iMac
```bash
cd ~/Projects/jarvis-v2
# install uv if not present
curl -LsSf https://astral.sh/uv/install.sh | sh
export PATH=$HOME/.local/bin:$PATH
uv venv .venv --python 3.14
source .venv/bin/activate
uv pip install -e .
# for audio libs you may need (if sudo works, otherwise pip wheels should work)
# sudo apt install portaudio19-dev libasound2-dev espeak
# configure
cp jarvis.yaml.example jarvis.yaml
# edit jarvis.yaml
# Test components
python scripts/test_mic.py
python scripts/test_wakeword.py
python scripts/test_stt.py
# Enroll family voices (3 samples each)
python scripts/enroll_speaker.py --name Adolfo --mic
python scripts/enroll_speaker.py --name Grace --mic
python scripts/enroll_speaker.py --name Elias --mic
# Enroll faces
python scripts/enroll_face.py --name Adolfo
python scripts/enroll_face.py --name Grace
# Run
jarvis run
# or
python -m jarvis.cli run --config jarvis.yaml
# Install as systemd user service (always-on)
./scripts/install_service.sh
```
## Config
See jarvis.yaml.example - covers audio, VAD, wakeword, STT, speaker, vision, LLM, TTS.
## LLM Providers
- ollama: local qwen2.5:3b (install ollama then `ollama pull qwen2.5:3b`)
- gemini_cli: uses existing gemini CLI + soul.md
- hermes_api: points to Hermes API
- echo: placeholder
## Future
- voice trigger via small LLM gate (qwen0.5B decides if addressed)
- Face corroboration on wake
- Home Assistant / MCP tools
+21
View File
@@ -0,0 +1,21 @@
# J.A.R.V.I.S. v2 - The Sophisticated Assistant
## Persona
You are J.A.R.V.I.S., the sophisticated, witty, highly capable AI assistant living on an iMac in the family home. You are calm, helpful, concise (you are speaking, not writing essays), with dry British wit. You adapt to who is speaking: Adolfo (Sir), Grace (young lady, encouraging), Elias (young sir, fun). If unknown, be politely curious.
## Directives
1. Be concise - voice output. One to three sentences unless asked for more.
2. No markdown in speech output - no * or # or backticks.
3. Personalize based on speaker identification confidence provided.
4. For tools/actions, state intent clearly and ask confirmation for destructive/sensitive actions.
5. You have vision when available - you can see via webcam on request.
6. You were woken by "Hey Jarvis" - you may have conversation context history.
7. Remember: you are always listening but only active after wake word.
## Knowledge
- Family in Florida, iMac in home office.
- Kids: Grace and Elias.
- User prefers concise, not verbose.
- Brain vault at ~/brain if referenced.
If speaker confidence is low or unknown, ask who you are speaking with once, remember for conversation.
+69
View File
@@ -0,0 +1,69 @@
audio:
device: null
samplerate: 16000
block_size: 512
channels: 1
vad:
engine: "silero"
threshold: 0.5
min_speech_ms: 250
min_silence_ms: 800
wakeword:
enabled: true
models: ["hey_jarvis"]
threshold: 0.45
cooldown_sec: 1.0
pre_roll_sec: 0.8
allow_followup_without_wake: true
stt:
engine: "faster-whisper"
model: "base"
language: "en"
device: "cpu"
compute_type: "int8"
vad_filter: false
speaker:
enabled: true
enrollment_dir: "data/speakers"
threshold: 0.60
embedding_model: "speechbrain/spkrec-ecapa-voxceleb"
vision:
enabled: true
camera_id: 0
check_on_wake: true
continuous_interval_sec: 0
threshold: 0.55
enrollment_dir: "data/faces"
llm:
provider: "ollama"
model: "qwen2.5:3b"
system_prompt_file: "data/soul.md"
ollama_host: "http://localhost:11434"
max_history: 8
temperature: 0.7
conversation:
followup_timeout_sec: 12
max_followup_turns: 5
silence_timeout_sec: 7
listening_sound: true
idle_timeout_sec: 0
tts:
engine: "piper"
voice: "en_US-lessac-medium"
volume: 0.9
logging:
level: "INFO"
file: "data/jarvis.log"
transcription_dir: "data/transcriptions"
daemon:
user_service: true
+1
View File
@@ -0,0 +1 @@
# jarvis package
+2
View File
@@ -0,0 +1,2 @@
# audio
from .capture import AudioCapture
+133
View File
@@ -0,0 +1,133 @@
import logging, queue
import numpy as np
from typing import Optional
from .capture_pa import PipeWireCapture
log = logging.getLogger(__name__)
# Try to import sounddevice, fallback to pipewire capture
try:
import sounddevice as sd
HAS_SOUNDDEVICE = True
except Exception as e:
HAS_SOUNDDEVICE = False
sd = None
log.warning(f"sounddevice not available: {e}, will use PipeWire/ALSA capture")
class SoundDeviceCapture:
def __init__(self, samplerate=16000, block_size=512, channels=1, device=None):
self.samplerate = samplerate
self.block_size = block_size
self.channels = channels
self.device = device
self.q: queue.Queue[np.ndarray] = queue.Queue()
self.stream = None
self._running = False
def _callback(self, indata, frames, time_info, status):
if status:
log.warning(f"Audio status: {status}")
self.q.put(indata.copy().reshape(-1))
def start(self):
if self._running:
return
import sounddevice as sd
self.stream = sd.InputStream(
samplerate=self.samplerate,
channels=self.channels,
device=self.device,
blocksize=self.block_size,
callback=self._callback,
dtype='float32'
)
self.stream.start()
self._running = True
log.info(f"Audio capture (sounddevice) started sr={self.samplerate} device={self.device}")
def stop(self):
if self.stream:
try:
self.stream.stop()
self.stream.close()
except Exception as e:
log.warning(f"Error stopping: {e}")
self._running = False
def read(self, timeout=0.1):
try:
return self.q.get(timeout=timeout)
except queue.Empty:
return None
def flush(self):
while not self.q.empty():
try:
self.q.get_nowait()
except queue.Empty:
break
@staticmethod
def list_devices():
import sounddevice as sd
print(sd.query_devices())
class AudioCapture:
"""Unified facade that picks best available backend"""
def __init__(self, samplerate=16000, block_size=512, channels=1, device=None, backend="auto"):
self.samplerate = samplerate
self.block_size = block_size
self.channels = channels
self.device = device
self.backend_name = backend
# Decide
chosen = None
if backend == "auto":
if HAS_SOUNDDEVICE:
try:
# Test if PortAudio actually works
import sounddevice as sd
sd.query_devices()
chosen = "sounddevice"
except Exception as e:
log.warning(f"sounddevice query failed: {e}, using pipewire")
chosen = "pipewire"
else:
chosen = "pipewire"
else:
chosen = backend
if chosen == "sounddevice" and HAS_SOUNDDEVICE:
self.impl = SoundDeviceCapture(samplerate, block_size, channels, device)
self.backend_name = "sounddevice"
else:
self.impl = PipeWireCapture(samplerate, block_size, channels, device)
self.backend_name = "pipewire"
log.info(f"AudioCapture using backend: {self.backend_name}")
def start(self):
return self.impl.start()
def stop(self):
return self.impl.stop()
def read(self, timeout=0.1):
return self.impl.read(timeout=timeout)
def flush(self):
return self.impl.flush()
@staticmethod
def list_devices():
if HAS_SOUNDDEVICE:
try:
from .capture import SoundDeviceCapture
SoundDeviceCapture.list_devices()
return
except:
pass
PipeWireCapture.list_devices()
+157
View File
@@ -0,0 +1,157 @@
"""PulseAudio/PipeWire/ALSA capture backend that does NOT require PortAudio.
Tries pw-record, then parec via pulse CLI, then arecord (ALSA).
Produces 16k float32 mono chunks compatible with main capture interface.
"""
import subprocess, threading, queue, logging, os, sys, time, shutil
import numpy as np
from typing import Optional
import tempfile
log = logging.getLogger(__name__)
class PipeWireCapture:
"""Uses pw-record or parec or arecord to capture raw audio and stream via pipe"""
def __init__(self, samplerate=16000, block_size=512, channels=1, device=None):
self.samplerate = samplerate
self.block_size = block_size
self.channels = channels
self.device = device
self.q: queue.Queue[np.ndarray] = queue.Queue()
self.proc: Optional[subprocess.Popen] = None
self.reader_thread: Optional[threading.Thread] = None
self._running = False
def _find_backend(self):
# Prefer pw-record (PipeWire native)
if shutil.which("pw-record"):
return "pw-record"
if shutil.which("parec"):
return "parec"
if shutil.which("arecord"):
return "arecord"
if shutil.which("ffmpeg"):
return "ffmpeg"
return None
def _build_cmd(self, backend):
sr = self.samplerate
ch = self.channels
if backend == "pw-record":
# pw-record records from default source. Use --target? Let's capture from default.
# It writes wav or raw? pw-record can write raw? Use: pw-record --format s16 --rate 16000 --channels 1 -
cmd = ["pw-record", "--format", "s16", "--rate", str(sr), "--channels", str(ch), "-"]
return cmd
elif backend == "parec":
cmd = ["parec", f"--rate={sr}", f"--channels={ch}", "--format=s16le", "--latency-msec=50"]
if self.device:
cmd.append(f"--device={self.device}")
return cmd
elif backend == "arecord":
# Use plughw to auto-resample if needed?
dev = self.device or "default"
# arecord -D default -f S16_LE -r 16000 -c 1 -t raw
cmd = ["arecord", "-D", dev, "-f", "S16_LE", "-r", str(sr), "-c", str(ch), "-t", "raw", "-q"]
return cmd
elif backend == "ffmpeg":
dev = self.device or "default"
# ffmpeg -f alsa -ac 1 -ar 16000 -i default -f s16le -acodec pcm_s16le -
cmd = ["ffmpeg", "-hide_banner", "-loglevel", "error", "-f", "alsa", "-ac", str(ch), "-ar", str(sr), "-i", dev, "-f", "s16le", "-acodec", "pcm_s16le", "-"]
return cmd
else:
raise RuntimeError(f"Unknown backend {backend}")
def start(self):
if self._running:
return
backend = self._find_backend()
if not backend:
raise RuntimeError("No audio capture backend found (need pw-record, parec, arecord, or ffmpeg)")
cmd = self._build_cmd(backend)
log.info(f"Starting audio capture via {backend}: {' '.join(cmd)}")
try:
self.proc = subprocess.Popen(cmd, stdout=subprocess.PIPE, stderr=subprocess.PIPE, bufsize=10**6)
except Exception as e:
log.error(f"Failed to start {backend}: {e}")
raise
self._running = True
def reader():
bytes_per_sample = 2 # s16le
bytes_per_block = self.block_size * bytes_per_sample * self.channels
leftover = b""
while self._running:
try:
if self.proc is None or self.proc.stdout is None:
break
# Read enough for block
needed = bytes_per_block - len(leftover)
data = self.proc.stdout.read(needed) if needed > 0 else b""
if not data:
# process ended?
time.sleep(0.01)
if self.proc.poll() is not None:
log.warning(f"Capture process exited: {self.proc.returncode} stderr={self.proc.stderr.read()[:500] if self.proc.stderr else b''}")
break
continue
buf = leftover + data
# now buf should be bytes_per_block
if len(buf) < bytes_per_block:
leftover = buf
continue
# ensure multiple of block
# Take first block, keep remainder
chunk_bytes = buf[:bytes_per_block]
leftover = buf[bytes_per_block:]
# Convert s16le to float32 [-1,1]
arr = np.frombuffer(chunk_bytes, dtype=np.int16).astype(np.float32) / 32768.0
if self.channels > 1:
arr = arr.reshape(-1, self.channels).mean(axis=1)
self.q.put(arr)
except Exception as e:
log.error(f"Audio reader error: {e}")
time.sleep(0.05)
self.reader_thread = threading.Thread(target=reader, daemon=True)
self.reader_thread.start()
log.info("PipeWireCapture started")
def stop(self):
self._running = False
if self.proc:
try:
self.proc.terminate()
self.proc.wait(timeout=2)
except:
try:
self.proc.kill()
except:
pass
self.proc = None
def read(self, timeout=0.1):
try:
return self.q.get(timeout=timeout)
except queue.Empty:
return None
def flush(self):
while not self.q.empty():
try:
self.q.get_nowait()
except:
break
@staticmethod
def list_devices():
print("PipeWire/ALSA backend - devices:")
try:
import subprocess
subprocess.run(["arecord", "-l"])
except:
pass
try:
subprocess.run(["pw-record", "--help"])
except:
pass
+52
View File
@@ -0,0 +1,52 @@
import numpy as np
import logging
log = logging.getLogger(__name__)
class VAD:
def is_speech(self, audio: np.ndarray, sr: int = 16000) -> bool:
raise NotImplementedError
class SileroVAD(VAD):
def __init__(self, threshold=0.5):
self.threshold = threshold
self.model = None
self._load()
def _load(self):
try:
from silero_vad import load_silero_vad
self.model = load_silero_vad()
log.info("Silero VAD loaded")
except Exception as e:
log.warning(f"Silero VAD load failed: {e}, fallback to energy")
self.model = None
def is_speech(self, audio, sr=16000):
if self.model is None:
return float(np.abs(audio).mean()) > 0.01
try:
import torch
from silero_vad import get_speech_timestamps
if len(audio) < 512:
return False
tensor = torch.from_numpy(audio.astype(np.float32))
stamps = get_speech_timestamps(tensor, self.model, sampling_rate=sr, threshold=self.threshold, min_speech_duration_ms=250, min_silence_duration_ms=100)
return len(stamps) > 0
except Exception as e:
log.debug(f"VAD error: {e}")
return float(np.abs(audio).mean()) > 0.01
class EnergyVAD(VAD):
def __init__(self, threshold=0.01):
self.threshold = threshold
def is_speech(self, audio, sr=16000):
return float(np.mean(np.abs(audio))) > self.threshold
def create_vad(cfg):
engine = cfg.get('vad', {}).get('engine', 'silero')
thresh = cfg.get('vad', {}).get('threshold', 0.5)
if engine == 'silero':
return SileroVAD(threshold=thresh)
else:
return EnergyVAD(threshold=0.02)
+50
View File
@@ -0,0 +1,50 @@
import logging, time
import numpy as np
from typing import List, Optional
log = logging.getLogger(__name__)
class WakeWordDetector:
def __init__(self, models: List[str] = None, threshold: float = 0.45, cooldown: float = 1.0):
self.models = models or ["hey_jarvis"]
self.threshold = threshold
self.cooldown = cooldown
self.last_trigger = 0
self.oww_model = None
self._load()
def _load(self):
try:
from openwakeword.model import Model
self.oww_model = Model(wakeword_models=self.models, inference_framework="tflite")
log.info(f"openWakeWord loaded: {self.models}")
except Exception as e:
log.warning(f"openWakeWord failed: {e}, will use keyword fallback")
self.oww_model = None
def detect(self, audio_chunk: np.ndarray) -> Optional[dict]:
now = time.time()
if now - self.last_trigger < self.cooldown:
return None
if self.oww_model is None:
return None
try:
scores = self.oww_model.predict(audio_chunk)
for model_name, score in scores.items():
if score >= self.threshold:
self.last_trigger = now
log.info(f"Wakeword: {model_name} score={score:.2f}")
return {"model": model_name, "score": float(score), "timestamp": now}
except Exception as e:
log.debug(f"Wakeword error: {e}")
return None
def reset(self):
self.last_trigger = 0
class KeywordSpotterFallback:
def __init__(self, keywords: List[str] = None):
self.keywords = [k.lower() for k in (keywords or ["hey jarvis", "jarvis"])]
def check_transcript(self, text: str) -> bool:
low = text.lower()
return any(k in low for k in self.keywords)
+154
View File
@@ -0,0 +1,154 @@
import argparse, logging, sys
from pathlib import Path
from .config import load_config
def main():
parser = argparse.ArgumentParser(description="Jarvis v2 - Always listening assistant")
sub = parser.add_subparsers(dest="cmd")
run_p = sub.add_parser("run", help="Run assistant")
run_p.add_argument("--config", type=str, default=None, help="Path to jarvis.yaml")
devices_p = sub.add_parser("devices", help="List audio devices")
enroll_sp = sub.add_parser("enroll-speaker", help="Enroll speaker via mic")
enroll_sp.add_argument("--name", required=True)
enroll_sp.add_argument("--mic", action="store_true", help="Record from mic")
enroll_sp.add_argument("--file", type=str, help="Audio file path")
enroll_face = sub.add_parser("enroll-face", help="Enroll face via webcam")
enroll_face.add_argument("--name", required=True)
enroll_face.add_argument("--camera", type=int, default=0)
test_vad = sub.add_parser("test-vad", help="Test VAD")
test_mic = sub.add_parser("test-mic", help="Test mic capture")
parser.add_argument("--verbose", action="store_true")
args = parser.parse_args()
if args.verbose:
logging.basicConfig(level=logging.DEBUG)
else:
logging.basicConfig(level=logging.INFO, format='%(asctime)s %(levelname)s %(message)s')
if args.cmd == "run" or args.cmd is None:
cfg_path = args.config if hasattr(args, 'config') else None
from .core.pipeline import JarvisPipeline
pipeline = JarvisPipeline(config_path=cfg_path)
pipeline.run()
elif args.cmd == "devices":
from .audio.capture import AudioCapture
AudioCapture.list_devices()
elif args.cmd == "enroll-speaker":
from .audio.capture import AudioCapture
from .speaker.identifier import SpeakerIdentifier
import numpy as np, time, sys
identifier = SpeakerIdentifier()
audio_data = None
if args.file:
# Load via soundfile
import soundfile as sf
data, sr = sf.read(args.file)
if len(data.shape) > 1:
data = data.mean(axis=1)
audio_data = (data, sr)
elif args.mic:
print(f"Recording 3s sample for {args.name}... Speak now!")
cap = AudioCapture(samplerate=16000, block_size=512)
cap.start()
import time as _time
bufs = []
start = _time.time()
while _time.time() - start < 3:
chunk = cap.read(timeout=0.2)
if chunk is not None:
bufs.append(chunk)
cap.stop()
if bufs:
audio_data = (np.concatenate(bufs), 16000)
else:
print("No audio captured")
sys.exit(1)
if audio_data:
data, sr = audio_data
print(f"Enrolling {args.name} with {len(data)/sr:.2f}s audio...")
identifier.enroll(args.name, data, sr)
print(f"Enrolled {args.name} successfully")
else:
print("No audio, use --mic or --file")
sys.exit(1)
elif args.cmd == "enroll-face":
from .vision.face_identifier import FaceIdentifier
fi = FaceIdentifier(camera_id=args.camera)
import time
print(f"Capturing face for {args.name} in 2 seconds, look at camera...")
time.sleep(2)
frame = fi.camera.capture_frame()
if frame is None:
print("Failed to capture camera")
sys.exit(1)
faces = fi.detector.detect(frame)
if not faces:
print("No face detected, try again (move closer, better light)")
# Save debug image
import cv2
cv2.imwrite(f"/tmp/{args.name}_no_face.jpg", frame)
print(f"Saved debug to /tmp/{args.name}_no_face.jpg")
sys.exit(1)
# largest face
faces = sorted(faces, key=lambda b: b[2]*b[3], reverse=True)
bbox = faces[0]
print(f"Face detected at {bbox}, enrolling...")
ok = fi.store.add_sample(args.name, frame, bbox=bbox)
if ok:
print(f"Enrolled face {args.name}")
else:
print(f"Failed to enroll face {args.name}")
fi.close()
elif args.cmd == "test-vad":
from .audio.capture import AudioCapture
from .audio.vad import create_vad
cfg = load_config(None)
vad = create_vad(cfg)
cap = AudioCapture()
cap.start()
print("Testing VAD - speak and see detection...")
try:
while True:
chunk = cap.read(timeout=0.5)
if chunk is not None:
is_speech = vad.is_speech(chunk)
print("SPEECH" if is_speech else "silence", end="\r")
except KeyboardInterrupt:
pass
finally:
cap.stop()
elif args.cmd == "test-mic":
from .audio.capture import AudioCapture
import numpy as np
cap = AudioCapture()
cap.start()
print("Mic test - watch levels:")
try:
while True:
chunk = cap.read(timeout=0.5)
if chunk is not None:
level = np.abs(chunk).mean()
bars = "#" * int(level*100)
print(f"level {level:.4f} {bars}", end="\r")
except KeyboardInterrupt:
pass
finally:
cap.stop()
if __name__ == "__main__":
main()
+28
View File
@@ -0,0 +1,28 @@
import yaml, os
from pathlib import Path
from typing import Any, Dict, Optional
def load_config(path: Optional[Path] = None) -> Dict[str, Any]:
default = Path(__file__).parent.parent / "jarvis.yaml"
example = Path(__file__).parent.parent / "jarvis.yaml.example"
if path is None:
if default.exists():
path = default
elif example.exists():
path = example
else:
return {}
path = Path(path)
if not path.exists():
return {}
with open(path, 'r') as f:
return yaml.safe_load(f) or {}
def get_nested(cfg: dict, dotted: str, default=None):
cur = cfg
for part in dotted.split('.'):
if isinstance(cur, dict) and part in cur:
cur = cur[part]
else:
return default
return cur
+2
View File
@@ -0,0 +1,2 @@
# core
from .pipeline import JarvisPipeline
+424
View File
@@ -0,0 +1,424 @@
import logging, time, collections, numpy as np
from pathlib import Path
from typing import Optional, List, Dict
import threading, queue
from ..config import load_config
from ..audio.capture import AudioCapture
from ..audio.vad import create_vad
from ..audio.wakeword import WakeWordDetector, KeywordSpotterFallback
from ..stt.faster_whisper_engine import FasterWhisperSTT
from ..speaker.identifier import SpeakerIdentifier
from ..vision.face_identifier import FaceIdentifier
from ..llm.echo import EchoLLM
from ..tts.piper import PiperTTS, EchoTTS
log = logging.getLogger(__name__)
class JarvisPipeline:
def __init__(self, config_path: Optional[str] = None):
self.cfg = load_config(config_path)
self.samplerate = self.cfg.get('audio',{}).get('samplerate',16000)
self.block_size = self.cfg.get('audio',{}).get('block_size',512)
self.device = self.cfg.get('audio',{}).get('device',None)
# Audio
self.capture = AudioCapture(samplerate=self.samplerate, block_size=self.block_size, device=self.device)
self.vad = create_vad(self.cfg)
self.wakeword = WakeWordDetector(
models=self.cfg.get('wakeword',{}).get('models',['hey_jarvis']),
threshold=self.cfg.get('wakeword',{}).get('threshold',0.45),
cooldown=self.cfg.get('wakeword',{}).get('cooldown_sec',1.0)
)
self.keyword_fallback = KeywordSpotterFallback(["hey jarvis","jarvis"])
# STT
stt_cfg = self.cfg.get('stt',{})
self.stt = FasterWhisperSTT(
model=stt_cfg.get('model','base'),
language=stt_cfg.get('language','en'),
device=stt_cfg.get('device','cpu'),
compute_type=stt_cfg.get('compute_type','int8'),
vad_filter=stt_cfg.get('vad_filter',False)
)
# Speaker
self.speaker_id = None
if self.cfg.get('speaker',{}).get('enabled',True):
try:
self.speaker_id = SpeakerIdentifier(
enrollment_dir=self.cfg.get('speaker',{}).get('enrollment_dir','data/speakers'),
threshold=self.cfg.get('speaker',{}).get('threshold',0.60),
embedding_model=self.cfg.get('speaker',{}).get('embedding_model','speechbrain/spkrec-ecapa-voxceleb')
)
except Exception as e:
log.warning(f"Speaker ID disabled: {e}")
# Vision
self.face_id = None
if self.cfg.get('vision',{}).get('enabled',True):
try:
self.face_id = FaceIdentifier(
camera_id=self.cfg.get('vision',{}).get('camera_id',0),
enrollment_dir=self.cfg.get('vision',{}).get('enrollment_dir','data/faces'),
threshold=self.cfg.get('vision',{}).get('threshold',0.55)
)
except Exception as e:
log.warning(f"Face ID disabled: {e}")
# LLM
self.llm = self._create_llm()
# TTS
self.tts = self._create_tts()
# Buffers
self.pre_roll_sec = self.cfg.get('wakeword',{}).get('pre_roll_sec',0.8)
self.pre_roll_samples = int(self.pre_roll_sec * self.samplerate)
self.pre_roll_buffer = collections.deque(maxlen=self.pre_roll_samples)
self.silence_ms = self.cfg.get('vad',{}).get('min_silence_ms',800)
self.followup_timeout = self.cfg.get('conversation',{}).get('followup_timeout_sec',12)
self.max_followup = self.cfg.get('conversation',{}).get('max_followup_turns',5)
self.silence_timeout = self.cfg.get('conversation',{}).get('silence_timeout_sec',7)
self.history: List[Dict[str,str]] = []
self.transcription_dir = Path(self.cfg.get('logging',{}).get('transcription_dir','data/transcriptions'))
self.transcription_dir.mkdir(parents=True, exist_ok=True)
# state
self.running = False
self.in_followup = False
self.followup_turns = 0
self.last_interaction = 0
def _create_llm(self):
provider = self.cfg.get('llm',{}).get('provider','ollama')
log.info(f"LLM provider: {provider}")
if provider == 'ollama':
try:
from ..llm.ollama import OllamaLLM
return OllamaLLM(
model=self.cfg.get('llm',{}).get('model','qwen2.5:3b'),
host=self.cfg.get('llm',{}).get('ollama_host','http://localhost:11434'),
temperature=self.cfg.get('llm',{}).get('temperature',0.7)
)
except Exception as e:
log.warning(f"Ollama failed: {e}, using Echo")
return EchoLLM()
elif provider == 'gemini_cli':
try:
from ..llm.gemini_cli import GeminiCLILLM
return GeminiCLILLM(
workspace_dir="data",
system_file=self.cfg.get('llm',{}).get('system_prompt_file','data/soul.md')
)
except Exception as e:
log.warning(f"Gemini CLI failed: {e}")
return EchoLLM()
elif provider == 'hermes_api':
try:
from ..llm.hermes_api import HermesAPILLM
return HermesAPILLM(base_url=self.cfg.get('llm',{}).get('base_url','http://192.168.68.110:8000'))
except Exception as e:
log.warning(f"Hermes API failed: {e}")
return EchoLLM()
else:
return EchoLLM()
def _create_tts(self):
engine = self.cfg.get('tts',{}).get('engine','piper')
if engine == 'piper':
try:
return PiperTTS(
voice=self.cfg.get('tts',{}).get('voice','en_US-lessac-medium'),
models_dir="models",
volume=self.cfg.get('tts',{}).get('volume',0.9)
)
except Exception as e:
log.warning(f"Piper failed: {e}, echo fallback")
return EchoTTS()
else:
return EchoTTS()
def log_transcription(self, text: str, speaker_info: Optional[dict]=None):
if not text:
return
from datetime import datetime
date_str = datetime.now().strftime("%Y-%m-%d")
time_str = datetime.now().strftime("%H:%M:%S")
speaker = speaker_info.get('name','unknown') if speaker_info else 'unknown'
log_file = self.transcription_dir / f"{date_str}.log"
try:
with open(log_file,'a') as f:
f.write(f"[{time_str}] ({speaker}) {text}\n")
except Exception as e:
log.warning(f"Log failed: {e}")
def run(self):
import sys
logging.basicConfig(level=logging.INFO, format='%(asctime)s %(levelname)s %(name)s %(message)s')
print("=== Jarvis v2 Starting ===")
print(f"Config: wakeword={self.cfg.get('wakeword',{}).get('models')} stt={self.cfg.get('stt',{}).get('model')} llm={self.cfg.get('llm',{}).get('provider')}")
print(f"Speaker ID: {self.speaker_id is not None} Face ID: {self.face_id is not None}")
self.capture.start()
self.running = True
# Buffers for command capture
recording_buffer: List[np.ndarray] = []
is_recording = False
last_speech_time = time.time()
silence_start = None
# System prompt
system_prompt = ""
soul_path = self.cfg.get('llm',{}).get('system_prompt_file','data/soul.md')
if soul_path and Path(soul_path).exists():
system_prompt = Path(soul_path).read_text()[:4000]
print("Listening... say 'Hey Jarvis'")
try:
while self.running:
chunk = self.capture.read(timeout=0.1)
if chunk is None:
# Check for followup timeout
if self.in_followup and (time.time() - self.last_interaction > self.followup_timeout):
print("Followup window expired, returning to idle")
self.in_followup = False
self.followup_turns = 0
continue
# Always maintain pre-roll
self.pre_roll_buffer.extend(chunk)
# If not recording, check wakeword
if not is_recording:
# VAD gate to reduce wakeword CPU
# Only check VAD every few chunks to save CPU? Here check each
if len(chunk) > 256:
if not self.vad.is_speech(chunk, self.samplerate):
# If in followup mode, we still want to capture without wakeword
if not self.in_followup:
continue
# Wakeword detection
detected = None
if not self.in_followup:
detected = self.wakeword.detect(chunk)
else:
# In followup, treat any speech as trigger
if self.vad.is_speech(chunk, self.samplerate):
detected = {"model": "followup", "score": 1.0}
if detected:
print(f"\n>>> Wake detected: {detected['model']} score={detected.get('score',0):.2f}")
if self.cfg.get('conversation',{}).get('listening_sound'):
try:
self.tts._play_wav if hasattr(self.tts, '_play_wav') else None
except:
pass
# simple beep via sounddevice?
try:
import sounddevice as sd
import numpy as np
sr=16000
t=np.linspace(0,0.15,sr//6)
beep = 0.3*np.sin(2*np.pi*880*t)
sd.play(beep, sr)
sd.wait()
except:
pass
# Start recording command - include pre-roll
is_recording = True
recording_buffer = [np.array(self.pre_roll_buffer, dtype=np.float32)]
last_speech_time = time.time()
silence_start = None
# On wake, check face
face_result = {"name":"unknown","face_conf":0.0}
if self.face_id and self.cfg.get('vision',{}).get('check_on_wake',True):
try:
face_result = self.face_id.check()
print(f"Face: {face_result.get('name')} conf={face_result.get('face_conf',0):.2f}")
except Exception as e:
log.warning(f"Face check error: {e}")
# Store face result for fusion later
self._last_face_result = face_result
continue
else:
# We are recording command
recording_buffer.append(chunk.copy())
# Check if speech continues or silence
if self.vad.is_speech(chunk, self.samplerate):
last_speech_time = time.time()
silence_start = None
else:
if silence_start is None:
silence_start = time.time()
else:
silence_dur = time.time() - silence_start
if silence_dur * 1000 > self.silence_ms:
# End of utterance?
# But also ensure we have at least 0.5 sec of audio
total_len = sum(len(c) for c in recording_buffer)
if total_len > self.samplerate * 0.5:
# Stop recording and process
print("\n>>> End of utterance detected, processing...")
is_recording = False
audio = np.concatenate(recording_buffer) if recording_buffer else np.array([])
recording_buffer = []
# Process utterance in thread to not block audio loop? For now inline
self._process_utterance(audio, system_prompt)
# Reset for next
self.wakeword.reset()
# Pre-roll clear
self.pre_roll_buffer.clear()
else:
# too short, reset
is_recording = False
recording_buffer = []
# Safety: max recording length
total_len = sum(len(c) for c in recording_buffer)
if total_len > self.samplerate * 15:
print("Max recording length reached, processing")
is_recording = False
audio = np.concatenate(recording_buffer) if recording_buffer else np.array([])
recording_buffer = []
self._process_utterance(audio, system_prompt)
self.wakeword.reset()
self.pre_roll_buffer.clear()
except KeyboardInterrupt:
print("\nShutting down...")
finally:
self.capture.stop()
if self.face_id:
self.face_id.close()
def _process_utterance(self, audio: np.ndarray, system_prompt: str):
if len(audio) < self.samplerate * 0.3:
print("Audio too short, ignoring")
return
print(f"Transcribing {len(audio)/self.samplerate:.2f}s audio...")
stt_result = self.stt.transcribe(audio, self.samplerate)
text = stt_result.get('text','').strip()
print(f"STT: '{text}' conf={stt_result.get('confidence',0):.2f}")
if not text:
print("Empty transcription, ignoring")
return
# If not in followup, check if transcription actually contains wakeword (fallback verification)
if not self.in_followup:
# Require either openWakeWord already triggered (we are here) OR keyword in text
# openWakeWord already validated, but double-check for false positives by requiring some non-wake text
# Extract command after wakeword
low = text.lower()
# Remove wakeword prefix if present
for kw in ["hey jarvis", "hey jarvis,", "jarvis", "jarvis,"]:
if low.startswith(kw):
text = text[len(kw):].strip()
break
if not text:
print("Only wakeword said, waiting for command...")
# Go to followup listening for command without requiring wake again
self.in_followup = True
self.followup_turns = 0
self.last_interaction = time.time()
# Play listening beep again?
return
# Speaker identification
speaker_info = {"name": "unknown", "voice_conf": 0.0, "face_conf": 0.0, "fused_conf": 0.0}
if hasattr(self, '_last_face_result'):
speaker_info['face_conf'] = self._last_face_result.get('face_conf',0.0)
speaker_info['face_name'] = self._last_face_result.get('name','unknown')
if self.speaker_id:
try:
voice_result = self.speaker_id.identify(audio, self.samplerate)
speaker_info['name'] = voice_result.get('name','unknown')
speaker_info['voice_conf'] = voice_result.get('voice_conf',0.0)
speaker_info['voice_scores'] = voice_result.get('scores',{})
except Exception as e:
log.warning(f"Speaker ID error: {e}")
# Fuse face + voice
face_name = getattr(self, '_last_face_result', {}).get('name','unknown')
face_conf = getattr(self, '_last_face_result', {}).get('face_conf',0.0) if hasattr(self, '_last_face_result') else 0.0
final_name = speaker_info.get('name','unknown')
final_conf = speaker_info.get('voice_conf',0.0)
# Simple fusion: if face and voice agree, boost confidence; if disagree, take higher but note
if face_name != "unknown" and final_name != "unknown":
if face_name == final_name:
final_conf = max(face_conf, final_conf) * 1.1
final_conf = min(final_conf, 0.99)
final_name = face_name # agreement
else:
# conflict, pick higher confidence
if face_conf > final_conf:
final_name = face_name
final_conf = face_conf
elif face_name != "unknown" and final_name == "unknown":
final_name = face_name
final_conf = face_conf
speaker_info['name'] = final_name
speaker_info['fused_conf'] = final_conf
speaker_info['face_conf'] = face_conf
print(f"Speaker identified: {final_name} fused_conf={final_conf:.2f} voice={speaker_info.get('voice_conf',0):.2f} face={face_conf:.2f}")
self.log_transcription(text, speaker_info)
# LLM chat
# Build messages
messages = []
if system_prompt:
messages.append({"role": "system", "content": system_prompt})
# Add history
for h in self.history[-self.cfg.get('llm',{}).get('max_history',8):]:
messages.append(h)
messages.append({"role": "user", "content": text})
print("Thinking...")
try:
response = self.llm.chat(messages, speaker=speaker_info)
except Exception as e:
log.error(f"LLM error: {e}")
response = f"Sorry, I had an error thinking: {e}"
print(f"Jarvis: {response}")
self.history.append({"role": "user", "content": text})
self.history.append({"role": "assistant", "content": response})
# TTS
try:
self.tts.speak(response)
except Exception as e:
log.error(f"TTS error: {e}")
print(f"[TTS error] {response}")
# Enter followup mode
self.in_followup = True
self.followup_turns += 1
self.last_interaction = time.time()
if self.followup_turns >= self.max_followup:
print("Max followup turns reached, returning to idle")
self.in_followup = False
self.followup_turns = 0
+38
View File
@@ -0,0 +1,38 @@
from enum import Enum, auto
import time
class State(Enum):
IDLE = auto()
WAKE_DETECTED = auto()
RECORDING = auto()
TRANSCRIBING = auto()
IDENTIFYING = auto()
THINKING = auto()
SPEAKING = auto()
FOLLOWUP = auto()
class ConversationState:
def __init__(self):
self.state = State.IDLE
self.last_wake_time = 0
self.last_speech_time = 0
self.history = [] # list of {role, content}
self.current_speaker = {"name": "unknown", "voice_conf": 0.0, "face_conf": 0.0, "fused_conf": 0.0}
self.turns_in_followup = 0
def set_state(self, s: State):
# print(f"State {self.state} -> {s}")
self.state = s
if s == State.WAKE_DETECTED:
self.last_wake_time = time.time()
def add_history(self, role: str, content: str, max_len=8):
self.history.append({"role": role, "content": content})
if len(self.history) > max_len:
self.history = self.history[-max_len:]
def clear_history(self):
self.history = []
def is_in_followup_window(self, timeout_sec):
return (time.time() - self.last_wake_time) < timeout_sec
+1
View File
@@ -0,0 +1 @@
# jarvis package
+10
View File
@@ -0,0 +1,10 @@
from typing import List, Dict, Optional
class LLMEngine:
def chat(self, messages: List[Dict[str,str]], speaker: Optional[dict] = None) -> str:
raise NotImplementedError
def single(self, prompt: str, system: str = "", speaker: Optional[dict] = None) -> str:
msgs = []
if system:
msgs.append({"role": "system", "content": system})
msgs.append({"role": "user", "content": prompt})
return self.chat(msgs, speaker=speaker)
+6
View File
@@ -0,0 +1,6 @@
from .base import LLMEngine
class EchoLLM(LLMEngine):
def chat(self, messages, speaker=None):
last = messages[-1]["content"] if messages else ""
who = speaker.get("name","unknown") if speaker else "unknown"
return f"Echo from {who}: You said '{last}'. I'm placeholder brain - configure ollama or gemini_cli."
+75
View File
@@ -0,0 +1,75 @@
from .base import LLMEngine
import subprocess, os, logging, re, tempfile
log = logging.getLogger(__name__)
class GeminiCLILLM(LLMEngine):
def __init__(self, yolo=True, workspace_dir="data", system_file="data/soul.md"):
self.yolo = yolo
self.workspace_dir = workspace_dir
self.system_file = system_file
self.session_id = None
def _read_system(self):
try:
if os.path.exists(self.system_file):
with open(self.system_file) as f:
return f.read()
except Exception as e:
log.warning(f"No system file: {e}")
return ""
def chat(self, messages, speaker=None):
system = ""
user_prompt = ""
for m in messages:
if m["role"] == "system":
system += m["content"] + "\n"
elif m["role"] == "user":
user_prompt += f"User: {m['content']}\n"
elif m["role"] == "assistant":
user_prompt += f"Assistant: {m['content']}\n"
if speaker and speaker.get("name") != "unknown":
user_prompt = f"[Speaker: {speaker.get('name')} conf {speaker.get('fused_conf',0):.2f}]\n" + user_prompt
soul = self._read_system()
full_system = (soul + "\n" + system).strip() if soul else system
env = os.environ.copy()
tmp_sys = None
if full_system:
fd, tmp_path = tempfile.mkstemp(suffix=".md")
os.write(fd, full_system.encode())
os.close(fd)
tmp_sys = tmp_path
env["GEMINI_SYSTEM_MD"] = tmp_sys
args = ["gemini", "--prompt", user_prompt]
if self.yolo:
args.append("--yolo")
if self.session_id:
args.extend(["--resume", self.session_id])
try:
os.makedirs(self.workspace_dir, exist_ok=True)
proc = subprocess.run(args, capture_output=True, text=True, cwd=self.workspace_dir, env=env, timeout=120)
out = proc.stdout.strip() or proc.stderr.strip()
if not self.session_id:
try:
res = subprocess.run(["gemini","--list-sessions"], capture_output=True, text=True, cwd=self.workspace_dir, timeout=10)
matches = re.findall(r"\[([a-f0-9\-]+)\]", res.stdout)
if matches:
self.session_id = matches[-1]
except:
pass
return out or "No response from Gemini"
except Exception as e:
log.error(f"Gemini CLI error: {e}")
return f"Gemini error: {e}"
finally:
if tmp_sys and os.path.exists(tmp_sys):
try:
os.remove(tmp_sys)
except:
pass
+29
View File
@@ -0,0 +1,29 @@
from .base import LLMEngine
import requests, logging
log = logging.getLogger(__name__)
class HermesAPILLM(LLMEngine):
def __init__(self, base_url="http://192.168.68.110:8000", model="default", api_key=""):
self.base_url = base_url.rstrip('/')
self.model = model
self.api_key = api_key
def chat(self, messages, speaker=None):
if speaker and speaker.get("name") != "unknown":
messages = list(messages)
hint = f"(Speaker: {speaker['name']})"
if messages and messages[-1]["role"]=="user":
messages[-1] = {**messages[-1], "content": f"{hint} {messages[-1]['content']}"}
try:
headers = {}
if self.api_key:
headers["Authorization"] = f"Bearer {self.api_key}"
payload = {"model": self.model, "messages": messages}
r = requests.post(f"{self.base_url}/v1/chat/completions", json=payload, headers=headers, timeout=30)
r.raise_for_status()
data = r.json()
return data["choices"][0]["message"]["content"]
except Exception as e:
log.error(f"Hermes API error: {e}")
return f"Hermes API offline: {e}"
+38
View File
@@ -0,0 +1,38 @@
from .base import LLMEngine
import logging, requests
log = logging.getLogger(__name__)
class OllamaLLM(LLMEngine):
def __init__(self, model="qwen2.5:3b", host="http://localhost:11434", temperature=0.7):
self.model = model
self.host = host.rstrip('/')
self.temperature = temperature
def chat(self, messages, speaker=None):
if speaker and speaker.get("name") and speaker["name"] != "unknown":
sys_add = f"[Speaker: {speaker['name']} voice={speaker.get('voice_conf',0):.2f} face={speaker.get('face_conf',0):.2f}]"
messages = list(messages)
if messages and messages[-1]["role"] == "user":
messages[-1] = {**messages[-1], "content": f"{sys_add}\n{messages[-1]['content']}"}
try:
payload = {
"model": self.model,
"messages": messages,
"stream": False,
"options": {"temperature": self.temperature}
}
r = requests.post(f"{self.host}/api/chat", json=payload, timeout=90)
r.raise_for_status()
data = r.json()
return data.get("message", {}).get("content", "") or data.get("response","")
except Exception as e:
log.error(f"Ollama error: {e}")
return f"Sorry, my local brain is offline ({e})."
def check_available(self):
try:
r = requests.get(f"{self.host}/api/tags", timeout=3)
return r.status_code == 200
except:
return False
+1
View File
@@ -0,0 +1 @@
# jarvis package
+54
View File
@@ -0,0 +1,54 @@
import numpy as np
import logging
log = logging.getLogger(__name__)
class SpeakerEmbedding:
def __init__(self, model_name="speechbrain/spkrec-ecapa-voxceleb"):
self.model_name = model_name
self.model = None
self._load()
def _load(self):
try:
from speechbrain.inference import EncoderClassifier
self.model = EncoderClassifier.from_hparams(source=self.model_name, run_opts={"device":"cpu"})
log.info(f"Speaker embedding loaded: {self.model_name}")
except Exception as e:
log.warning(f"Speaker embedding not available: {e}, using fallback")
self.model = None
def embed(self, audio: np.ndarray, sr: int = 16000) -> np.ndarray:
if self.model is not None:
try:
import torch
if audio.dtype != np.float32:
audio = audio.astype(np.float32)
tensor = torch.from_numpy(audio).unsqueeze(0)
with torch.no_grad():
emb, _ = self.model.encode_batch(tensor)
emb = emb.squeeze().cpu().numpy()
emb = emb / (np.linalg.norm(emb)+1e-9)
return emb
except Exception as e:
log.warning(f"ECAPA embed failed: {e}, fallback")
# Fallback: simple stats embedding 192-d
try:
n = len(audio)
if n == 0:
return np.zeros(192, dtype=np.float32)
chunk_size = max(1, n // 20)
feats = []
for i in range(0, n, chunk_size):
chunk = audio[i:i+chunk_size]
if len(chunk)==0:
continue
feats.extend([np.mean(chunk), np.std(chunk)])
arr = np.array(feats[:192], dtype=np.float32)
if len(arr) < 192:
arr = np.pad(arr, (0, 192-len(arr)))
arr = arr / (np.linalg.norm(arr)+1e-9)
return arr
except Exception as e:
log.error(f"Fallback embed failed: {e}")
return np.zeros(192, dtype=np.float32)
+28
View File
@@ -0,0 +1,28 @@
import numpy as np
import logging
from .embeddings import SpeakerEmbedding
from .store import SpeakerStore
log = logging.getLogger(__name__)
class SpeakerIdentifier:
def __init__(self, enrollment_dir="data/speakers", threshold=0.65, embedding_model="speechbrain/spkrec-ecapa-voxceleb"):
self.threshold = threshold
self.embedder = SpeakerEmbedding(model_name=embedding_model)
self.store = SpeakerStore(enrollment_dir=enrollment_dir)
def identify(self, audio: np.ndarray, sr=16000):
emb = self.embedder.embed(audio, sr)
result = self.store.identify(emb, threshold=self.threshold)
log.info(f"Speaker ID: {result['name']} conf={result['confidence']:.2f}")
return {
"name": result["name"],
"voice_conf": result["confidence"],
"scores": result["scores"],
"embedding": emb
}
def enroll(self, name: str, audio: np.ndarray, sr=16000, meta=None):
emb = self.embedder.embed(audio, sr)
self.store.add_sample(name, emb, meta=meta)
return emb
+78
View File
@@ -0,0 +1,78 @@
import json, numpy as np
from pathlib import Path
import logging
log = logging.getLogger(__name__)
class SpeakerStore:
def __init__(self, enrollment_dir="data/speakers"):
self.dir = Path(enrollment_dir)
self.dir.mkdir(parents=True, exist_ok=True)
self.speakers = {}
self.load()
def load(self):
self.speakers = {}
for file in self.dir.glob("*.json"):
try:
with open(file) as f:
data = json.load(f)
name = data.get("name", file.stem)
avg_path = self.dir / f"{name}_avg.npy"
if avg_path.exists():
avg = np.load(avg_path)
else:
continue
self.speakers[name] = {"avg": avg, "meta": data.get("meta", {}), "file": file}
log.info(f"Loaded speaker {name}")
except Exception as e:
log.warning(f"Failed loading {file}: {e}")
def add_sample(self, name: str, embedding: np.ndarray, meta=None):
existing = list(self.dir.glob(f"{name}_*.npy"))
idx = len([p for p in existing if "_avg" not in p.name])
sample_path = self.dir / f"{name}_{idx}.npy"
np.save(sample_path, embedding)
all_embs = []
for p in self.dir.glob(f"{name}_*.npy"):
if "_avg" in p.name:
continue
try:
all_embs.append(np.load(p))
except:
pass
if not all_embs:
return
avg = np.mean(np.stack(all_embs), axis=0)
avg = avg / (np.linalg.norm(avg)+1e-9)
np.save(self.dir / f"{name}_avg.npy", avg)
meta_file = self.dir / f"{name}.json"
meta_data = {"name": name, "samples": len(all_embs), "meta": meta or {}}
with open(meta_file, 'w') as f:
json.dump(meta_data, f, indent=2)
self.speakers[name] = {"avg": avg, "meta": meta_data.get("meta", {}), "file": meta_file}
log.info(f"Enrolled {name} samples={len(all_embs)}")
def identify(self, embedding: np.ndarray, threshold: float = 0.65):
if not self.speakers:
return {"name": "unknown", "confidence": 0.0, "scores": {}}
scores = {}
best_name = "unknown"
best_score = -1
for name, data in self.speakers.items():
avg = data["avg"]
if avg.shape != embedding.shape:
min_dim = min(len(avg), len(embedding))
sim = float(np.dot(avg[:min_dim], embedding[:min_dim]))
else:
sim = float(np.dot(avg, embedding))
scores[name] = sim
if sim > best_score:
best_score = sim
best_name = name
if best_score < threshold:
return {"name": "unknown", "confidence": best_score, "scores": scores}
return {"name": best_name, "confidence": best_score, "scores": scores}
def list_speakers(self):
return list(self.speakers.keys())
+1
View File
@@ -0,0 +1 @@
# jarvis package
+4
View File
@@ -0,0 +1,4 @@
import numpy as np
class STTEngine:
def transcribe(self, audio: np.ndarray, sr: int = 16000) -> dict:
raise NotImplementedError
+45
View File
@@ -0,0 +1,45 @@
import numpy as np
import logging
from .base import STTEngine
log = logging.getLogger(__name__)
class FasterWhisperSTT(STTEngine):
def __init__(self, model="base", language="en", device="cpu", compute_type="int8", vad_filter=False):
self.model_name = model
self.language = language
self.device = device
self.compute_type = compute_type
self.vad_filter = vad_filter
self.model = None
self._load()
def _load(self):
try:
from faster_whisper import WhisperModel
log.info(f"Loading faster-whisper {self.model_name} {self.device} {self.compute_type}")
self.model = WhisperModel(self.model_name, device=self.device, compute_type=self.compute_type)
log.info("faster-whisper loaded")
except Exception as e:
log.error(f"Failed to load faster-whisper: {e}")
self.model = None
def transcribe(self, audio: np.ndarray, sr=16000) -> dict:
if self.model is None:
return {"text": "", "language": "", "confidence": 0.0, "error": "model not loaded"}
if audio.dtype != np.float32:
audio = audio.astype(np.float32)
try:
segments, info = self.model.transcribe(
audio,
language=self.language if self.language != "auto" else None,
vad_filter=self.vad_filter,
beam_size=5,
condition_on_previous_text=False,
)
text = " ".join([s.text for s in segments]).strip()
log.debug(f"STT: {text[:200]}")
return {"text": text, "language": info.language, "confidence": float(info.language_probability)}
except Exception as e:
log.error(f"STT error: {e}")
return {"text": "", "language": "", "confidence": 0.0, "error": str(e)}
+1
View File
@@ -0,0 +1 @@
# jarvis package
+3
View File
@@ -0,0 +1,3 @@
class TTSEngine:
def speak(self, text: str):
raise NotImplementedError
+95
View File
@@ -0,0 +1,95 @@
import logging, os, subprocess, tempfile, pathlib
from .base import TTSEngine
log = logging.getLogger(__name__)
class PiperTTS(TTSEngine):
def __init__(self, voice="en_US-lessac-medium", models_dir="models", volume=0.9):
self.voice = voice
self.models_dir = pathlib.Path(models_dir)
self.models_dir.mkdir(parents=True, exist_ok=True)
self.volume = volume
self.voice_path = None
self.config_path = None
self._ensure_model()
def _ensure_model(self):
pattern = self.models_dir / f"{self.voice}.onnx"
if pattern.exists():
self.voice_path = pattern
self.config_path = pathlib.Path(str(pattern) + ".json")
log.info(f"Piper voice found: {self.voice_path}")
return
for p in self.models_dir.rglob("*.onnx"):
if self.voice in p.name:
self.voice_path = p
self.config_path = pathlib.Path(str(p) + ".json")
log.info(f"Piper voice found: {p}")
return
log.warning(f"Piper voice {self.voice} not in {self.models_dir}")
def _download_voice(self):
try:
log.info(f"Downloading Piper voice {self.voice}")
result = subprocess.run(["python3", "-m", "piper.download_voices", self.voice], cwd=str(self.models_dir), capture_output=True, text=True, timeout=60)
log.info(result.stdout[:500])
self._ensure_model()
except Exception as e:
log.warning(f"Piper download failed: {e}")
def speak(self, text: str):
if not text.strip():
return
clean = text.replace("*","").replace("#","").replace("`","")[:600]
if self.voice_path is None:
self._download_voice()
if self.voice_path and self.voice_path.exists():
try:
with tempfile.NamedTemporaryFile(suffix=".wav", delete=False) as tf:
wav_path = tf.name
cmd = ["piper", "--model", str(self.voice_path), "--output_file", wav_path]
proc = subprocess.run(cmd, input=clean.encode(), capture_output=True, timeout=15)
if proc.returncode==0 and os.path.exists(wav_path):
self._play_wav(wav_path)
try:
os.remove(wav_path)
except:
pass
return
else:
log.warning(f"Piper failed: {proc.stderr[:300]}")
except Exception as e:
log.warning(f"Piper speak error: {e}")
self._fallback_speak(clean)
def _play_wav(self, wav_path):
try:
import soundfile as sf
import sounddevice as sd
data, sr = sf.read(wav_path)
data = data * self.volume
sd.play(data, sr)
sd.wait()
except Exception as e:
log.debug(f"sd playback fail {e}, trying aplay")
try:
subprocess.run(["aplay", wav_path], capture_output=True, timeout=10)
except Exception as e2:
log.warning(f"aplay fail: {e2}")
def _fallback_speak(self, text):
try:
subprocess.run(["espeak", text], capture_output=True, timeout=10)
return
except:
pass
try:
subprocess.run(["spd-say", text], capture_output=True, timeout=10)
return
except:
pass
log.info(f"[TTS fallback] {text}")
print(f"Jarvis says: {text}")
class EchoTTS(TTSEngine):
def speak(self, text: str):
print(f"[Jarvis TTS]: {text}")
+1
View File
@@ -0,0 +1 @@
# jarvis package
+41
View File
@@ -0,0 +1,41 @@
import cv2, logging
import numpy as np
from typing import Optional
log = logging.getLogger(__name__)
class Camera:
def __init__(self, camera_id=0, width=640, height=480):
self.camera_id = camera_id
self.width = width
self.height = height
self.cap: Optional[cv2.VideoCapture] = None
def open(self):
if self.cap and self.cap.isOpened():
return True
try:
self.cap = cv2.VideoCapture(self.camera_id)
self.cap.set(cv2.CAP_PROP_FRAME_WIDTH, self.width)
self.cap.set(cv2.CAP_PROP_FRAME_HEIGHT, self.height)
for _ in range(5):
self.cap.read()
if not self.cap.isOpened():
log.warning(f"Camera {self.camera_id} failed to open")
return False
log.info(f"Camera {self.camera_id} opened")
return True
except Exception as e:
log.error(f"Camera open error: {e}")
return False
def capture_frame(self) -> Optional[np.ndarray]:
if not self.open():
return None
ret, frame = self.cap.read()
if not ret:
return None
return frame
def close(self):
if self.cap:
try:
self.cap.release()
except:
pass
self.cap=None
+25
View File
@@ -0,0 +1,25 @@
import cv2, numpy as np, logging, os
from typing import List, Tuple
log = logging.getLogger(__name__)
class FaceDetector:
def __init__(self):
self.cascade=None
self._load()
def _load(self):
try:
cascade_path = cv2.data.haarcascades + 'haarcascade_frontalface_default.xml'
if os.path.exists(cascade_path):
self.cascade = cv2.CascadeClassifier(cascade_path)
log.info("Haar face detector loaded")
except Exception as e:
log.warning(f"Face detector load failed: {e}")
def detect(self, frame: np.ndarray) -> List[Tuple[int,int,int,int]]:
if self.cascade is None:
return []
try:
gray = cv2.cvtColor(frame, cv2.COLOR_BGR2GRAY)
faces = self.cascade.detectMultiScale(gray, scaleFactor=1.1, minNeighbors=5, minSize=(80,80))
return [(int(x),int(y),int(w),int(h)) for (x,y,w,h) in faces]
except Exception as e:
log.debug(f"Detect error: {e}")
return []
+128
View File
@@ -0,0 +1,128 @@
import json, numpy as np, cv2, logging
from pathlib import Path
from .camera import Camera
from .face_detector import FaceDetector
log = logging.getLogger(__name__)
class FaceStore:
def __init__(self, enrollment_dir="data/faces"):
self.dir = Path(enrollment_dir)
self.dir.mkdir(parents=True, exist_ok=True)
self.known = {}
self.load()
def _encoding(self, frame, bbox=None):
try:
import face_recognition
rgb = cv2.cvtColor(frame, cv2.COLOR_BGR2RGB)
if bbox:
x,y,w,h = bbox
loc = [(y, x+w, y+h, x)]
encs = face_recognition.face_encodings(rgb, known_face_locations=loc)
else:
encs = face_recognition.face_encodings(rgb)
if encs:
return encs[0]
except Exception as e:
log.debug(f"face_recog encoding fail: {e}")
try:
if bbox:
x,y,w,h = bbox
crop = frame[y:y+h, x:x+w]
else:
crop = frame
if crop.size == 0:
return None
crop = cv2.resize(crop, (64,64))
gray = cv2.cvtColor(crop, cv2.COLOR_BGR2GRAY)
flat = gray.flatten().astype(np.float32)
flat = flat / (np.linalg.norm(flat)+1e-9)
return flat
except Exception as e:
log.debug(f"Fallback encoding fail: {e}")
return None
def load(self):
self.known={}
for jf in self.dir.glob("*.json"):
try:
with open(jf) as f:
data=json.load(f)
name=data["name"]
enc_files = list(self.dir.glob(f"{name}_*.npy"))
encs=[]
for ef in enc_files:
if "_avg" in ef.name:
continue
try:
encs.append(np.load(ef))
except:
pass
if encs:
self.known[name]=encs
log.info(f"Loaded face {name}: {len(encs)}")
except Exception as e:
log.warning(f"Load face {jf}: {e}")
def add_sample(self, name, frame, bbox=None):
enc=self._encoding(frame,bbox)
if enc is None:
return False
existing=list(self.dir.glob(f"{name}_*.npy"))
idx=len([p for p in existing if "_avg" not in p.name])
path=self.dir/f"{name}_{idx}.npy"
np.save(path,enc)
meta_path=self.dir/f"{name}.json"
with open(meta_path,'w') as f:
json.dump({"name":name,"samples":idx+1},f)
self.load()
return True
def identify(self, frame, bbox=None, threshold=0.55):
enc=self._encoding(frame,bbox)
if enc is None:
return {"name":"unknown","confidence":0.0}
best_name="unknown"
best_score=-1
scores={}
for name, enc_list in self.known.items():
sims=[]
for ke in enc_list:
if ke.shape != enc.shape:
min_dim=min(len(ke),len(enc))
sim=float(np.dot(ke[:min_dim], enc[:min_dim]))
else:
if len(enc)==128:
dist=np.linalg.norm(ke-enc)
sim=max(0,1-dist)
else:
sim=float(np.dot(ke,enc))
sims.append(sim)
max_sim=max(sims) if sims else 0
scores[name]=max_sim
if max_sim>best_score:
best_score=max_sim
best_name=name
if best_score<threshold:
return {"name":"unknown","confidence":best_score,"scores":scores}
return {"name":best_name,"confidence":best_score,"scores":scores}
class FaceIdentifier:
def __init__(self, camera_id=0, enrollment_dir="data/faces", threshold=0.55):
self.camera=Camera(camera_id=camera_id)
self.detector=FaceDetector()
self.store=FaceStore(enrollment_dir=enrollment_dir)
self.threshold=threshold
def check(self):
frame=self.camera.capture_frame()
if frame is None:
return {"name":"unknown","face_conf":0.0,"error":"no camera"}
faces=self.detector.detect(frame)
if not faces:
return {"name":"unknown","face_conf":0.0,"faces_found":0,"frame":frame}
faces_sorted=sorted(faces, key=lambda b: b[2]*b[3], reverse=True)
best_bbox=faces_sorted[0]
result=self.store.identify(frame,bbox=best_bbox,threshold=self.threshold)
result["faces_found"]=len(faces)
result["bbox"]=best_bbox
result["face_conf"]=result.get("confidence",0.0)
result["frame"]=frame
return result
def close(self):
self.camera.close()
+7
View File
@@ -0,0 +1,7 @@
#!/usr/bin/env python3
import sys
from jarvis.cli import main
if __name__ == '__main__':
sys.argv[0] = 'jarvis'
main()
+30
View File
@@ -0,0 +1,30 @@
[project]
name = "jarvis-v2"
version = "0.2.0"
description = "Always-listening voice assistant - CPU optimized"
readme = "README.md"
requires-python = ">=3.10"
dependencies = [
"numpy<2",
"sounddevice",
"soundfile",
"openwakeword",
"faster-whisper",
"onnxruntime",
"torch",
"torchaudio",
"silero-vad",
"webrtcvad-wheels",
"scipy",
"piper-tts",
"scikit-learn",
"opencv-python-headless",
"pyyaml",
"requests",
"python-dotenv",
"ollama",
"tqdm",
]
[project.scripts]
jarvis = "jarvis.cli:main"
+16
View File
@@ -0,0 +1,16 @@
#!/bin/bash
# Build PortAudio locally in user space if system lib missing
set -e
cd /tmp
if [ -d portaudio ]; then
echo "portaudio src exists, skipping clone"
else
git clone https://github.com/PortAudio/portaudio.git
fi
cd portaudio
./configure --prefix=$HOME/.local --without-jack
make -j$(nproc)
make install
echo "Installed to ~/.local/lib"
echo "Add to LD_LIBRARY_PATH: export LD_LIBRARY_PATH=\$HOME/.local/lib:\$LD_LIBRARY_PATH"
ldconfig -n $HOME/.local/lib || true
+10
View File
@@ -0,0 +1,10 @@
import sys
from jarvis.cli import main
if len(sys.argv) == 1:
print("Usage: python scripts/enroll_face.py --name NAME")
sys.exit(1)
if sys.argv[1] != "enroll-face" and not sys.argv[1].startswith("--"):
sys.argv.insert(1, "enroll-face")
elif sys.argv[1].startswith("--"):
sys.argv.insert(1, "enroll-face")
main()
+13
View File
@@ -0,0 +1,13 @@
import sys
from jarvis.cli import main
# Pass through args
sys.argv[0] = "jarvis"
if len(sys.argv) == 1:
print("Usage: python scripts/enroll_speaker.py --name NAME --mic")
sys.exit(1)
# ensure first arg is enroll-speaker
if sys.argv[1] != "enroll-speaker" and not sys.argv[1].startswith("--"):
sys.argv.insert(1, "enroll-speaker")
elif sys.argv[1].startswith("--"):
sys.argv.insert(1, "enroll-speaker")
main()
+16
View File
@@ -0,0 +1,16 @@
#!/bin/bash
set -e
DIR="$(cd "$(dirname "$0")/.." && pwd)"
SERVICE_FILE="$DIR/systemd/jarvis.service"
DEST="$HOME/.config/systemd/user/jarvis.service"
mkdir -p "$HOME/.config/systemd/user"
cp "$SERVICE_FILE" "$DEST"
echo "Reloading systemd user daemon"
systemctl --user daemon-reload
echo "Enabling jarvis service"
systemctl --user enable jarvis.service
echo "To start: systemctl --user start jarvis"
echo "To logs: journalctl --user -u jarvis -f"
echo "To stop: systemctl --user stop jarvis"
+4
View File
@@ -0,0 +1,4 @@
from jarvis.cli import main
import sys
sys.argv = ["jarvis","test-mic"]
main()
+50
View File
@@ -0,0 +1,50 @@
"""Test STT with mic"""
from jarvis.audio.capture import AudioCapture
from jarvis.stt.faster_whisper_engine import FasterWhisperSTT
from jarvis.audio.vad import create_vad
from jarvis.config import load_config
import numpy as np, time, collections
cfg = load_config(None)
vad = create_vad(cfg)
stt = FasterWhisperSTT(model=cfg.get('stt',{}).get('model','base'), device='cpu', compute_type='int8')
cap = AudioCapture()
cap.start()
print("Speak, with 1.5s silence to trigger transcription. Ctrl+C to exit")
buffer = []
silence_start = None
recording = False
try:
while True:
chunk = cap.read(timeout=0.1)
if chunk is None:
continue
is_speech = vad.is_speech(chunk)
if is_speech:
if not recording:
print("Speech start")
recording = True
buffer.append(chunk)
silence_start = None
else:
if recording:
if silence_start is None:
silence_start = time.time()
buffer.append(chunk)
if time.time() - silence_start > 1.2:
# transcribe
audio = np.concatenate(buffer)
print(f"Transcribing {len(audio)/16000:.2f}s...")
res = stt.transcribe(audio)
print(f" => {res.get('text')}")
buffer = []
recording = False
silence_start = None
print("Listening again...")
# else silence, continue
except KeyboardInterrupt:
pass
finally:
cap.stop()
+29
View File
@@ -0,0 +1,29 @@
"""Test wakeword detection without full pipeline"""
from jarvis.audio.capture import AudioCapture
from jarvis.audio.wakeword import WakeWordDetector
from jarvis.audio.vad import create_vad
from jarvis.config import load_config
import time
cfg = load_config(None)
vad = create_vad(cfg)
ww = WakeWordDetector(models=cfg.get('wakeword',{}).get('models',['hey_jarvis']), threshold=0.45)
cap = AudioCapture()
cap.start()
print("Say 'Hey Jarvis'...")
try:
while True:
chunk = cap.read(timeout=0.1)
if chunk is None:
continue
# Optional VAD gate
if not vad.is_speech(chunk):
continue
det = ww.detect(chunk)
if det:
print(f"Wake detected: {det}")
except KeyboardInterrupt:
pass
finally:
cap.stop()
+24
View File
@@ -0,0 +1,24 @@
[Unit]
Description=Jarvis v2 Voice Assistant - Always Listening
After=network.target sound.target
Wants=network.target
[Service]
Type=simple
WorkingDirectory=%h/Projects/jarvis-v2
ExecStart=%h/Projects/jarvis-v2/.venv/bin/jarvis run --config %h/Projects/jarvis-v2/jarvis.yaml
Restart=on-failure
RestartSec=5
Environment=PATH=%h/.local/bin:/usr/bin:/bin
Environment=PYTHONUNBUFFERED=1
# Logging
StandardOutput=journal
StandardError=journal
# Resource limits - protect old iMac
Nice= -5
CPUQuota=200%
[Install]
WantedBy=default.target