Files
whisper-translation/README.md
T
Adolfo Reyna 80f0bf309f feat: Apple Speech v3 + Freeflow polish + draft streaming
- Apple SpeechAnalyzer (macOS 26+) binary: --bench (31x RTF), --pipe
  (persistent process, 150ms finals), --live (word-by-word drafts)
- Pipe protocol: 4-byte BE length + wav payload, emits JSONL
  {event:draft|final, text, isFinal, chunk} — 31 drafts for 6s audio (~60ms granularity)
- engine_apple_transcribe.py: ApplePipeTranscriber with
  transcribe() + transcribe_with_draft_callback(), VAD + draft
  queue, new flags --apple-stream (on), --apple-stream-interval,
  --apple-pipe (on). Fixes PIL/transformers import crash by lazy import.
- main_v3.py: engine selector {whisper,apple}, passthrough translate
  when no -es/-fr/-ar, freeflow flags same as v2
- Freeflow polish: deterministic punctuation commands (comma,
  question mark, new paragraph, at sign), filler stripping,
  <keep> protection, skip-clean heuristic, freeflow/qwen/legacy
  prompt styles. Much better final readability vs raw Apple/Whisper.
- main_v2.py, engine_llm.py, engine_distribute.py: integrate freeflow
- bench: Apple 2.12% WER vs Whisper Small 3.74% (Inscribe), CPU
  0mW ANE (measured via powermetrics), 196M EN cryptex per locale.
- Verified: 31 word-by-word drafts, 2 finals, exit 0, bench regression ok.

Freeflow still much better for final polish — Apple wins on speed
and raw accuracy, freeflow wins on readable paragraph output.

Co-authored-by: internal-model
2026-07-13 21:08:18 -04:00

141 lines
6.8 KiB
Markdown

# Python Whisper Live Transcription & Translation
A real-time, low-latency audio transcription and translation tool utilizing OpenAI's Whisper (via `mlx-whisper` for Apple Silicon optimization), Silero VAD for speech detection, and Helsinki-NLP's Opus-MT models for translation.
## Features
- **Live Transcription:** Real-time speech-to-text with automatic language detection.
- **On-the-fly Translation:** Bridge translations to English and then to target languages (Spanish, Arabic, French, etc.).
- **Voice Activity Detection (VAD):** Intelligent audio buffering using Silero VAD to process only actual speech.
- **Apple Silicon Optimized:** Uses MLX for high performance on Mac (MPS).
- **Background Ingestion:** Optional background thread to send JSON payloads to a remote server.
- **Configurable:** Command-line parameters to select languages, devices, and ingestion.
## Installation
1. **Clone the repository:**
```bash
git clone <repository-url>
cd pythonwhisper
```
2. **Install dependencies:**
Ensure you have Python 3.9+ and the required libraries:
```bash
pip install mlx-whisper numpy sounddevice torch requests transformers silero-vad
```
*Note: On Apple Silicon, ensure `mlx` and `torch` with MPS support are correctly installed.*
## Usage
Run the script using `python3 transcribe.py` with optional flags.
### LLM Prompt Testing
You can now send the line-correction or paragraph-refinement prompts directly without starting live transcription. Every LLM request is appended as JSON Lines to `logs/llm_requests.jsonl` by default.
- **Test the line-correction prompt:**
```bash
python3 main_v2.py --post-correct-model qwen3.5:0.8b --llm-test-line "so um we should probably ship it tomorrow" --llm-test-prev1 "We finished the staging deploy." --llm-test-prev2 "QA signed off this morning."
```
- **Test the paragraph-refinement prompt:**
```bash
python3 main_v2.py --post-correct-model qwen3.5:0.8b --llm-test-context "We reviewed the launch checklist." --llm-test-segments "we confirmed monitoring we confirmed rollback and then talked about the release window"
```
- **Change where requests are logged:**
```bash
python3 main_v2.py --llm-request-log-path tmp/my_llm_requests.jsonl --llm-test-line "example text"
```
- **Replay the most recent logged request:**
```bash
python3 main_v2.py --llm-test-from-log -1
```
- **Replay a logged request using the original model from the log entry:**
```bash
python3 main_v2.py --llm-test-from-log 12 --llm-test-use-logged-model
```
Each log entry includes the timestamp, provider, model, mode, temperature, and full `messages` payload that was sent to the LLM.
### Freeflow-style Local Polish
This implementation includes a Python port of the portable parts of Freeflow's polish pipeline:
- spoken punctuation commands such as `comma`, `question mark`, `new paragraph`, `at sign`, and `hashtag`
- filler/noise stripping before any LLM call
- protected `<keep>...</keep>` symbols so the LLM does not reinterpret dictated symbols
- a clean-transcript skip heuristic for low-latency local runs
- Freeflow-inspired English/minimal/Qwen line-polish prompts
Run deterministic cleanup only:
```bash
python3 main_v2.py --post-correct
```
Run deterministic cleanup plus a local Ollama line-polish model when needed:
```bash
python3 main_v2.py --post-correct-llm --post-correct-model qwen3.5:0.8b --post-correct-prompt-style qwen
```
Force every line through the LLM, even if the deterministic output already looks clean:
```bash
python3 main_v2.py --post-correct-llm --no-post-correct-skip-clean
```
### Common Commands
- **List available audio devices:**
```bash
python3 transcribe.py -l
```
- **Caption system audio (speakers) using a loopback device:**
```bash
python3 transcribe.py --loopback -es
```
- **Transcribe and translate to Spanish (screen only):**
```bash
python3 transcribe.py -es
```
- **Enable Spanish, Arabic, and French with English bridging, and send to server:**
```bash
python3 transcribe.py -es -ar -fr -en -i
```
- **Use a specific input device (e.g., index 3) and translate to Spanish:**
```bash
python3 transcribe.py -d 3 -es
```
### Arguments
- `-es`: Enable Spanish translation.
- `-en`: Enable English (shows original if detected, or bridged if not).
- `-ar`: Enable Arabic translation.
- `-fr`: Enable French translation.
- `-i`, `--ingest`: Enable data transmission to the remote server.
- `-l`, `--list-devices`: Show available audio devices and exit.
- `-d`, `--device [ID]`: Input device index (bypasses selection prompt).
- `--loopback`: Automatically select a loopback device (e.g., BlackHole, Stereo Mix) to caption system audio.
- `-q`, `--quantize`: Use 4-bit quantized Whisper model for faster transcription (Strategy 3).
- `-s`, `--stream`: Enable real-time streaming transcription/draft mode (Strategy 1).
- `-c`, `--context`: Enable prompt caching/rolling context to help the model maintain sentence continuity across chunks.
- `--lang [CODE]`: Hardcode the source language (e.g., `en`, `es`) to bypass automatic language detection for faster processing.
- `--silence [MS]`: Set the minimum silence duration in milliseconds to end a chunk. Defaults to 1000ms. Increase to force longer sentences before translation.
- `--max-buffer [SEC]`: Maximum buffer duration in seconds before forcing a flush (default: 20).
- `--channels [N]`: Number of input channels (default: 1).
- `--pick-channel [0|1]`: If stereo, select channel 0 (Left) or 1 (Right) to focus transcription.
- `--filter-lang`: If used with `--lang`, discards segments that do not match the target language.
---
## Technical Note: Universal Translation Models
While the current implementation uses specialized, per-language models from the **Helsinki-NLP Opus-MT** project (e.g., `opus-mt-en-es`, `opus-mt-en-ar`), there is an alternative approach: **Universal Models**.
### Universal Model Alternative (e.g., Meta's NLLB-200)
The current per-language model approach is highly accurate and memory-efficient if you only need 1 or 2 target languages. However, if you require support for many languages simultaneously, loading multiple specialized models can consume significant RAM/VRAM.
We have the option to switch the translation engine to a single, universal model such as **NLLB-200 (No Language Left Behind)**:
- **Model ID:** `facebook/nllb-200-distilled-600M`
- **Benefits:**
- Supports over **200 languages** in a single model.
- Simplified code: no need to load/manage multiple model objects.
- More efficient for complex multilingual environments.
- **Trade-off:** Slightly higher memory footprint for the single model compared to a single specialized model, but more efficient than 3+ specialized models.
If you wish to switch to a universal model, the `transcribe.py` logic can be updated to use a single `M2M100` or `NLLB` pipeline instead of the current `MarianMT` loop.