Files
mcp_screen/websocket_implementation_plan.md

3.4 KiB

WebSocket Audio Streaming Implementation Plan

Implement a WebSocket-based real-time voice integration between the ESP32 board and the Hermes Voice Gateway (Phase 1: WebSocket streaming with server-side batch fallback) to solve the RAM/Flash storage constraints and lower latency.

User Review Required

Important

  • This plan adds a new route /api/esp32/voice/ws to the Hermes Gateway. The old POST route /api/esp32/voice will remain completely intact to act as a stable fallback.
  • A lightweight WebSocket client websocket_client.py will be added to the MicroPython lib/ directory since standard MicroPython doesn't bundle a WebSocket client by default.

Open Questions

Note

None at the moment. The proposed wire protocol and incremental phased approach minimize risk.


Proposed Changes

1. Hermes Voice Gateway

[MODIFY] api_server_endpoint.py

  • Add a _handle_esp32_voice_websocket handler method.
  • Listen for a JSON start frame to read parameters like device_id, reply_mode, and screen_url.
  • Accept incoming binary frames (raw PCM s16le mono 16 kHz chunks) and write them directly to a temporary WAV file on the host's disk.
  • Upon receiving the JSON end frame (or a disconnection/timeout):
    • Close the WAV file writer.
    • Invoke transcribe_audio(wav_path) using the local faster-whisper engine.
    • Run the agent turn via _run_esp32_voice_turn.
    • Synthesize TTS audio and send the resulting WAV file back as a binary frame to the client.
  • Update register_esp32_voice_route to bind the new WebSocket endpoint:
    app.router.add_get("/api/esp32/voice/ws", api_server._handle_esp32_voice_websocket)
    

[MODIFY] README.md

  • Update documentation to cover the new /api/esp32/voice/ws WebSocket client contract and protocol.

2. ESP32 MicroPython Firmware

[NEW] websocket_client.py

  • Implement a minimal, robust, memory-efficient WebSocket client helper class (WebSocketClient) in MicroPython.
  • Support HTTP handshake, sending/receiving text frames (opcode 1), and binary frames (opcode 2) with proper masking (client-to-server requirement).

[NEW] demo_websocket_voice.py

  • Create a dedicated firmware script demonstrating how the board communicates with the new gateway:
    • Connect to ws://<ip>:8642/api/esp32/voice/ws.
    • Send the JSON start frame.
    • Record continuously from I2S and stream 100ms chunks (3200 bytes mono) over the socket.
    • Send the JSON end frame on screen release.
    • Listen for the server's response WAV frame and play it using the ES8311 speaker codec.

Verification Plan

Automated Tests

  • Run python3 hermes_voice_gateway/smoke_test.py to ensure local helper methods (normalization, MIME mapping) remain fully functional.

Manual Verification

  1. Run the updated Hermes server with the new WebSocket route registered.
  2. Run demo_websocket_voice.py on the ESP32 board.
  3. Verify that the ESP32 connects, streams audio chunks on-the-fly without running out of RAM (SRAM memory checks via gc.mem_free()), transcribes the speech, and plays back the agent's voice response.