Speakrail is an open‑source, fully local full‑duplex voice agent that runs on a single RTX 4090 and matches GPT‑Live performance on select benchmarks. It combines Voxtral Realtime with a turn‑taking head, a microturn‑fine‑tuned Gemma 4 12B language model, and Breeze TTS 2 for speech synthesis. The system eliminates non‑local pipeline components to lower latency and address privacy concerns. Source code is available at github.com/speakrail/speakrail.
Read original
reddit/r/LocalLLM