| ▲ | abdik 3 hours ago | |
Actually, we measured exactly this recently. The strongest open-weights speech-to-speech model we have run is NVIDIA's NemotronLabs VoiceChat 11B - no provider serves it, so we hosted it ourselves and ran the same scripted call every model on our board gets. Remarkably stable, 40+ sessions with zero errors - but by turn thirteen it was answering nine tries in ten without saying a word. Stable engine, but degrades on long calls. Wider picture: we rebuilt the speech-to-speech test around a hard 20-turn concierge call, and today only the closed models finish it cleanly. Board: https://benchmarks.speko.ai/s2s shared some write-up: https://benchmarks.speko.ai/blog/can-s2s-replace-the-cascade And jakswa's Gemma datapoint matches what we measure, the fastest rows on our LLM board are the small models like Gemma 4 is performing incredibly well for voice agents | ||