Remix.run Logo
webo 5 hours ago

The benchmarks page seems interesting and something I can use to help make an informed decision. Can you talk about how you're measuring some of these? I imagine it needs to involve some human input.

https://benchmarks.speko.ai/turntaking

abdik 3 hours ago | parent [-]

Turn-taking specifically does not need a listening panel, but it is measured mechanically. 200+ real human clips, and we score end-vs-wait decisions: did the model decide the caller finished speaking, or just paused mid-thought. The best detector gets 94.0% of those right; a plain VAD silence timer gets 46.9%. Results are published here: https://benchmarks.speko.ai/turntaking

You are right about human input for naturalness, that one we did not automate away with yet. We run blind A/B listening rounds with native speakers.