| ▲ | webo 5 hours ago | |
The benchmarks page seems interesting and something I can use to help make an informed decision. Can you talk about how you're measuring some of these? I imagine it needs to involve some human input. | ||
| ▲ | abdik 3 hours ago | parent [-] | |
Turn-taking specifically does not need a listening panel, but it is measured mechanically. 200+ real human clips, and we score end-vs-wait decisions: did the model decide the caller finished speaking, or just paused mid-thought. The best detector gets 94.0% of those right; a plain VAD silence timer gets 46.9%. Results are published here: https://benchmarks.speko.ai/turntaking You are right about human input for naturalness, that one we did not automate away with yet. We run blind A/B listening rounds with native speakers. | ||