Remix.run Logo
mariano54 3 hours ago

Just added this to my benchmark site: https://multilingualsttbench.com/

It doesn't reach the frontier in either latency or accuracy for ai multilingual conversations.

adamgoodapp 2 hours ago | parent | next [-]

Thanks for this, really helpful.

I would also like to see benchmark for translation. I'm looking for live translated subtitles so my Japanese wife can enjoy any show with out waiting months for official VOD streams to release them.

Kokouane 2 hours ago | parent | prev [-]

I'm confused, doesn't your leaderboard clearly show it is the most accurate model? It's number one in the leaderboard. Am I missing something?

Kokouane 2 hours ago | parent [-]

Figured it out. 3.5 Flash and 3.5 Transcribe are different models