Remix.run Logo
WarmWash 8 hours ago

The benchmark also doesn't include speed. You almost think something has gone wrong when using it because it returns full responses so incredibly fast.

sotix 24 minutes ago | parent | next [-]

This one uses that as a priority weight: https://winstonrc.github.io/ai-coding-agents-leaderboard/

scrlk 8 hours ago | parent | prev [-]

Not just speed, also reliability. IME, Gemini's speed and quality doesn't degrade badly during weekday working hours compared to OAI, and especially Anthropic.

ford 7 hours ago | parent [-]

I've had Gemini model API use degrade the most out of OAI/Anthropic/Google (often "over capacity" vs true failures)

Not sure on consumer/product use though

scrlk 7 hours ago | parent [-]

That's interesting to hear. I should have added that I use Gemini through Google AI Studio as my general chat model, which probably explains our wildly different experiences.