Remix.run Logo
satvikpendem 8 hours ago

We'll see about that. I suspect benchmaxxing as all the labs do as I haven't found Gemini models to be nearly as good in agentic engineering compared to Claude or GPT models.

NitpickLawyer 8 hours ago | parent | next [-]

If anything, gemini models are the least benchmaxxed out of any lab, IMO.

onlyrealcuzzo 8 hours ago | parent | prev [-]

And the benchmarks agreed with you... until now.

So, yes, maybe it's still not - but this would be the only time it would be highly suspicious / obvious benchmaxxing / obviously bad benchmarks.