| ▲ | satvikpendem 8 hours ago | |
We'll see about that. I suspect benchmaxxing as all the labs do as I haven't found Gemini models to be nearly as good in agentic engineering compared to Claude or GPT models. | ||
| ▲ | NitpickLawyer 8 hours ago | parent | next [-] | |
If anything, gemini models are the least benchmaxxed out of any lab, IMO. | ||
| ▲ | onlyrealcuzzo 8 hours ago | parent | prev [-] | |
And the benchmarks agreed with you... until now. So, yes, maybe it's still not - but this would be the only time it would be highly suspicious / obvious benchmaxxing / obviously bad benchmarks. | ||