Remix.run Logo
dudisubekti a day ago

Artificialanalysis benchmark is a combination of a several benchmarks which might or might not represent realistic coding:

"Artificial Analysis Intelligence Index combines performance across 10 evaluations: AA-Briefcase v1.1, GDPval-AA v2.1, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience, and AA-LCR v1.1."

Not saying it doesnt have any value but it's probably irrelevant if you use these AIs for a specific use case. Like for example Humanity Last Exam tests general knowledge, which is not very useful for coding.

It's best to go to the specific coding benchmarks and compare there.

aucisson_masque 19 hours ago | parent [-]

There is too many money involved, benchmark can’t be trusted.

dudisubekti 14 hours ago | parent [-]

I had a favorite benchmark, SWE-rebench, but sadly it's no longer maintained.

But yeah, I'll just take these benchmarks with a grain of salt. Only hands-on experience matters in the end, and these days it's very easy to switch models.