Remix.run Logo
pants2 3 hours ago

We're gonna need some new benchmarks...

ARC-AGI-3 might be the only remaining benchmark below 50%

Leynos 2 hours ago | parent | next [-]

Opus 4.6 currently leads the remote labor index at 4.17. GPT-5.4 isn't measured on that one though: https://www.remotelabor.ai/

GPT 5.4 Pro leads Frontier Maths Tier 4 at 35%: https://epoch.ai/benchmarks/frontiermath-tier-4/

randomtoast 2 hours ago | parent | prev [-]

[dead]