Remix.run Logo
guessmyname 4 hours ago

> No benchmarks, no info on which models are used, […]

The benchmarks are here → https://echo.tracerml.ai/eval/

They are not good benchmarks but at least they exist.

seizethecheese 3 hours ago | parent [-]

I've been working on a similar project and I found that it's easy to replicate Fable results if you use saturated benchmarks.

In my project, I wasted a huge amount of time trying to improve GPQA Diamond results above ~93% range. I realized my mistake when Fable dropped and made no improvement on this benchmark vs. Opus.

yorwba 31 minutes ago | parent [-]

I wouldn't be surprised if ≈7% of GPQA Diamond questions simply have the wrong answer in the ground truth data, so that getting such a question correct is graded as an error. Most machine-learning benchmarks are rather badly validated.

seizethecheese 29 minutes ago | parent [-]

Yep! I found this interesting article after banging my head against a wall for a long time: https://epoch.ai/gradient-updates/gpqa-diamond-whats-left