| ▲ | seizethecheese 3 hours ago | |||||||
I've been working on a similar project and I found that it's easy to replicate Fable results if you use saturated benchmarks. In my project, I wasted a huge amount of time trying to improve GPQA Diamond results above ~93% range. I realized my mistake when Fable dropped and made no improvement on this benchmark vs. Opus. | ||||||||
| ▲ | yorwba 32 minutes ago | parent [-] | |||||||
I wouldn't be surprised if ≈7% of GPQA Diamond questions simply have the wrong answer in the ground truth data, so that getting such a question correct is graded as an error. Most machine-learning benchmarks are rather badly validated. | ||||||||
| ||||||||