| ▲ | guessmyname 4 hours ago | ||||||||||||||||
> No benchmarks, no info on which models are used, […] The benchmarks are here → https://echo.tracerml.ai/eval/ They are not good benchmarks but at least they exist. | |||||||||||||||||
| ▲ | seizethecheese 3 hours ago | parent [-] | ||||||||||||||||
I've been working on a similar project and I found that it's easy to replicate Fable results if you use saturated benchmarks. In my project, I wasted a huge amount of time trying to improve GPQA Diamond results above ~93% range. I realized my mistake when Fable dropped and made no improvement on this benchmark vs. Opus. | |||||||||||||||||
| |||||||||||||||||