| ▲ | astro1234 4 hours ago | ||||||||||||||||||||||
I work in AI evaluation, lots of problems and leakage is an issue as is ecological validity, but they definitely do not explain the progress we see. I think Epoch has the best analysis I’ve seen on evaluation trends; they use IRT to basically model a variety of benchmark difficulties, and then model a capability parameter for each model. This is as robust a sort of “meta-study” of evaluations as I’ve seen and the trend in capabilities show no sign of slowing down. So I think people’s feelings clash with reality, and that’s because releases are more frequent and the jumps between releases are smaller, but the growth in capabilities _over time_ has not changed for the better or worse over a very very long period of time. | |||||||||||||||||||||||
| ▲ | otterdude 4 hours ago | parent [-] | ||||||||||||||||||||||
Benchmarks saturate around 80-90%? This is not "Acing" a test, this is hitting a wall. | |||||||||||||||||||||||
| |||||||||||||||||||||||