Remix.run Logo
▲ CuriouslyC 5 hours ago

It's not so much that they're hitting a plateau in capability, as we're saturating long horizon benchmarks and it's not greatly improving general usability. On the other hand, newer models have been amazing for people interested in 3d, graphics, video editing, etc. The difference between Opus 5.5/Astra and earlier models is night and day even if for many coding tasks they're not a revolution.

▲omalled 3 hours ago | parent [-]

I agree that they're not hitting a plateau and I see it in my reserach. I had a math/code benchmark paper [1] at NeurIPS last year that is still unsaturated. At the time of writing the paper, the best model was o3, which was scoring 3-4%. By the time NeurIPS came around, GPT-5.2 was the latest model but it was getting similar scores to o3. The models were still in the flat part of the usual hockey stick curve. The newer models are getting into the steep part. I evaluated gpt-5.6-sol+codex a week or two ago and it got ~16%. Astra+codex got ~24%.

On some tasks in this benchmark, the models seem to be coming up with novel solutions. For example, Astra came up with a relatively simple formula for a sequence that only has 8 terms in OEIS and is considered "hard" [2]. It produced a lean proof that the formula is correct, but I'm just starting to learn lean and don't have enough expertise to check it.

[1] https://proceedings.neurips.cc/paper_files/paper/2025/hash/c... [2] https://oeis.org/A000530