Remix.run Logo
simonw 9 hours ago

This doesn't look like a plateau to me: https://artificialanalysis.ai/evaluations/artificial-analysi...

I do agree that they're investing heavily in brute force methods though. I've been trying out GPT-5.6 Sol "Ultra" recently and that thing fires up a bunch of subagents and crunches for hours.

supermdguy 5 hours ago | parent [-]

Here's the performance of frontier models without reasoning, to more directly address the claim that raw performance is plateauing:

https://artificialanalysis.ai/evaluations/artificial-analysi...

I don't have any insider info, but if model sizes actually have increased exponentially since GPT 4.1, there's an argument to be made that there are diminishing returns in scaling pretraining alone.

Also interesting thing I haven't noticed before, Opus models have followed a really consistent linear improvement, while it looks like OpenAI struggled with base model performance until 5.5/5.6 (EDIT - 5.5 was their first new pretraining run in over a year).

simonw 4 hours ago | parent [-]

The trend I've found most interesting is models of the same size getting better.

I'm very much looking forward to seeing how Qwen 3.8 27B compares to Qwen 3.6 27B next week, for example.

And the latest DeepSeek v4 Flash has extremely impressive performance for a 304B model.

asadotzler 2 hours ago | parent [-]

The trends you found don't support my goals so I've got some other trends I find more interesting than yours.