| ▲ | simonw 9 hours ago | ||||||||||||||||
This doesn't look like a plateau to me: https://artificialanalysis.ai/evaluations/artificial-analysi... I do agree that they're investing heavily in brute force methods though. I've been trying out GPT-5.6 Sol "Ultra" recently and that thing fires up a bunch of subagents and crunches for hours. | |||||||||||||||||
| ▲ | supermdguy 5 hours ago | parent [-] | ||||||||||||||||
Here's the performance of frontier models without reasoning, to more directly address the claim that raw performance is plateauing: https://artificialanalysis.ai/evaluations/artificial-analysi... I don't have any insider info, but if model sizes actually have increased exponentially since GPT 4.1, there's an argument to be made that there are diminishing returns in scaling pretraining alone. Also interesting thing I haven't noticed before, Opus models have followed a really consistent linear improvement, while it looks like OpenAI struggled with base model performance until 5.5/5.6 (EDIT - 5.5 was their first new pretraining run in over a year). | |||||||||||||||||
| |||||||||||||||||