| ▲ | jackb4040 an hour ago | |
Yes, diminishing returns. Not overall, they've still been able to create more intelligent models even up to today. But the strategy for scaling that intelligence has shifted. From the initial ChatGPT release to GPT-4.1, they were basically scaling up compute training compute / model size. Then 4.5 flopped, while o1 demonstrated that gains could continue by reasoning (scaling up compute at inference time). o1 is now the ancestor of all their flagship models from GPT-5 on. This is why I'm trying so hard to drill down on the theory of scaling, and not just talk about improvement in general, hand-wavy terms. If the bottleneck of current scaling strategies is training data, or something fundamental about the model architecture, then just throwing more harnessed chatbots at it won't lead to an exponential increase in performance. Now you could argue that the AI we have now will help us find that change in architecture, and I would agree. But that means we're firmly outside the singularity for the time being, and what people are in fact talking about is a hypothetical. | ||
| ▲ | famouswaffles 25 minutes ago | parent [-] | |
>Then 4.5 flopped, while o1 demonstrated that gains could continue by reasoning (scaling up compute at inference time). o1 is now the ancestor of all their flagship models from GPT-5 on. That's not quite right. They are still scaling model size and have had several new base pre-trains, just nothing so big as 4.5 (as far as we're aware). o1/4o has not been the base for some time now. Data is obviously a bottleneck for some regimes and LLMs will have to get their hands dirty experimenting but it doesn't look like an insurmountable wall either. | ||