| ▲ | mixdup 5 hours ago |
| Another piece of evidence on the pile that the sudden panic and desire to "slow down" is because they're hitting the plateau on capability Which, honestly, is fine. A lot of juice to squeeze in efficiency and even if models got zero more capable, making the capability that is already here cheaper is a huge win for everyone (except Nvidia) |
|
| ▲ | luma 5 hours ago | parent | next [-] |
| Some version of this claim has been made for the past 4 years. There's a data cliff, there's no more compute to buy, the financials don't make sense and all of these orgs will be out of business by end of quarter. Not once has any of these predictions come true, the pace of progress has continued on it's exponential trajectory since ChatGPT first came to the public's attention. So why now? What is special about today that suggests all of this is coming to a screeching halt despite all evidence to the contrary? |
| |
| ▲ | OliveronData 4 hours ago | parent | next [-] | | > ... the pace of progress has continued on it's exponential trajectory since ChatGPT first came to the public's attention. Did it? Model wise? I would understand agents wise, sure. But model wise? The attention to detail from the model? The ability to recall minute things? Improvements are there, yes, but mostly on Fable and Astra. Opus still isn't as attentive as Fable in long term writing for example. Sure, Opus 5.5 benchmarks better than Fable. Sure. But is that the model, or is that the RL for agentic work? From where I'm standing, the model work has not been exponential at all, and more and more it looks like the latest and greatest is getting too expensive too fast. Both 5.5 and 5.6 chat models got nerfed, actually nerfed not the tea leaves kind. In mid 5.5 cycle the chat model lost the ability to substitute names if given an outline. 5.6 cycle the chat model lost the ability to use paragraphs after a few hundred words (coinciding with Chat/Work split). There's a race from OpenAI to serve dumber models on chat. I'm not even sure who they are racing against, but the fact that Astra, Sol 6.0, and now Sol 6.1 not being available for chat, should tell you that those models are expensive, and not the kind of models that can be freely "chatted" with on a subscription. OpenAI much prefers you use Work and limit the chat usage, much like Grok and Claude. I'm guessing they will announce that later during the dev days. That could be cost cutting too, true, but really? That's the only explanation? And nothing else? Sure, the progress did not stop. But it is nowhere near close being exponential when it comes to LLMs themselves. Agents are separate. | | |
| ▲ | luma 4 hours ago | parent [-] | | I didn't use the word LLM. I'm talking AI capability, you're focused on this or that current approach to AI. I think it's fair to assume that the approach will change as new ideas are learned, new and more hardware will be purchased and applied to the problem, and then capabilities will (for now) continue on their exponential curve, same as it has gone for the past several years. These things are knocking down Millennium Prize problems while a substantial subset of commenters here are still thinking about stochastic parrots. | | |
| |
| ▲ | chamomeal 2 hours ago | parent | prev | next [-] | | Has it been exponential this whole time? I feel like GPT-4 was pretty dang good. Maybe it’s rose tinted glasses cause I could finally have a bot write my dockerfiles and bash scripts, which knocked my socks off | |
| ▲ | john_strinlai 4 hours ago | parent | prev | next [-] | | >Not once has any of these predictions come true, the pace of progress has continued on it's exponential trajectory since ChatGPT first came to the public's attention. do you think it will be exponential forever? | | |
| ▲ | RobCat27 4 hours ago | parent [-] | | I think we'll eventually hit an information theoretic type of wall with physical hardware and GPUs and need a similar AI breakthrough as well as the development refinement of logical/physical qubits in the quantum computing space with some analogue to the transformer architecture to continue accelerating. However, I think there must be many years of development and refinement that can take place before that paradigm shift to overcome the physical compute wall is necessary. This is just my theory, but I'm young enough that I'm expecting with the rate that we are advancing, I will see AI / LLM analogues developed and run on a quantum computer in my lifetime. | | |
| ▲ | spathi_fwiffo 3 hours ago | parent | next [-] | | I think the bottleneck will be the current one. Fabs. Either needing more fabs, new types of fabs, retooling existing fabs. All of that takes years. maybe we can design our way out of that too. But, I suppose that would be the similar breakthrough you are mentioning. | |
| ▲ | 2 hours ago | parent | prev [-] | | [deleted] |
|
| |
| ▲ | trentnix 4 hours ago | parent | prev | next [-] | | Yep. I've made the claim (and been wrong). I was convinced the data cliff was going to be a real problem. Now I feel like we are on the cusp of having Tony Stark's Jarvis at our fingertips. What a time to be alive. | | |
| ▲ | neta1337 3 hours ago | parent [-] | | Incredible how many times I read similar comments over the years, containing 'on the cusp' and 'what a time to be alive'. Indeed, what a time - not a single user-facing thing on the internet has improved since then, considering the power tool we got. The most used web services get drowned in generated stuff and so are the users | | |
| ▲ | trentnix 3 hours ago | parent [-] | | Not a single thing? In my house, we are using LLMs to: - plan youth soccer practices - develop well-formatted soccer game substitution schedules - build and ship software in languages I haven't used in 25 years on platforms I've never programmed for - do meal planning and build shopping lists - prepare grocery shopping carts - solicit medical advice - perform Garmin watch data analysis - administer devices (with SSH access) using natural language - avoid counterfeit soccer jersey purchases - create "Warrior Cat" graphic novels - make cartoon strips - troubleshoot appliances - manage finances - review accounting ledgers - diagnose malware infections - so much more And we do it all from a simple prompt that we can talk to if we choose. I've built more (and better) software in the past month than I did in any given year in the 30+ years I've been programming. I can understand pessimism regarding how this affects society. I can understand pessimism regarding how this gets abused. But for the life of me there's no good reason at all to be pessimistic about how quickly this has improved. | | |
| ▲ | FiberBundle 2 hours ago | parent [-] | | > I've built more (and better) software in the past month than I did in any given year in the 30+ years I've been programming I feel similarly, but I think it's a valid question. Why is all the software I'm using not getting better? To be honest, I feel it's more buggy than it's ever been. |
|
|
| |
| ▲ | digdugdirk 4 hours ago | parent | prev | next [-] | | The difference now is that they've hit the "good enough" point. LLMs are a tool, and that tool is useful but not incredibly valuable unto itself. To make a manufacturing analogy - ChatGPT was a manual machining mill, and in the years after we've gone from that to a 3-axis CNC mill. Now we've added a 4th and 5th axis, which is great for the 2% of parts that need that functionality. But the big win was that initial jump from manual control to CNC. Why would I pay an extra $2 million for my CNC machine when I could just design my parts to be simpler to produce instead? The AI labs are trying to make these incredibly complex tools, but the market doesn't want/need them so they're competing on price for the tools that people do use. By selling their metaphorical CNC machines for half of what they cost to produce. Oh, and we've bet the entire economy on the hope that fancier CNC machines will magically solve all our problems in all industries, from healthcare to the legal system. So - will AI progress continue to improve? Sure. Will we continue lighting money on fire in order to make it happen? That remains to be seen. | | |
| ▲ | famouswaffles 4 hours ago | parent | next [-] | | >The difference now is that they've hit the "good enough" point. In some aspects sure, but in others no. Open AI's goal is to build "highly autonomous systems that outperform humans at most economically valuable work." and Astra was a big jump in that. There still isn't a better model for computer use and vision/spatial work. Driving, Operating Robots, Video Editing, 3D modelling, graphics are all things Astra was >>> at than any other model. I'm sure you don't care about any of that so it's easy enough to slip by you but this analogy - "Now we've added a 4th and 5th axis, which is great for the 2% of parts that need that functionality." is dead wrong. | |
| ▲ | willchis 4 hours ago | parent | prev [-] | | This is how I feel about it. I've stopped looking at all the scores of new releases and just look at the price to see how much usage I can get in a month. Seems like I'm not the only one either, from comments above like > "Opus 5.5 is so good that I don't want it to be replaced anytime soon. Stop training models[...]"_ |
| |
| ▲ | interestpiqued 4 hours ago | parent | prev | next [-] | | 4 years is not that long in the grand scheme of things to be fair | |
| ▲ | dgellow 5 hours ago | parent | prev | next [-] | | Those points were true at the time and most are still true now. But they aren’t predictions. - it’s correct there isn’t much fresh data anymore - it’s correct that compute is scarce, that was 100% the case and a huge issue at the beginning of the year, it is better now but still scarce, and hardware is now way, way more expensive - it’s correct the finances don’t make sense But there is no way to know when a bubble pop, because it’s a psychological phenomenon across an extremely complicated distributed system (ie the stock and bonds markets) | | |
| ▲ | agoodusername63 12 minutes ago | parent | next [-] | | The amount of irrationality I see in the economy with AI makes me more convinced that the wall street bankrollers know very well they're throwing money into a pit, but it's a pit they're gambling will turn into some world hunger ending AI (that will somehow also keep them making money off of scarcity) never mind that theres no guarantee we'll get that mythical AI. Never mind that the societal reformations would also impact their revenue numbers. | |
| ▲ | moosehater 4 hours ago | parent | prev | next [-] | | I was thinking the same thing in terms of running out of data a few months ago. But aren't most gains in the past year+ due to reinforcement learning in some form? Which doesn't need "fresh data" per se, as the model effectively creates the data as it goes. As long as engineers can come up with proper environments, tasks/goals, rewards, and actions, I don't really see data being a limit to model improvement in an agentic sense. Maybe as a knowledge base | |
| ▲ | JacobAsmuth 4 hours ago | parent | prev | next [-] | | The new hardware (TPU v8 and VR) are more expensive but they are significantly cheaper per flop. e.g. many multiples more performance for only 2x the price. If I have some ML workload to run I can buy $x of Blackwell chips or I can buy significantly less $ worth of Vera Rubin chips to get the same performance. That's the key thing to keep in mind when you're talking about financials. | |
| ▲ | dumberquestions 4 hours ago | parent | prev [-] | | [dead] |
| |
| ▲ | dcchambers 3 hours ago | parent | prev [-] | | [dead] |
|
|
| ▲ | CuriouslyC 5 hours ago | parent | prev | next [-] |
| It's not so much that they're hitting a plateau in capability, as we're saturating long horizon benchmarks and it's not greatly improving general usability. On the other hand, newer models have been amazing for people interested in 3d, graphics, video editing, etc. The difference between Opus 5.5/Astra and earlier models is night and day even if for many coding tasks they're not a revolution. |
| |
| ▲ | omalled 3 hours ago | parent [-] | | I agree that they're not hitting a plateau and I see it in my reserach. I had a math/code benchmark paper [1] at NeurIPS last year that is still unsaturated. At the time of writing the paper, the best model was o3, which was scoring 3-4%. By the time NeurIPS came around, GPT-5.2 was the latest model but it was getting similar scores to o3. The models were still in the flat part of the usual hockey stick curve. The newer models are getting into the steep part. I evaluated gpt-5.6-sol+codex a week or two ago and it got ~16%. Astra+codex got ~24%. On some tasks in this benchmark, the models seem to be coming up with novel solutions. For example, Astra came up with a relatively simple formula for a sequence that only has 8 terms in OEIS and is considered "hard" [2]. It produced a lean proof that the formula is correct, but I'm just starting to learn lean and don't have enough expertise to check it. [1] https://proceedings.neurips.cc/paper_files/paper/2025/hash/c...
[2] https://oeis.org/A000530 |
|
|
| ▲ | sebzim4500 5 hours ago | parent | prev | next [-] |
| Is there anything that could happen that you wouldn't use as evidence that they are hitting a plateau? It just seems like these claims are constant and looking back the calls of 'plateau' between 2023 and 2025 were clearly false, why should we think it's different now? |
|
| ▲ | LPisGood 5 hours ago | parent | prev | next [-] |
| Nvidia can start putting weights in silicon if model development slows down. |
|
| ▲ | theturtletalks 5 hours ago | parent | prev | next [-] |
| I think they are hitting compute restrictions. And buying compute right now can be 3-4X. And the costs are increasing. If they train a larger model and demand is high, that’s a lot of compute for Codex subscriptions, which is a loss leader for them. Especially Pro 20X which they just nerfed to 10X. |
|
| ▲ | serf 5 hours ago | parent | prev | next [-] |
| >Another piece of evidence on the pile that the sudden panic and desire to "slow down" is because they're hitting the plateau on capability if true then LLM related AI (post-post AI winter AI?) is probably one of the fastest inception-to-plateau tech sectors to have ever existed. We're still improving transistors on a somewhat routine basis. |
| |
| ▲ | mixdup 5 hours ago | parent [-] | | The plateau doesn't have to be perfectly flat, but it's not a straight line upward anymore either (kind of like our work on transistors, where we've kind of hit the bounds of speed in clock cycles but are improving on miniaturization and power efficiency) | | |
| ▲ | password54321 5 hours ago | parent [-] | | It took 6 years to solve ARC-AGI 1, 1 year to solve ARC-AGI 2 and 6 months to solve ARC-AGI 3. | | |
| ▲ | delillos 5 hours ago | parent [-] | | Those version numbers don't necessarily correspond to equal increases in "difficulty", though. | | |
| ▲ | password54321 5 hours ago | parent | next [-] | | Correct, the benchmark became exponentially more difficult as it progressed from pattern matching puzzles to games. | |
| ▲ | JacobAsmuth 4 hours ago | parent | prev [-] | | Very true, the sharp increase in difficulty (as measured by human passrate plummeting from 1->2 and again from 2->3) gives an even more stark view of AI capabilities over time. |
|
|
|
|
|
| ▲ | colechristensen 5 hours ago | parent | prev | next [-] |
| >Another piece of evidence on the pile that the sudden panic and desire to "slow down" is because they're hitting the plateau on capability I think it's more a token-cost-demand plateau. They've reached the scale and investor trillions to which they can't 10x the hardware cost of inference any more. They can't afford to compete by eating costs and there isn't appetite for more expensive inference. So in order that they don't bankrupt each other they're looking for the legal cartel behavior coordinating a stop to growth by convincing governments to regulate them into stopping. There's a lot of juice to squeeze in efficiency but only so much whereas it seemed like capability was going to continue to scale with parameter count. Maybe it's good news for everyone that model capability is now going to scale on semiconductor cost meaning huge players are going to be very motivated to make semiconductors cheap. |
|
| ▲ | semiquaver 5 hours ago | parent | prev | next [-] |
| What universe do you live in that you can look at the past six months and see anything like a plateau in capability? Edit: removed a comment that was uncharitable and rude, for which I apologize. |
| |
| ▲ | arctic-true 5 hours ago | parent | next [-] | | Most of the impressive accomplishments we’ve seen in the last few months have been the result of huge agent swarms working together and brute-forcing solutions, not massive leaps in intelligence from standalone models. That is still an improvement in the usefulness and power of the technology, but it is NOT evidence that model intelligence is increasing faster than before. | | |
| ▲ | famouswaffles 5 hours ago | parent | next [-] | | I don't have any access to any agent swarms (and neither do most) and i still think the models have obviously improved massively in standalone intelligence. Of course they have, agent swarms are not magic. You can swarm all you want around GPT-4 era models and you'll get nowhere. And i've never seen the term 'brute-force' more abused than these LLM discussions. Basically none of the results have been brute force. | | |
| ▲ | semiquaver 5 hours ago | parent [-] | | Agreed. You can’t “brute force” reality, which has an infinitely large state space. A million monkeys won’t write Shakespeare and all that |
| |
| ▲ | CamperBob2 5 hours ago | parent | prev [-] | | "This machine-intelligence stuff is overrated, they are just using <insert particular machine-intelligence technique here>" isn't the resounding verdict it may have sounded like when you typed it. |
| |
| ▲ | mixdup 5 hours ago | parent | prev | next [-] | | Not that they've hit it but that they are approaching it. The time to panic and steer the narrative is before you hit the iceberg, not after | | | |
| ▲ | phoghed 5 hours ago | parent | prev | next [-] | | People have been saying this since GPT-4. | |
| ▲ | ActionHank 5 hours ago | parent | prev [-] | | Have we honestly seen that great a leap in the last 6 months, or just better application of what we had 6 months before that. We are seeing multiple frontier models dropping on the same day and no one bats an eye, because it's more of the same. | | |
| ▲ | CuriouslyC 5 hours ago | parent [-] | | The difference between 6 months ago frontier and now frontier in 3d modelling, graphics and video editing is night and day. | | |
| ▲ | ActionHank 5 hours ago | parent [-] | | Just because there are new capabilities, doesn't mean they've pushed passed the plateau, they've just expanded where the previous solutions work. We've gone from 80% in some places to 80% in some more places. | | |
| ▲ | CuriouslyC 4 hours ago | parent [-] | | > We've gone from 80% in some places to 80% in some more places. Any area that is verifiable will trend inexorably towards 100% over time. In unverifiable areas, it'll always be "80%" because the ubiquity of "AI" style erodes its value, and ">80%" for unverifiable things involves fashion, cachet and "vibes" that humans will probably never knowingly let it have. | | |
|
|
|
|
|
| ▲ | xienze 5 hours ago | parent | prev | next [-] |
| > sudden panic and desire to "slow down" is because they're hitting the plateau on capability I don't think that's the motivation, it's because both companies want to IPO and the _only_ way to even hope to be profitable is to do a whole lot less training, which costs a fortune. But unless Chinese labs go along with this gentleman's agreement (they won't), slowing down on training will bring about the inevitable Chinese model parity date more rapidly. At which point the game is well and truly over for OpenAI and Anthropic. Bit of a pickle they've gotten themselves into with the emphasis on being best, with premium prices to match. |
| |
|
| ▲ | azan_ 5 hours ago | parent | prev [-] |
| > Another piece of evidence on the pile that the sudden panic and desire to "slow down" is because they're hitting the plateau on capability People were talking about plateau for years already. |