| ▲ | ranger_danger 2 hours ago |
| I figured they're just admitting AI models have plateaued and are coming up with some fake story about self restraint so they don't lose VC money |
|
| ▲ | lukan 2 hours ago | parent | next [-] |
| Not sure about the level of irony here, but I keep hearing models have plateaued since a while now, but I keep being impressed with the latest model performance. |
| |
| ▲ | 0xcde4c3db an hour ago | parent | next [-] | | I don't think "plateaued" is the right word, but I do feel like there's been something like a logistic curve compression in the difference between smaller and larger models as the field evolves. For inference at least, the scale of practical difference between a single high-VRAM GPU or SFF UMA box, a whole rack, and a whole data center seems to be falling far short of what we might have imagined just a few years ago. The conversations I've heard have largely turned away from breathless anticipation of the next frontier model and toward attempts at hard-nosed evaluation of which tokens are worth the cost. | | |
| ▲ | antupis 21 minutes ago | parent | next [-] | | I think it’s more that pushing frontier is extremely costly and there is no free lunches in same way as 2024. | |
| ▲ | eru an hour ago | parent | prev [-] | | Maybe, though that's another kind of progress in itself. Very impressive progress! |
| |
| ▲ | unleashhale 2 hours ago | parent | prev | next [-] | | Anything in particular? My experience has been like seeing the addition of retractable cupholders, but maybe different domains. | | |
| ▲ | qskousen an hour ago | parent [-] | | I have a pet project I have been working away on for some time that involves building GPU backends for various cards in Zig, lots of complex stuff in it. Lately I mostly use Opus 5, it can pretty reliably plug away at things but it does mess stuff up occasionally. For this codebase, Fable 5.1 was noticeably better at getting things right and doing things in a good reliable way. Of course, I can only use Fable for a bit before I hit the usage cap for the week, so I save it for the tougher things. That said, I absolutely abhor the way recent Anthropic models write prose, especially comments. I recently tried doing a fairly normal task for this codebase with codex, as I have seen a lot of people talking it up on here. A single task running for ~1-2 hours burned through over half of my usage for the week on the $125/month plan, not on a top model (I don't remember which one specifically I used). It struggled to get the basics done, then got absolutely stuck on a follow up. Handed it over to Claude and it 1-shot it. | | |
| ▲ | eru an hour ago | parent [-] | | I really liked codex in the last few weeks, especially its ability to clean up after Claude's (prose) messes and do reviews. But in the last few days something seems to have happened that made Codex's models massively stupider (for what I am doing). Really weirdly, it suddenly refused to even run tests it previously wrote itself (and previously ran), because of some false positive about cybersecurity. That by itself is not evidence of stupidity. Trying to make a 200+ file PR full of research notes is, and the PR didn't even solve the problem I asked it to. |
|
| |
| ▲ | f4dd 29 minutes ago | parent | prev | next [-] | | They've not plataued but they're certainly not as impressive as the hype would have them to be. The reality is, it doesnt matter if LLMs keep getting more powerful because they still need a human to steer it. Without the human providing inputs to the LLM it just sits there and does nothing. | | |
| ▲ | arcanemachiner 21 minutes ago | parent [-] | | You don't need human input. Any coherent input will do the trick. You can, for example, hook it up to a logging system and have it fix errors as they occur on your platform. |
| |
| ▲ | whateveracct an hour ago | parent | prev | next [-] | | astra is more parlor tricks than real gains tbh i swear they trained in on threejs in particular so those idiots on twitter could spam their garbage demos | |
| ▲ | dismalaf 22 minutes ago | parent | prev | next [-] | | Impressed with the model performance or the chatbot/agent performance? | |
| ▲ | drTobiasFunke 2 hours ago | parent | prev [-] | | Really? My employer rolled back to opus 4.8 because 5 was expensive AND crap. Didnt even consider fable because it didn’t add any additional value. For most software eng and design work opus 4.6-4.8 just works fine. For everyday joe asking ai to plan a trip or home diy work even sonnet works fine. Any cybersecurity or other areas are niches that cannot support trillion $ valuations. What am I missing? Genuinely curious | | |
| ▲ | lukan an hour ago | parent | next [-] | | No idea what you are missing and yes, Opus is quite solid, but Fable is clearly way better for me. I just did a direct comparison, big change in a quite complex codebase. Same prompt for Opus, same for Fable. Fable clearly won and delivered very good results, while Opus delivered mediocre, so I did not let it finish. I expected both to fail and was prepared to do lots of manual steering, but not necessary with Fable one shotting it, and all this with 35$ of credits for fable. I am still impressed. If I would have had to hire a human, it would have cost me thousands of dollar for the same task - and a way longer time. So maybe the valuations are overblown, but they clearly provide value for me. | |
| ▲ | bpodgursky 2 hours ago | parent | prev [-] | | If Fable doesn't add additional value in your workplace, it means you aren't being ambitious enough in how you integrate agents into your workstream. Yes, it's probably comparable to 4.8 if you are just using it to write code and put up a couple pull requests. That's not where things are now. | | |
| ▲ | nradov an hour ago | parent | next [-] | | You shouldn't be down voted, AI native companies have already moved up to the next level beyond writing individual PRs. | |
| ▲ | bix6 an hour ago | parent | prev | next [-] | | And where are things now? | |
| ▲ | nozzlegear 41 minutes ago | parent | prev [-] | | This is just "you're holding it wrong" with a little smooch of condescension. If only we plebeians could comprehend what magnificent works those who have ambitiously integrated agents into the workstream have wrought! | | |
| ▲ | plorkyeran 2 minutes ago | parent | next [-] | | Sometimes you actually are holding it wrong. It's pretty reasonable to think that Fable isn't worth the massive increase in cost, but if you think it outright doesn't have any benefits over Opus 4.8 then your workflow is probably not making good use of the tools. | |
| ▲ | bpodgursky 27 minutes ago | parent | prev [-] | | I mean it's very easy to comprehend and you don't need me for it. Just download claude code or codex and ask it to give suggestions about where to integrate agents into your workstream. | | |
|
|
|
|
|
| ▲ | nullbio an hour ago | parent | prev [-] |
| The models have not plateaued, and they are not even mildly close to any sort of ceiling. Right now the barrier is data and compute. Quality data can be created synthetically at an exponential rate as models improve. Humans are actively feeding them with private IP. Compute advancements will begin to skyrocket as we unlock photonic computing and materials science advancements and scale up chip fabs. This is also compounding because the AI is accelerating the pace of research, testing, development, manufacturing, etc. It's a big self-accelerating feedback loop. There is no plateau. |
| |
| ▲ | konmok an hour ago | parent | next [-] | | > Quality data can be created synthetically at an exponential rate as models improve No it can't? Every time the labs try this we see model collapse, e.g. shoving goblins into every conversation. And I have seen zero evidence that AI is accelerating materials science in any meaningful way, let alone photonic computing. | | |
| ▲ | dorolow 9 minutes ago | parent | next [-] | | We use large amounts of synthetic data for training at work and have not observed any sort of model collapse when done properly. Edit:
https://arxiv.org/abs/2404.01413
https://arxiv.org/abs/2406.07515 | |
| ▲ | nullbio 23 minutes ago | parent | prev | next [-] | | > Every time the labs try this we see model collapse The latest studies demonstrate model collapse is not a given and synthetic data can be used just fine. The latest models are proof of that, they're all trained on large swathes of synthetic data. It can't be used as the -only- data source of course, but that's not how it is being used. This is an obvious conclusion, too, because there's no difference between synthetic data and the data people can create, the difference is whether that data is revealing new information about the thing the model is trying to learn. If the synthetic data is just teaching the model the same thing over and over again it results in overfitting, so it needs to be done intelligently. For example, if I have an example of a puzzle, I can generalize that example and create thousands of synthetic data examples, with different rotations/perspectives, rather than having to find the data naturally. It's not that the models are just generating data out of thin air, they're generating the synthetic data on top of real world data. The smarter the models get, the better they are at generating quality synthetic variations and finding valid synthetic variations. > And I have seen zero evidence that AI is accelerating materials science in any meaningful way, let alone photonic computing. It is accelerating how quickly researchers and engineers can do their jobs. https://news.mit.edu/2026/ai-helps-design-new-materials-that... This is only the beginning, too... Look ahead a year or two. | | |
| ▲ | konmok 7 minutes ago | parent [-] | | > The latest studies demonstrate model collapse is not a given Which studies? [edit: I'll assume you mean these two given by @dorolow: https://arxiv.org/abs/2404.01413 https://arxiv.org/abs/2406.07515] > It can't be used as the -only- data source of course, but that's not how it is being used Right, so human data creation would also have to scale up exponentially, and that's not gonna happen. > because there's no difference between synthetic data and the data people can create I mean, that's obviously false, otherwise model collapse wouldn't exist. The difference is statistical, but it's there. > It is accelerating how quickly researchers and engineers can do their jobs.
> https://news.mit.edu/2026/ai-helps-design-new-materials-that... That's pretty clearly a hype article, the headline even says "The CrysVCD tool developed at MIT COULD cut the huge amounts of time and money spent". I'm asking for empirical measurements of timelines, not hypotheticals. > This is only the beginning, too... Look ahead a year or two. Lol that excuse is getting really old |
| |
| ▲ | f4dd 33 minutes ago | parent | prev [-] | | There's a lot of deluland posts about. | | |
| ▲ | nullbio 8 minutes ago | parent [-] | | You're not well informed. Helps to keep an open mind if you want to keep up to date. |
|
| |
| ▲ | ranger_danger 27 minutes ago | parent | prev [-] | | Can you provide sources for these claims? | | |
| ▲ | nullbio 9 minutes ago | parent [-] | | What claim do you have a problem with? There are plenty of research papers on synthetic data that show its value, do a search on arxiv for "synthetic data". There are plenty of open-source post-training pipelines that incorporate synthetic data. As for the claim about accelerating the progress of hardware or materials science, I've seen quite a number of news articles from teams at universities using AI in their work with high quality outcomes, and they're becoming more frequent. https://openai.com/index/jalapeno-first-results/ > We used AI to design the chip, and designed the chip so AI could program it
AI played a direct role in Jalapeño’s development, enabling the team to move from initial design to tapeout in nine months by exploring implementations, shortening design, measurement, and verification loops, and continuously iterating on model workloads. AI also helped optimize the chip’s arithmetic circuits, allowing the team to fit more compute performance into the chip on schedule. https://www.anl.gov/article/scientists-deploy-ai-agents-to-a... > An AI-driven system automates a powerful simulation method used to discover new materials. The system can potentially reduce discovery time from months or years to just days. |
|
|