| ▲ | SyneRyder 6 hours ago |
| You've got to think the comments about "good enough" are people who have not yet tried Opus 5.5. I haven't been this struck by a step change since 4.5/4.6. It's a bigger jump even than when Fable first arrived. As for Mistral - I got really excited when they said Large 4 was focusing on being #1 in cybersecurity, because that's somewhere that they genuinely could edge out Anthropic & OpenAI. Have it actually solve problems, instead of Anthropic flagging "you tried to find a null pointer exception bug in your own code, we're now reporting you to the US government". But on the Mistral benchmarks I'm seeing, this looks very disappointing, but at least they haven't entirely given up. I genuinely thought Mistral had given up on new general models. They need to learn the bitter lesson all over again. |
|
| ▲ | Systemerror7A69 6 hours ago | parent | next [-] |
| I feel like the "good enough" argument isn't about how big the gap between models is but about how good they are at solving the tasks at hand. The capabilities of all models increasing so much all the time means there are simply less and less tasks you need a frontier model for. Even if Opus 5.5 is 500x better than Deepseek, if deepseek can solve all my problems, why do I need to pay for more? |
| |
| ▲ | user43928 6 hours ago | parent | next [-] | | Many on HN still have the opinion that you must understand every line of code in the project, and that all is lost should you merge code that wasn't reviewed. Obviously any model will do if you use it as a better autocomplete. I believe that there is a large gap in expectations between different workflows. Until the AI like reads my mind and produces perfectly production ready apps with minimal intervention from my side, there is still going to be room for improvement. | |
| ▲ | kragen 5 hours ago | parent | prev [-] | | If Opus 5.5 is 500x better than Deepseek, but Deepseek can solve all your problems, maybe you need to work on better problems. If you don't, and you're in business, your competitors will work on the better problems. If you're an employee, your employer might prefer to pay Anthropic instead of you. If you're doing projects you're interested in, you can tackle more ambitious projects with a more capable model. This morning I elicited a microkernel operating system from Opus 5.5. Well, mostly. It doesn't implement task switching yet; we'll see if it runs into a wall at some point. But it boots in QEMU, and it's running a user process in ring 3 and serving web pages. | | |
| ▲ | fultonn 3 hours ago | parent | next [-] | | > maybe you need to work on better problems I have enough real problems in life. I don't need to invent new ones just because a new technology is available. Many of my problems in life are fully solved far past my satiation point by a 3b model that costs me nothing to run. Many others are not. But in either case, when I am acting and living wisely, almost all of my problems exist prior to the existence of technological solutions to those problems. This is also true for the customers and employers that I care to work with. This has changed in me over time, but I now try my best to avoid inventing new problems. The world has enough big, important problems already. > This morning I elicited a microkernel operating system from Opus 5.5. This is cool but also a good example. I don't need a personalized microkernel just because it's possible to have one. Maybe I need one and I don't know it, but the problem statement definitely isn't "I have inherent desire for a personalized microkernel". | |
| ▲ | breuleux 3 hours ago | parent | prev | next [-] | | The best problems to work on are not necessarily the hardest ones, nor the ones that need the most intelligence. They're the problems you, or other people, actually have. Are you going to give up on painting your deck because it's too easy and you don't need a 500x genius to do it? | |
| ▲ | sashank_1509 5 hours ago | parent | prev [-] | | Yeah no, if you elicit Opus 5.5 , anyone else can, and you have no moat either. But if on the other hand, I mostly use my human intelligence and just need a dumb model to complement my human intelligence at low cost and high speed (say review every commit to catch obvious bugs), I have a much better chance of building an actual moat than you do. But outside of coding, it’s even more clear that you don’t need frontier intelligence. My customer service agent is very happy with a 100B param Deepseek flash model, thank you! | | |
| ▲ | TuxSH 3 hours ago | parent [-] | | > say review every commit to catch obvious bugs I’m using subscription models for exactly that, better models catch more subtle bugs, and they catch them faster. It works out far better in terms of work-hours saved. Also A/ then OAI slashed token pricing by 2x~5x on their latest models |
|
|
|
|
| ▲ | eigenspace 6 hours ago | parent | prev | next [-] |
| I use Opus 5.5 daily for my job. I am aware (and in awe of) it's capabilities. Look at the context in which I used that term 'good enough'. What i was saying is that there are tasks for which a dumber model can be good enough, and for organizations with sovereignty/ privacy concerns, those concerns can be strong enough to incentivize the use of a dumber model. |
|
| ▲ | skerit 6 hours ago | parent | prev [-] |
| > I haven't been this struck by a step change since 4.5/4.6. It's a bigger jump even than when Fable first arrived. I had the exact same experience. And unlike Fable, it doesn't gobble up your entire usage limit in a few hours. I always wonder what the "good enough" people are actually using it for. |
| |
| ▲ | sashank_1509 5 hours ago | parent | next [-] | | “Good enough” as in we don’t see any point in 1-shotting everything we want to build in lightning speed. If Opus 5.5 can 1 shot it then that product is essentially commodified, no point in any one building it except as an internal tool. If Opus can’t 1-shot it, then it must rely on our human intelligence which can be complemented well enough with a dumb model as a frontier model. | | |
| ▲ | nonadhocproblem 3 hours ago | parent [-] | | I assume that you'd also argue in favour of working in a team with an average IQ of 80 as opposed to 120. |
| |
| ▲ | blahblaher 3 hours ago | parent | prev [-] | | The thing is, for how long? Are they going to keep giving you "so much intelligence" for a "small" subscription dollar amount? When they really, for real, need to start making money to cover their costs, what do you think it's going to happen? Suddenly you will start having tasks that a "good enough" model is going to be fine. |
|