|
| ▲ | dd8601fn 4 hours ago | parent | next [-] |
| At half the price and less likely to auto-downgrade, it sounds like a reasonable claim. |
| |
| ▲ | benjiro29 an hour ago | parent | next [-] | | > At half the price and less likely to auto-downgrade, it sounds like a reasonable claim Two benchmarks (artificial analysis and vals) show a increase in cost (a insane increase for vals compared to Opus 4.8). Already posted this before, so here is the link. https://news.ycombinator.com/item?id=49041158 | |
| ▲ | binsquare 4 hours ago | parent | prev [-] | | given that i couldn't even use fable without it downgrading to Opus, this is just a straight upgrade for me | | |
| ▲ | Freedumbs 2 hours ago | parent [-] | | Opus 5 also downgrades. it's now Fable -> Opus 5 ; Opus 5 -> Opus 4.8.
Unclear why they want to nerf their own products with sometimes right classifiers. I guess the government ban might've been real and not coordinated marketing? |
|
|
|
| ▲ | ceejayoz 4 hours ago | parent | prev | next [-] |
| Best can describe multiple things. Almost as good for half the cost is something I'm very comfortable describing that way. |
| |
| ▲ | lelanthran 3 hours ago | parent | next [-] | | > Almost as good for half the cost is something I'm very comfortable describing that way. It's also not unusual in this context - many people describe the Chinese models as "best", because it's 80% as good for 20% of the price (or similar). | | |
| ▲ | moffkalast an hour ago | parent [-] | | Hopefully it's not like old Opus, where it was actually more expensive than Fable cause it thought for half an hour, got it wrong, and then thought until you ran out of credits trying to come up with a correction, while Fable just went for it and did it in one go, getting it right the first time without thinking more than a few seconds. Got an endless list of stuff done with Fable, Opus 4.8 was like a flailing braindead idiot in comparison. Maybe this one is a bit better if it's distilled. |
| |
| ▲ | ProofHouse 4 hours ago | parent | prev [-] | | Best marketing |
|
|
| ▲ | tshaddox 4 hours ago | parent | prev | next [-] |
| The blog posts figure cites Frontier-Bench for its agentic coding score, and shows Opus 5 beating Fable 5 43.3% to 33.7%. |
|
| ▲ | ActivePattern 4 hours ago | parent | prev | next [-] |
| I think you're being overly cynical here. First, I don't see any claim that is the world's best model for agentic coding. Second, it is absolutely the best model in terms of coding performance vs. dollar, and it's raw performance seems very close to the frontier. |
| |
| ▲ | adam_arthur 4 hours ago | parent | next [-] | | GPT 5.6 is far more token efficient at most tasks with similar performance. Especially so for Opus 4.8, still to be seen with Opus 5. Where are you getting cheaper per dollar? | | | |
| ▲ | shwaj 4 hours ago | parent | prev [-] | | It would still be the best model per dollar if the score was 2% lower instead of 0.1% lower. Would it be ok to still give it the highlight color then? How big of a lie is too big? Especially when no lie needed to be told at all: many including myself would have noticed the tiny 0.1% deficit and been suitably impressed by the Opus 5 result. I’ll admit this is a small deception by today’s standards. I’m one of those who believes in truth for truth’s sake. Edit: typo | | |
| ▲ | ai-x 4 hours ago | parent [-] | | we don't know if it is 0.1% deficit, could be 0.05% | | |
|
|
|
| ▲ | manojlds 4 hours ago | parent | prev | next [-] |
| Which numbers are you seeing? It does show that it's better than Fable 5 in most things related to coding? |
|
| ▲ | dbbk 3 hours ago | parent | prev | next [-] |
| Yeah I spotted this immediately too. I'm sorry. You're supposed to be a multi billion dollar company and you can't even highlight your chart honestly? |
|
| ▲ | jsLavaGoat 4 hours ago | parent | prev | next [-] |
| In my opinion, the frontier is passed what is really needed for coding. Fable is good as a supervisor. |
|
| ▲ | 4 hours ago | parent | prev | next [-] |
| [deleted] |
|
| ▲ | toephu2 3 hours ago | parent | prev | next [-] |
| Also it scored worse on DeepSWE than chatgpt 5.6 sol |
|
| ▲ | Aurornis 4 hours ago | parent | prev | next [-] |
| Using the most expensive model for all of your agentic coding work hasn’t been good practice for a long time. Not unless you have infinite money to spend. Fable is typically used for key planning, architecting, and review tasks. I think this is a case where you don’t understand the use case, not that the marketing department is making mistakes. |
| |
| ▲ | airstrike 4 hours ago | parent | next [-] | | They cost the same if you're already at $200/mo | | |
| ▲ | Aurornis 3 hours ago | parent | next [-] | | Fable consumes your usage at a higher rate. If you bought the $200/mo plan and you don’t use it much, using Fable for everything is fine. | |
| ▲ | maineldc 3 hours ago | parent | prev [-] | | I am not a tokenmaxxer per se but I blow through my weekly quota on my max plan in 3-4 days… fable would make that worse. |
| |
| ▲ | akmarinov 4 hours ago | parent | prev [-] | | Eh, not really. Fable does a lot better on coding than Opus 4.8. Just this past week Fable was able to figure out a couple of small issues for me where Opus was failing to. Also both are still somewhat bad at UI implementation. Opus more so |
|
|
| ▲ | unclebucknasty 3 hours ago | parent | prev | next [-] |
| Recent releases have said something to the effect (paraphrasing here): "Use <less expensive or older model> for everyday tasks and <other non-critical stuff>. Use <more expensive or recent model> for complex coding tasks, refactoring large code bases, etc.". Then, the next model/release emerges and the previous "best for complex" gets demoted to "everyday". Obviously, it's all relative. But, it does beg the question: was the previous model really good for complex coding tasks or no? I mean, how is it now suddenly only good for the "easy" stuff? |
| |
| ▲ | hvb2 3 hours ago | parent [-] | | > I mean, how is it now suddenly only good for the "easy" stuff? Because your expectations have changed. | | |
| ▲ | unclebucknasty an hour ago | parent [-] | | I'm sure the marketeers would love for the public's assessment of complex versus easy to conveniently shift per their release cycles; or for the public to simply forget their prior marketing. |
|
|
|
| ▲ | entropicdrifter 4 hours ago | parent | prev [-] |
| I mean that certainly makes it best-in-class |