| ▲ | dannyw 6 hours ago |
| Haiku 5.5 is noticeably smarter than GPT-6 Luna, so I can see their pricing strategy here. For a while Anthropic has lacked a cost effective “cheap” LLM for summarisation, compacting, RAG helpers, etc. These ‘ephemeral’ workloads are often under 100k tokens, or can be structured to be under 100k. In some coding benchmarks, Haiku 5.5 beats Sonnet 5! (Especially implementation; do a well defined Jira ticket; etc), it’s really impressive how much intelligence per dollar has grown in just a few short months. |
|
| ▲ | RussianCow 2 hours ago | parent | next [-] |
| I said this in another comment, but Artificial Analysis has the cost per task of Haiku on max roughly equal to that of Sol on medium, and the latter is significantly more intelligent. (And I'd wager that Sol probably finishes tasks more quickly, even with Haiku inference being faster.) So Haiku really only makes sense on lower reasoning levels, and only if you care about intelligence and speed more than you do about cost effectiveness (where Luna currently dominates). And that's without even bringing Chinese models into the mix. |
|
| ▲ | WinstonSmith84 5 hours ago | parent | prev | next [-] |
| noticeably smarter remains to be seen in practice. For now, Haiku is a bit more expensive than Luna on < 100k token, but I just don't have any agentic work below 100k, so this is going to be 5x more expensive than shown on these charts. It's hardly competitive ... |
| |
| ▲ | goosejuice 20 minutes ago | parent [-] | | No? Don't use these lower end models to work on small well defined tasks with a frontier model orchestrating? This approach works very well for me and don't have any issue staying under 100k. I have no idea if it's cheaper but it does seem to be much faster for tasks like QA. |
|
|
| ▲ | flockonus 3 hours ago | parent | prev | next [-] |
| > it’s really impressive how much intelligence per dollar has grown in just a few short months. Open weights models giving a distant salute from afar |
| |
| ▲ | pimeys 2 hours ago | parent [-] | | Yes, it was weird to see MiMo and DeepSeek missing in the article's comparison... | | |
| ▲ | stavros an hour ago | parent | next [-] | | Is Mimo good? I've never tried it, but I've seen it mentioned three times in this subthread alone. DSv4.1 is my daily driver. | |
| ▲ | RussianCow 2 hours ago | parent | prev [-] | | It's not that weird. Most companies considering paying Anthropic are probably not considering Chinese models as alternatives. Many don't even realize they exist. | | |
| ▲ | doodlesdev 2 hours ago | parent | next [-] | | The thing is: availability of near-SOTA cheap Chinese models is forcing OAI and Anthropic to bring prices down and offer more efficient models, instead of simply focusing on super expensive SOTA LLMs. | | |
| ▲ | RussianCow an hour ago | parent [-] | | Is it? I would guess that it's much more about the race to get customers as they both near IPO than anything to do with the Chinese models. | | |
| ▲ | verdverm 42 minutes ago | parent [-] | | we are actively preparing to move our devs from closed to open models, take it as a piece of anecdata the trend in industry is clear by now though |
|
| |
| ▲ | epolanski 2 hours ago | parent | prev [-] | | "companies" is a meaningless metric. If you want to make it about 99% of real world companies, they are all on Gemini or Copilot anyway, nobody is going through legal and procurement to get models from dubious silicon valley startups when you have relations with Microsoft or Google or Amazon from ages because some benchmark is showing some minor digit benefit when vibe coding GTA 6. | | |
| ▲ | RussianCow an hour ago | parent | next [-] | | I said "most companies considering paying Anthropic", which is not the same as "most companies". I also don't agree that "nobody" is doing this; I have lots of anecdata suggesting otherwise. Maybe the majority of companies are using the easy option of Copilot or Gemini like you said, but it's nowhere near 99%. | |
| ▲ | tranceylc an hour ago | parent | prev [-] | | Vibe coding gta 6 haha |
|
|
|
|
|
| ▲ | JacobAsmuth 3 hours ago | parent | prev [-] |
| The benchmarks are very long form logic, knowledge, and coding tasks though. I'm very interested in Haiku 5.5's performance on ObviousBench where Luna 6 is currently SotA. |