| ▲ | minimaxir 5 hours ago |
| Pricing is...a bit weird. Input
$0.10 / MTok for prompts up to 100,000 tokens
$0.50 / MTok for prompts over 100,000 tokens
Output
$0.50 / MTok for prompts up to 100,000 tokens
$2.50 / MTok for prompts over 100,000 tokens
100k tokens is an absurdly low cutoff and it is only applicable to Haiku and not Sonnet or Opus. It's a low enough cutoff that it will be quickly exceeded if you are doing anything with Agents; for typical generation or Jev-like classifiers, it's a good value and as noted in this article, that is apparently the vast majority of Haiku use.In both cases, still much cheaper than Haiku 4.5's $1 input / $5 output and these prices better compete with GPT-6 Luna. ($0.10 input / $0.50 output, but with no token threshold [EDIT: the threshold for Luna is apparently 272k]) |
|
| ▲ | dannyw 5 hours ago | parent | next [-] |
| Haiku 5.5 is noticeably smarter than GPT-6 Luna, so I can see their pricing strategy here. For a while Anthropic has lacked a cost effective “cheap” LLM for summarisation, compacting, RAG helpers, etc. These ‘ephemeral’ workloads are often under 100k tokens, or can be structured to be under 100k. In some coding benchmarks, Haiku 5.5 beats Sonnet 5! (Especially implementation; do a well defined Jira ticket; etc), it’s really impressive how much intelligence per dollar has grown in just a few short months. |
| |
| ▲ | RussianCow an hour ago | parent | next [-] | | I said this in another comment, but Artificial Analysis has the cost per task of Haiku on max roughly equal to that of Sol on medium, and the latter is significantly more intelligent. (And I'd wager that Sol probably finishes tasks more quickly, even with Haiku inference being faster.) So Haiku really only makes sense on lower reasoning levels, and only if you care about intelligence and speed more than you do about cost effectiveness (where Luna currently dominates). And that's without even bringing Chinese models into the mix. | |
| ▲ | WinstonSmith84 4 hours ago | parent | prev | next [-] | | noticeably smarter remains to be seen in practice. For now, Haiku is a bit more expensive than Luna on < 100k token, but I just don't have any agentic work below 100k, so this is going to be 5x more expensive than shown on these charts. It's hardly competitive ... | |
| ▲ | flockonus 2 hours ago | parent | prev | next [-] | | > it’s really impressive how much intelligence per dollar has grown in just a few short months. Open weights models giving a distant salute from afar | | |
| ▲ | pimeys an hour ago | parent [-] | | Yes, it was weird to see MiMo and DeepSeek missing in the article's comparison... | | |
| ▲ | RussianCow an hour ago | parent | next [-] | | It's not that weird. Most companies considering paying Anthropic are probably not considering Chinese models as alternatives. Many don't even realize they exist. | | |
| ▲ | doodlesdev an hour ago | parent | next [-] | | The thing is: availability of near-SOTA cheap Chinese models is forcing OAI and Anthropic to bring prices down and offer more efficient models, instead of simply focusing on super expensive SOTA LLMs. | |
| ▲ | epolanski an hour ago | parent | prev [-] | | "companies" is a meaningless metric. If you want to make it about 99% of real world companies, they are all on Gemini or Copilot anyway, nobody is going through legal and procurement to get models from dubious silicon valley startups when you have relations with Microsoft or Google or Amazon from ages because some benchmark is showing some minor digit benefit when vibe coding GTA 6. | | |
| ▲ | RussianCow 3 minutes ago | parent | next [-] | | I said "most companies considering paying Anthropic", which is not the same as "most companies". I also don't agree that "nobody" is doing this; I have lots of anecdata suggesting otherwise. Maybe the majority of companies are using the easy option of Copilot or Gemini like you said, but it's nowhere near 99%. | |
| ▲ | tranceylc 32 minutes ago | parent | prev [-] | | Vibe coding gta 6 haha |
|
| |
| ▲ | stavros 24 minutes ago | parent | prev [-] | | Is Mimo good? I've never tried it, but I've seen it mentioned three times in this subthread alone. DSv4.1 is my daily driver. |
|
| |
| ▲ | JacobAsmuth 2 hours ago | parent | prev [-] | | The benchmarks are very long form logic, knowledge, and coding tasks though. I'm very interested in Haiku 5.5's performance on ObviousBench where Luna 6 is currently SotA. |
|
|
| ▲ | Tiberium 5 hours ago | parent | prev | next [-] |
| There's also a tokenizer efficiency difference: modern Claude's 100K tokens are about ~60-65K modern GPT tokens, so in reality the Luna cutoff is much further away than the Haiku one. You can test with Anthropic's count_tokens endpoint or with https://crates.io/crates/tokwc |
| |
| ▲ | AtNightWeCode 5 hours ago | parent [-] | | > ...this tokenizer, the same input text produces approximately 30% more tokens on Claude Haiku 5.5 than on Claude Haiku 4.5. So, it is might be even worse. | | |
| ▲ | Tiberium 5 hours ago | parent [-] | | No, it's just Haiku 4.5 is so old that it predates the new Claude tokenizer change in Claude 4.7+ |
|
|
|
| ▲ | jeremyjh an hour ago | parent | prev | next [-] |
| If you can't get any coding done with 100K context that is either a broken model, a broken harness or a skill issue. I would mostly use Haiku in task or explorer subagents. I'm not saying I stay under that on every task, but I do have quite a few sessions that cap out well below that, so that price difference would be very meaningful. I use Luna for this day in and out and its excellent - if Haiku is that much better I will be changing things up. |
| |
| ▲ | serf 41 minutes ago | parent [-] | | >If you can't get any coding done with 100K context that is either a broken model, a broken harness or a skill issue. "less context is better and if you can't get stuff done with less yur bad" is the worst argument ever. it might be pure luxury to your eyes, but it's great to not require the use of a special custom harness that transcribes everything into emoji and compresses everything into barcode images. it's great to have a million token context to throw a large project into. If I need 100k just about any current gen consumer GPU in the world has very good models that I can self host for 100k context, limiting myself to 100k on someone elses machine seems to be missing a lot of the point unless the model itself is extraordinary. |
|
|
| ▲ | Eridrus 5 hours ago | parent | prev | next [-] |
| It's actually existing flat per-token pricing that is weird. Neither encode nor decode are linear in compute, so providers need to price for average expected length. This is just getting closer to the true cost of generating tokens. |
| |
| ▲ | foota 4 hours ago | parent | next [-] | | My theory here is that providers cover the non-constant costs of output tokens as context length caries using the cache input fees. | |
| ▲ | hgoel 2 hours ago | parent | prev | next [-] | | Flat per-token pricing is likely just logistically easier, particularly if these closed models are also picking up the kv cache efficiency improvements seen in recent open weight models. | |
| ▲ | sebzim4500 4 hours ago | parent | prev [-] | | Flat pricing is weird too but jumping up 5x at one cutoff is surprising in the other direction IMO |
|
|
| ▲ | HarHarVeryFunny 5 hours ago | parent | prev | next [-] |
| Notable that one suggested use case for Haiku is "classification requests", i.e. Jev competitor, and the pricing matches GPT-6 Luna which is behind OpenAI's "Decisions API" Jev competitor. For this application 100K token input is plenty. Of course Anthropic and OpenAI, both at $0.10/M, are still 2.5x the cost of Jev's $0.04/M. |
| |
| ▲ | martianvoid 5 hours ago | parent [-] | | I think the 2.5 times cost but actually pays off in terms of intelligence compared to jev and the general capability of using it beyond classification | | |
| ▲ | HarHarVeryFunny 5 hours ago | parent [-] | | The classification performance remains to be seen, but presumably we'll soon start to see classification benchmarks. For other tasks like summaries (another suggested usage) it's good to see Haiku and Luna now competing against each other on cost. I'd love to know how the business automation market breaks down by volume of call type though - hard to imagine that decision making (e.g. branching, triage) isn't a very large part of it, greater than these other suggested Haiku use cases. |
|
|
|
| ▲ | alexchamberlain 4 hours ago | parent | prev | next [-] |
| Isn't it less than a year since Claude models went from 100k token limit to 1M limit? Don't get me wrong - my main agent normally gets to 25% or so before I clear it these days, but as a subagent, doing research or summarisation, I don't think 100k is "absurdly low". |
|
| ▲ | tr4656 5 hours ago | parent | prev | next [-] |
| Luna does as well, but just at a higher limit. From OpenAI's website: Prompts with more than 272K input tokens are priced at 2x input and cache rates and 1.5x output for the full request. |
| |
| ▲ | minimaxir 5 hours ago | parent | next [-] | | Huh, that disclaimer is on the model page (https://developers.openai.com/api/docs/models/gpt-6-luna) but not the pricing page. Annoying. Fixed. | |
| ▲ | tripleee 4 hours ago | parent | prev [-] | | So even at the 1.5x/2x rate luna is still half the price of this. Weird pricing strategy from Anthropic. I'm sticking with Luna if I don't need a super smart model | | |
| ▲ | usef- 3 hours ago | parent [-] | | You're judging purely by token cost I assume, not cost per completed task? The benchmark in the article showed it as lower per completed task than luna, but I guess we'll find out how representative that is. Anthropic has generally been fairly honest in their benchmarking though. | | |
| ▲ | RussianCow 2 hours ago | parent | next [-] | | The cost per task from Artificial Analysis is roughly 3x higher at every reasoning effort level for Haiku than Luna. Sol 6.1 on medium has the same cost per task as Haiku with significantly higher intelligence. According to those numbers (which you should take with a grain of salt), from a pure cost vs intelligence standpoint, you're better off using Luna for economics and Sol for intelligence. With that said, the real reason to use Haiku is that it's faster than all of these models. OpenRouter is showing an average so far of 93 tokens/sec, and AA got at least 137 in each of their benchmarks. So it might be valuable for speed at lower thinking levels. (At higher thinking levels, it's likely going to take longer to produce results than Sol on low/medium.) https://artificialanalysis.ai/models/releases/comparisons/cl... | |
| ▲ | tripleee 3 hours ago | parent | prev [-] | | yes, that's true. I should be looking at the $/completed task |
|
|
|
|
| ▲ | port3000 5 hours ago | parent | prev | next [-] |
| They are targeting businesses/API use for fast decision making and agent integration. Plus they now need to be competitive with Jev-type models in that space. |
| |
| ▲ | cogman10 3 hours ago | parent [-] | | I think they are also trying to make sure Deepseek and other chinese models don't eat their lunch. They need something price competitive. |
|
|
| ▲ | mkotlikov 3 hours ago | parent | prev | next [-] |
| If you look at how different reasoning levels can easily exceed task cost of sonnet 5.5 you will see that you will basically never fall into that under 100,000 token threshold. I mean maybe you can choose low and do a basic summary task, but then you could choose something much cheaper instead. I don't know what Anthropic is thinking with its dumber models. |
|
| ▲ | mnicky 4 hours ago | parent | prev | next [-] |
| You could also use it as a subagent prompted eg by Sonnet/Opus orchestrator agent and for many agentic workflows significant part of the dispatched tasks might be under 100k budget. |
|
| ▲ | giancarlostoro 5 hours ago | parent | prev | next [-] |
| I with they'd give Haiku like 400k tokens roughly, I think between 400k or even 600k tokens is a sweet spot, but Haiku is basically designed to be for small edits is my understanding, but it sucks because any time I ask Opus to "try" letting Haiku do the work, it just falls apart and Opus comes back and tells me it switched to Sonnet (even before Sonnet finally jumped up to 5.x). I will try the new Haiku, but it would be worthwhile if Haiku could take sane instructions and do all file editing for Opus / Sonnet / Fable then it would be worth using. |
|
| ▲ | enraged_camel 5 hours ago | parent | prev | next [-] |
| >> 100k tokens is an absurdly low cutoff and it is only applicable to Haiku and not Sonnet or Opus. It's a low enough cutoff that it will be quickly exceeded if you are doing anything with Agents Your vibes don't appear to be supported by facts. From the announcement: >> Claude Haiku 5.5 is priced 90% lower than Claude Haiku 4.5 for requests up to 100,000 tokens, and 50% lower for requests over 100,000 tokens. On Haiku 4.5, 90% of requests fell into the former category. |
| |
| ▲ | Philpax 5 hours ago | parent | next [-] | | People weren't using Haiku 4.5 for agents before. 5.5 is good enough that it might be. | | | |
| ▲ | StilesCrisis 4 hours ago | parent | prev [-] | | Haiku 4.5 users were using it for Kleenex requests because that was the best it could do. | | |
| ▲ | enraged_camel 3 hours ago | parent [-] | | Not really. We use Haiku 4.5 to turn users' natural language queries and requests into fairly complex structured specs for interior design and construction. It has near perfect accuracy. | | |
| ▲ | dotancohen 2 hours ago | parent [-] | | How many examples are in your prompt? How large is that prompt? Or do you have some other way of tuning the output? I'm asking to learn for a similar project, not to discount anything you're saying. |
|
|
|
|
| ▲ | solenoid0937 an hour ago | parent | prev | next [-] |
| This is pretty good tbh |
|
| ▲ | insanitybit 5 hours ago | parent | prev | next [-] |
| I mostly use Haiku for really, really basic stuff, never for actual engaging work. I've used it for first-pass analysis to triage bugs, for example - all it does is related N bugs together to see if any potentially relate. Then I have Sonnet investigate further. |
|
| ▲ | sixtyj 3 hours ago | parent | prev | next [-] |
| Chatbot could be < 100k tokens. |
|
| ▲ | AustinDev 5 hours ago | parent | prev | next [-] |
| encode and decode tok/s which is ($/s) when it comes to pricing drops heavily above 100k tokens. There are plenty of workflows like translations where you'd easily be under the cap. |
|
| ▲ | system2 5 hours ago | parent | prev | next [-] |
| Who in their right mind would use haiku while Mimo or GLM cost 10% of what they are charging with much smarter models? |
| |
| ▲ | mrngld 4 hours ago | parent | next [-] | | That's not what any benchmarks that look at cost per task or similar says in terms of cost. The Chinese models, generally speaking, might be cheaper per token but need a lot more tokens to get there. | | |
| ▲ | RussianCow 2 hours ago | parent [-] | | Except for the new MiMo V2.6 models, which appear to give some of the best value right now, at least on paper. (I haven't tried them so I can't speak from experience.) |
| |
| ▲ | wyrdcurt 5 hours ago | parent | prev | next [-] | | Some people/organizations are ideologically opposed to using Chinese models. Not me, I use GLM-5.3-Flash for almost everything (the subscription-subsidized pricing on a legacy Z.ai plan makes it the best value model by a wide margin), along with some MiMo and DeepSeek. Still, I use Luna for certain tasks where speed is more valuable than performance; I can see this new Haiku displacing Luna for those. If you mean Haiku 4.5 though I agree, that model was a waste of time and money. | | |
| ▲ | pimeys 2 hours ago | parent | next [-] | | Luna is not really the fastest. You need to use it in high/max to get the good output for what it is good for: summarizing. And that is already close to two minutes per task... | |
| ▲ | girvo 2 hours ago | parent | prev [-] | | I’m on the Legacy v2 plan and same: nothing comes close to 5.3 Flash’s value on it. It’s crazy, no wonder they discontinued them! |
| |
| ▲ | nharada an hour ago | parent | prev | next [-] | | Isn't the point of this release that it's comparable? AAI Index // Input // Output Haiku 5.5: 43 // $0.10 // $0.50 Mimo 2.6 Pro: 46 // $0.43 // $0.87 Mimo 2.6 Flash: 38 // $0.10 // $0.28 Seems competitive to me? Plus then I don't have to manage multiple providers | |
| ▲ | pkulak 3 hours ago | parent | prev | next [-] | | Where do you get this 10% number? Checking providers I know/respect, and GLM 5.3 flash is $0.15/m. Haiku is $0.10/m. | |
| ▲ | user43928 5 hours ago | parent | prev | next [-] | | Presumably everyone who doesn't bother integrating a third party API key into their harness, which would probably be most of the Claude Code users. | |
| ▲ | usef- 3 hours ago | parent | prev | next [-] | | On subscription pricing a $20 Anthropic subscription gives >$500 equivalent tokens, which is not so different, and you get smarter models. API pricing has decent margins. And Opus 5.5 is really good. | |
| ▲ | skeledrew 4 hours ago | parent | prev | next [-] | | Well, unless you're using OpenCode Go, it's per-token costs (even if already super low), while Haiku falls under the Claude sub. It's just more straight forward and you aren't feeling a "loss" with the sub. | |
| ▲ | aesthesia 4 hours ago | parent | prev | next [-] | | There really aren't any models at 10% of the price of Luna or Haiku. | |
| ▲ | ray_kay777 3 hours ago | parent | prev [-] | | People who are stuck using Bedrock in-geo due to their company policy (me). |
|
|
| ▲ | esafak 5 hours ago | parent | prev | next [-] |
| It's their creative way of 'matching' Luna's prices. |
|
| ▲ | j45 5 hours ago | parent | prev [-] |
| It could be to incentivize people to not be lazy users of tokens. |