| ▲ | ainch 2 hours ago | |
That's a very fair critique. I don't mean to imply that Kimi is not at all cheaper than U.S frontier models. I more wrote this because I believe - since Chinese LLMs entered the public consciousness via DeepSeek R1, which was genuinely ~20x cheaper than o1 - there's a bit of a halo effect around Chinese models which causes people to overestimate the scale of the discount. And relative to that price anchor, Kimi is less extraordinarily cheap. At the moment Kimi is ~10% cheaper than GPT-5.6 on the AA benchmark, and as you say that could go down to 20-30% cheaper (although I don't know how inference provider discounts play out on real world usage once you account for quantisation etc...). I'm not trying to suggest that that's nothing, but I do think some of the people driving the Chinese AI discourse would have a harder time pitching their conclusions if they were saying "this new Chinese model is 10% cheaper on some tasks, and it might get another 20% cheaper in the future". | ||
| ▲ | sroerick 32 minutes ago | parent | next [-] | |
Sir, it's 3X cheaper on coding. 3X cheaper in any industry is earth shattering. 10% is significant. 3X is really big. | ||
| ▲ | coder543 an hour ago | parent | prev | next [-] | |
But if you compare to Anthropic's models? The cost difference is huge. Anthropic is clearly concerned that people are realizing they are expensive, since the Opus 5 blog post dedicated a lot of time to talking about how cheap the model was compared to the competition... but this doesn't hold water when I haven't seen any independent benchmarks claiming Opus 5 is cheaper than GPT-5.6-Sol, even if it is supposedly closer. GPT-5.6-Sol is pretty competitively priced, but not all American frontier models are, and even 10% to 30% is still significant for any commodity that's as fungible as frontier models often are. > as you say that could go down to 20-30% cheaper I never said anything about 20% to 30%. We don't know how much it actually costs to host this model yet, and that will determine the final price. It could be just a little less, or it could be a lot less. > once you account for quantisation There will be no need to account for quantization. Kimi models have been 4-bit only since at least K2.5. They don't release or serve models in higher precision than that. This isn't one of those situations where LLM inference providers are debating between serving 16-bit, 8-bit, or 4-bit, and I have never seen a publicly hosted, paid model that was hosted in less than 4-bit, even if hobbyists will use sub-4-bit quantizations sometimes locally. | ||
| ▲ | an hour ago | parent | prev [-] | |
| [deleted] | ||