Remix.run Logo
gruez 14 hours ago

>What work was attempted? https://cursor.com/evals shows the previous generation of open models (Kimi K2.7) trading blows with the others, cost effectively

Because you're comparing retail price whereas the parent commenter (and the article) is talking about marginal (ie. inference) costs. American labs are providing a premium product and they're charging accordingly. Meanwhile for chinese models they're open weight so they're limited to how much they can charge without competitors undercutting them.

If we use tokens as a rough proxy of inference costs (rough approximation, I know) and look at artifical analysis benchmarks, you see that all the open models are behind the pareto frontier in terms of efficiency.

striking 14 hours ago | parent [-]

I'm arguing we can't trust retail prices because the marginal pricing isn't meaningfully connected to it anyway.

But if we have to look at what we think margins might look like, DeepSeek continues to host v4 Flash at the existing price despite competitors beating it in price (https://openrouter.ai/deepseek/deepseek-v4-flash), so there's at least one example of a Chinese lab charging a predetermined price despite competition. And no one but Moonshot is hosting Kimi K3 yet (https://openrouter.ai/moonshotai/kimi-k3). Perhaps there's room in the market for those who release their models to make margin on them.

And I believe my Composer example speaks for itself. The open models are behind but there's tangible proof they can be tuned for pareto frontier efficiency. See "Cost per Task" at https://artificialanalysis.ai/agents/coding-agents.

gruez 13 hours ago | parent | next [-]

>DeepSeek continues to host v4 Flash at the existing price despite competitors beating it in price (https://openrouter.ai/deepseek/deepseek-v4-flash),

their competitors are discounted at around 33%, so it's safe to say that's the margin, maybe less if their competitors have worse caching or quantization. Meanwhile claude code/codex resellers selling tokens for 90% off API price, presumably by reselling usage from fixed consumption plans, which gives an idea on how fat the american labs' margins are.

>And I believe my Composer example speaks for itself. The open models are behind but there's tangible proof they can be tuned for pareto frontier efficiency. See "Cost per Task" at https://artificialanalysis.ai/agents/coding-agents.

But composer is a closed model? If it's really that easy to get better coding performance, why haven't the chinese labs replicated it? And this is all assuming the performance boost is real and not from benchmaxxing. Moreover if you apply the "street price" discount I mentioned above, American labs look far more favorable.

striking 13 hours ago | parent | next [-]

The fixed consumption plans are offering several times their worth compared to API pricing with completely free cache reads: https://she-llac.com/claude-limits

I look at that and think that they must be losing money hand over fist on something like this, not that this shows what their margins are like. If their margins are like this then I don't see why they'd be raising money and shuffling it around in circles.

> If it's really that easy to get better coding performance, why haven't the chinese labs replicated it?

Nobody said it would be easy! I just think it's possible, and that presumably they will get around to doing it at some point.

ycui7 12 hours ago | parent | prev | next [-]

discounted competitor could cheat. they can offer subpar model response and sell it as deepseek-v4. it is uneconomical to prove inference providers are cheating, so they get away with it. cheating inference provider does not care if their customers stay.

c0brac0bra 13 hours ago | parent | prev | next [-]

The 33% discounted competitors have no non-retention policy

byzantinegene 8 hours ago | parent | prev [-]

the reselling of fixed price plans by resellers are causing american labs to lose alot of money and that's why they're trying really hard to stamp it down.

metalspot 6 hours ago | parent | prev [-]

Deepseek is still charging cached input at 1/10th the price of any competitor.

For an example of my real token usage for a day with DS: Input (Cache hit) 530,949,760, Input (Cache miss) 7,875,004, Output 1,389,685 - it is still 1/5th the price of Baidu (the cheapest) and 1/7th-1/10th the price of US hosts.

Also, Deepseek platform is not the same thing as Deepseek open weights. There is a major misconception that the existence of an open weights model means that it is the same thing as the proprietary platform offering, but that is definitely not the case.