Remix.run Logo
energy123 19 hours ago

None of my use cases require frontier capabilities but I still pay $200/month to a frontier lab. I value the additional time saved at more than $200/month. If I had to pay actual API rates, then I'm not sure what I would do, but it would not be an easy decision.

m_ke 19 hours ago | parent | next [-]

Sure, but anthropic is charging businesses based on usage now and tried hard to pull Fable from the consumer subscriptions before Sol and K3 dropped.

Even now on the $200 plan I use up my Fable credits in a single day and had to start using codex and openrouter for more usage because Fable burns $100s an hour when billed on usage.

surgical_fire 18 hours ago | parent | prev | next [-]

I thought similarly until I decided to try DeepSeek.

It became an easy decision, even the $200/month by Anthropic sounds like a bad deal.

cmrdporcupine 19 hours ago | parent | prev [-]

Yes, the reckoning here will happen in a year or two when the (probably subsidized, maybe?) coding plans become either unavailable or much more costly.

It's also already the case that in larger companies people do not have access to these plans as employees and must use API rates.

There's also the political angle. When Anthropic and OpenAI held back their top of the line models because of Bessent & Trump's bullshit, and threatened to deny unwashed foreigners like me access... I dropped my Codex plan and made do purely with GLM 5.2 for three weeks before OpenAI finally released 5.6 Sol. Feels inevitable that this will happen again.

Or, somebody will come up with a way to serve e.g. Kimi K3 or the new Qwen model in an extremely cheap way. Or DeepSeek releases a competitive model at their cut-throat rates. And then the cost argument just wins.

m_ke 18 hours ago | parent [-]

k3 costs will go down at least 3x within a week of the weights dropping.

we'll get new quants, dspark speculators, distills and optimized kernels

as long as there are near frontier models available there will be inference providers selling them at or below cost of inference in attempt to get market share.

cmrdporcupine 17 hours ago | parent [-]

I have not seen that kind of significant drop with GLM 5.2 yet? so curious why you think it will happen for K3.

This is a very large model. Much larger (3x) than GLM. The resources to run it are very expensive.

tokai 13 hours ago | parent [-]

There's been a price war going on openrouter between providers of GLM 5.2. NovitaAI, DeepInfra, and StreamLake keeps underbidding each other in waves. Yesterday evening both input and output $/M was ~$0.3. Output was especially cheap.

cmrdporcupine 13 minutes ago | parent [-]

I had the opposite experience. I bought the $100 monthly sub from Neuralwatt last month because it was the only economical provider for GLM 5.2. They raised their rates halfway through, and it simultaneously became too slow to use.

I just looked at DeepInfra -- I've got an account there already etc -- and it's at FP4 quant. How much that effects the quality of inference for GLM, I can't say. I could see using it as a backup when other things run out but don't think I'd trust it.