Remix.run Logo
▲ vishvananda 2 hours ago

The reason people aren’t freaking out is because most people are using heavily subsidized subscriptions.

I tried the cheapest provider on openrouter and burned through $50 in a few days. Quality was ok, seems slightly above Luna quality perhaps? But that $50 is 1/4 of my codex subscription where I could have burned that many tokens or more using Astra within my weekly reset.

This won’t last forever but as long as the frontier labs are subsidizing this heavily the open models won’t matter.

▲pimeys an hour ago | parent | next [-]

Yes. You have to find the provider with pricing that suits your usage.

I am having 98% my input in cache, so using Coralbricks makes sense due to them giving cache reads for free — you only pay for writes. I spend maybe 5-10 dollars a day and my agents basically work day and night implementing things for me.

If your tasks are write-heavy, find a provider with cheaper output.

If you build a customer-facing app, pay a bit extra for 400+ tok/s e.g. on Lithos.

▲sheeshkebab 13 minutes ago | parent [-]

What are your agents implementing day and night…sheesh. Built anything useful for anyone yet?

▲georgel an hour ago | parent | prev | next [-]

I am curious how you managed to spend that much on Deepseek via OpenRouter. I loaded $100 back in July while using v4-flash or whatever the cheap good model was at the time, and have upgraded as the new ones came out from Deepseek. I still have $16 and some of that spend also goes towards the AI usage from my customers (the context they need to load in is quite large too).

And I am using the Claude Code harness with DS as the endpoint. And I use it ~5-8hrs a day to do my coding.

▲whstl 28 minutes ago | parent | next [-]

It's wild how different usage patterns are between users.

I have seen the Cursor leaderboard on my company and the vibe coders consume about 5x more tokens than the developers. They and other office workers also have Claude and their limits are often over around Wednesday.

People are using millions of tokens to do very simple HTML reports. I have seen someone asking the LLM to download the entire data into the context and asking it to sort.

Those usage patterns don't correlate to output.

▲lifeisloving 42 minutes ago | parent | prev [-]

Likely the user doesn't know what they're doing or has extermely bad workflows. They're prob not managing their cache, and dont use compaction.. Letting context get to 500k and invalidating their cache every 10 tool calls because they have no providor fallback settings.

I was running deepseek v4.1 pretty much non stop during work hours, with heavy tool/mcp usage and finding it very difficult to spend more than $75 in a month.

Also the cheapest providers on Openroutrr can often have terrible cache hit %, short TTLs resulting in their effective price being much more expensive than people realize. 75% cache pretty much destroys any savings from a super cheap token perspective.

▲lionkor an hour ago | parent | prev | next [-]

How's the caching? I have 99.5% cache hit rate with deepseek when using their own API, it's dirt cheap.

▲eikenberry 37 minutes ago | parent | prev | next [-]

But aren't you developing bad habits and learning patterns that won't work long term? Or do you think things will get cheap enough that you will be able to keep going with your current patterns post-subsidies?

▲irjustin 34 minutes ago | parent | next [-]

> But aren't you developing bad habits and learning patterns that won't work long term?

2 reasons - there's an advantage now, use it. 2nd the frontier providers, this is the "early cheap days" like when uber was initially cheap to compete vs standard cabs. they want you to become hooked and boy are we hooked.

▲denkmoon 25 minutes ago | parent | prev | next [-]

I use the frontier openai/anthropic models at work but exclusively open weight models (on cloud/hosted inference) for personal stuff and I think about it like this; 1) I don't see any reason GLM and DeepSeek won't eventually be as good as Claude, it's just a matter of time and 2) the open models are well and truly capable enough for most of what I'd want to do. I don't need nor want an LLM chewing away on a horrible enterprise spaghetti codebase, my employers can pay for that privilege.

▲FromTheFirstIn 35 minutes ago | parent | prev | next [-]

No one involved in this is thinking about the long term

▲pmontra 22 minutes ago | parent | prev | next [-]

Long term, we will see what happens and adapt. At worst we all go back coding by hand. Meanwhile what can I do, tell my customers that I'm raising my fee because I have to pay for token? The Claude Pro $20 plan is good enough for me and even in auto mode I never had to wait for the 5 hours reset.

▲jaggederest 31 minutes ago | parent | prev [-]

I expect by that point we'll have local models that can do a decent job, I would guess give it a decade and we'll be running custom accelerators that are smarter than current frontier models.

In the same way that only supercomputers used to have multiple processors and caches but it's now standard.

▲onlyrealcuzzo an hour ago | parent | prev | next [-]

OpenRouter is complete garbage.

Buy directly from DeepSeek's API.

You can literally get overcharged 100x on DeepSeek on OpenRouter (or more).

▲jorvi 3 minutes ago | parent | next [-]

Or.. BYOK Deepseek because OpenRouter's UX is much nicer?

▲patwolf 8 minutes ago | parent | prev | next [-]

One of the reasons I use OpenRouter is because they offer zero data retention. As far as I can tell, DeepSeek's own API doesn't support ZDR.

▲georgel an hour ago | parent | prev | next [-]

I'm all in for saving money and _can_ move to using DS directly from them, but maybe I am missing something here:

OpenRouter Pricing:

$0.02/M input tokens $0.60/M output tokens

DeepSeek Pricing (cache miss, off-peak):

$0.15/M Input $0.60/m output

▲girvo an hour ago | parent | next [-]

When 98.5% of my requests are cache hits (according to Pi for the last week), the cache miss price isn’t that important to me, and $0.003-0.006 per 1M input tokens is shockingly cheap.

It’s also the major difference between using DeepSeek directly vs other providers also serving it, though I have not looked lately: it’s possible other providers have matched its cache hit pricing better?

▲georgel an hour ago | parent [-]

Interesting, if the cache hit is that good, I think HN convinced me to toss $20 at DS official, and see how long that lasts.

▲mswphd an hour ago | parent | prev | next [-]

I've heard that certain inference providers may have different quality of caching implementations, so even if the listed numbers are as you say, the practical cache hit % you get might be significantly different/incur significantly different costs.

▲ckdot 17 minutes ago | parent | prev [-]

There’s a big difference in speed & quality between using DeepSeek API directly with DSH vs. DeepSeek in Opencode Go with Opencode CLI. Can’t tell if it’s the provider or the harness - but worth to give it a try.

▲BeetleB 9 minutes ago | parent | prev [-]

DeepSeek trains on your inputs. That's why people go on OpenRouter and choose ZDR providers.

▲seunosewa an hour ago | parent | prev | next [-]

Which provider was that?

▲0xbadcafebee 32 minutes ago | parent | prev | next [-]

You should basically never pay API prices, they are always several times higher than subscriptions.

There are several open weight subscription providers. OpenCode Go used to be good but now it's complete shit. Charm Hyper is really great and the best value. Other subscriptions have a more limited model selection or provide less value but are still decent.

▲mensetmanusman 2 hours ago | parent | prev [-]

Also DeepSeek usage is subsidized as well, it’s a power hungry model.

▲jchw 2 hours ago | parent | next [-]

Interesting. Are all of the providers on OpenRouter simply losing money? How does that even work out?

▲arjie 32 minutes ago | parent | next [-]

No, it’s outrageously profitable above x% without stealing any prompts. Provider economics still pretty good. Acquiring hardware is the current limiter.

▲worldsavior an hour ago | parent | prev | next [-]

You're the RLHF.

▲edflsafoiewq 44 minutes ago | parent | next [-]

Not if you're not giving feedback.

▲jchw an hour ago | parent | prev [-]

With ZDR-only enabled? That seems illegal.

▲ 43 minutes ago | parent [-]
[deleted]
▲jansan an hour ago | parent | prev [-]

Don't ask and dance as long as the music keeps playing.

▲kennywinker 2 hours ago | parent | prev [-]

Are you sure about that? My impression was most providers on openrouter were purely selling tokens for profit...

▲thesnarkitecht an hour ago | parent [-]

Have y'all tried an Ollama Cloud subscription? Their off-hours pricing for V4.1 Flash is extremely competitive.