Remix.run Logo
▲ proxysna 4 hours ago

I am yet to spend $200 on deepseek this year. Not sure what kind of usage can justify $200/month of either openai or anthropic, i'm not even talking about $500. Deepseek is faster, IMO intelligence difference is negligible and it so much cheaper that i no longer care about how much i use it. I never hit any daily/weekly quota or anything like that while working or tinkering. At this point i am OK with being 6 months behind the "frontier", purely on bang-for-buck basis and who cares which shadowy government gets my data.

▲rapind 3 hours ago | parent | next [-]

So I took Deepseek V4.1 Flash for a spin maybe 2 weeks ago now (before Luna 6 and Sol 6 were announced), and I racked up $100+ in about 2-3 days. It was pretty great, but it uses way more tokens (TPS is fast, but it's way more tokens per turn) than Sol 5.6 which I found to be about it's equivalent at the time (on medium or high, with DS on max). My cache rate was around 98-99%.

It would definitely cost me more per month than a x20 ChatGPT or Claude plan, probably around $400+ was my estimate at the time. This was with Fireworks (ZDR) which has since increased their prices (and got slower!).

That being said, very impressed with the model, and looking forward to what comes next. As the frontier models become less subsidized, the open models will become more appealing.

P.S. There are subscription plans for open models, but I've found most of them to be extremely slow, have model throttling (only so much of model X), and also very sketchy about training and data retention. No thanks! If you want to share your data, just use Muse Spark contributor. Seems impossible to beat that on price per task if you don't mind feeding your data to the Meta machine (spoiler: I won't).

▲jbellis 3 hours ago | parent | next [-]

There's a bunch of skepticism in the replies but I ran over 100 tasks against DeepSeek 4.1 Flash and Sol (among others) and I can confirm, it is in fact a little smarter than Sol and a little more expensive than Luna. https://slopcop.com/power-ranking?pricing=api

I also spent $280 on DeepSeek doing the tests (direct to DS, not OpenRouter). I suggest that if you can't conceive of anyone spending $200 on DeepSeek, you're not being ambitious enough!

▲cpursley 2 hours ago | parent [-]

Is slopcop your domain, because that is awesome. Wishing you great success with it.

▲jbellis an hour ago | parent [-]

It is. TY!

▲taylorfinley 3 hours ago | parent | prev | next [-]

Were you using OpenRouter? I've used 1.8bn tokens in the past week from DeepSeek themselves and 99.2% were cache hits. Total cost was $18.13 usd.

▲faitswulff 3 hours ago | parent | next [-]

For readers wondering, OpenRouter isn’t capable of caching as effectively as DeepSeek is because they will, for instance, switch inference providers in the middle of a session.

▲Computer0 3 hours ago | parent [-]

The way I use openrouter is I find a model/provider combination I like then pin all requests for that model to that single provider.

▲swingboy 2 hours ago | parent | next [-]

If you disable all other providers but DeepSeek in your OpenRouter guardrails, is that effectively the same thing?

▲roarkeful 2 hours ago | parent | prev [-]

How do you do this?

▲pests 2 hours ago | parent | next [-]

provider: { order: ['deepinfra/turbo'], allowFallbacks: false, },

https://openrouter.ai/docs/guides/routing/provider-selection

▲pwython 2 hours ago | parent | prev [-]

What pests said. And you can make a preset and pass "model": "@preset/deep-seek"

▲rapind an hour ago | parent | prev [-]

Fireworks directly. At the time they were the best value of cost, speed, ZDR. They got slower on me though, but I think they are retooling, so maybe things have or will get better again. I think fireworks is primarily for when you want to do your own training on top, which I wasn't doing.

▲ampdot 22 minutes ago | parent [-]

DeepSeek, DigitalOcean, GMICloud, and NovitaAI were the only OpenRouter providers that didn't lead to major performance degradation for me

▲christophilus 3 hours ago | parent | prev | next [-]

This is my experience, too. It's a great model, but it burns tokens if you use it heavily for work on complex domains.

Edit: others have noted the provider and harness matters. My experience is with opencode.

▲InsideOutSanta 2 hours ago | parent | prev | next [-]

Same experience. I often see people say how little they spend on DeepSeek v4.1 flash, but when I put 60 bucks into my account, it was gone in a few days of non-exclusive use. I'm actually curious what the difference is. I used it through pi and opencode, but the harness seemed to have no obvious impact on usage.

▲esafak 34 minutes ago | parent [-]

Maybe they just use it less. If you code all day you can go through a billion tokens.

▲pimeys 3 hours ago | parent | prev | next [-]

How on earth you can do 100 dollars in 2-3 days with DeepSeek? I have 7 agents in omp running 24/7 every day. I use maybe 10-15 dollars a day. A rarely see a session going over 2 dollars. My maximum is maybe 3.5 dollars and that session took three days.

What harness you are using?

▲rapind 2 hours ago | parent | next [-]

pi. I wasn't even going that hard. I checked the logs for Sep 18 and I did just shy of 3b input with approx 98.5% cache and 5.8m output, which cost around $35. Most of the was a Rust code review exercise with 1 driving agent and a varying number of subagents (up to 6 some times). I do the same with with Sol med/high driving and Luna x-high reviewing and get at least as much done if not more in a day, but I'd use up two x20 weekly allowances for the week. Worth noting that token cost isn't super meaningfull on it's own, because DS is super token heavy (but also great at caching) compared to Sol. (my stats show DS uses 3x the tokens as Sol)

The shape of my work changes obviously, so it'll vary, sometimes more, sometimes less. For example, fixing all of the bugs and defects I found that week was 2-3 times the effort and chewed through my ChatGPT allowance, but I had banked resets...

Also worth noting that codex models have been kind of all over the place recently with their usage... and it looks like costs are changing again.

▲ricardobeat an hour ago | parent [-]

You might be overusing subagents. Especially with a chatty model like DS, you’ll be wasting millions of tokens on re-discovering the project and facts instead of actual reasoning.

▲rapind 16 minutes ago | parent [-]

I agree, and I started sending well defined review packages to Luna x-high instead, which is how I have this setup when using Codex (Sol drives, Luna async reviews, and Sol keeps moving). Same process with DS really dropped my DS usage by a lot and I think that might actually be the secret sauce. Especially if you use Luna through a subscription (probably the entry pro level would be enough). I'm not sure what a Luna equivalent open model is though. 5.6 Luna max was catching a lot of issues while my implementer keeps rolling. Last couple of days I had 3-4 Sol Mediums running with Luna x-high reviewers (async reviews) and one Sol medium orchestrator and Astra X-high to plan everything out.

Gotta be honest though, I don't love fiddling with this all the time. I would rather be working on my projects than evaluating my usage. Having DS in my back pocket should i need it is a relief. The providers get fiddly though too.

▲the__alchemist 2 hours ago | parent | prev [-]

What are those agents doing? I am out of the loop. Bitcoin mining? Blogging? Reddit bots?

▲drewnick 2 hours ago | parent | next [-]

I have deepseek agents doing email responses, with real tools (think running quotes, gathering info, scheduling things) and running business processes that used to be done by $35/hr administrative type people. And the capabilities are expanding every day as I learn how to build scaffolding around the model.

▲esperent an hour ago | parent | next [-]

> I have deepseek agents doing email responses

So, spam?

▲abustamam 30 minutes ago | parent [-]

Not all communications done by a non-human is necessarily spam.

▲ducktoysleftout an hour ago | parent | prev [-]

I am a tech lead for a useful but non-essential PaaS my company subscribes to. Their product manager occasionally sends me clearly AI-authored emails. I ignore them. I am a believer in the usefulness of AI, but nothing says I don’t care more clearly than sending me slop.

I can’t overstate how bad of an idea I think using an AI for customer interaction is.

▲abustamam 28 minutes ago | parent [-]

I think it depends on the industry. In my company (insurtech) a large percentage of our sales come from AI engagements (phone and text). The customers know they're talking to an AI and they can elect to talk to a human at any time but many times they don't.

But yeah, send me a non solicited AI slop email or worse, political ad, and you dont get the dignity of me saying stop to unsubscribe. Straight to spam for you.

▲pimeys an hour ago | parent | prev | next [-]

Work for my company. Research, code, analysis.

▲Mistletoe 2 hours ago | parent | prev [-]

I’m lost when I read these sort of comment chains. Free Gemini works just fine for me. Maybe it’s because I don’t use it for programming? How many programmers really exist out there? Surely it can’t support the weight of investment that exists in AI already. It’s just such a small pool of the human race.

▲internetter 2 hours ago | parent | next [-]

First, they come for the programmers, and next the mathematicians. Then it will be the biologists, lawyers and doctors. Humanities will stake it out a little bit longer because AI isn’t human, but AI companies would be dammed if they don’t try. Eventually, with advancements in robotics, stabs at increasingly more physical sciences will also be attempted. Eventually, AI will have its hand in the pie of all knowledge work, if it is possible. Not to mention all the roles like tech support and customer service. Once they have gotten as far as they think they can go, they will try to turn up the prices. However, they might struggle to do so as models are becoming a commodity. This is why they are arguing for regulation and stating that only they can tame these beasts.

▲abustamam 26 minutes ago | parent [-]

I think as long as we have open weighted models there will always be competition. Sure Anthropic could raise their prices but it wont take long for someone else to undercut them. The quality may not be as good but people would probably prefer to pay 1% of the cost for 75% of the quality.

▲ac29 42 minutes ago | parent | prev [-]

> How many programmers really exist out there? Surely it can’t support the weight of investment that exists in AI already. It’s just such a small pool of the human race

I dont consider myself a programmer but use LLMs almost exclusively for coding.

The number of people able to create useful software today is much much larger than it used to be and arguably a minority of these people are/were "programmers"

▲MisterMunchkin 2 hours ago | parent | prev | next [-]

How have you spent hundreds of dollars? I’ve only spent 11 and I’ve been using it for four months!

▲guluarte 3 hours ago | parent | prev [-]

same, used a wrapper around cc and i was spending up to $30 a day with basic stuff

▲taylorfinley 3 hours ago | parent [-]

Maybe CC does something that breaks the cache? I cannot recommend Oh My Pi enough. Every default is galaxy brained, and it plays incredibly well with deepseek flash 4.1. My favorite coding harness rn for sure.

▲rapind 2 hours ago | parent [-]

I found omp used quite a bit more tokens than my fairly basic pi setup... but most of those tokens would be cached with DS V4.1, so maybe worth if there are gains elsewhere.

▲sillysaurusx 4 hours ago | parent | prev | next [-]

It’s easy to hit your quota. “Speed up the compilation time of this C++ codebase. Feel free to use several subagents to search through the files in parallel.” That’ll cost you about $200 for a codebase of ~1,000 files.

Subagents are like trading derivatives. You can lose as much as you want.

▲abixb 4 hours ago | parent | next [-]

What bothers me about this whole AI tokenomics situation is the lack of transparency. OpenAI and Anthropic have to perhaps be the most opaque companies in existence wrt their offerings. There's like a thousand variables that they can change on the backend at the push of a button which can wildly swing API spends within the same model (partly also due to the non-deterministic nature of LxMs, but still), and there's no objective way to measure them other than vibes.

When the regulations do arrive, I think they should really focus on AI companies and API providers being more transparent wrt how they're billing their customers. Because right now, it's a totally vibes-dependent and a mess.

▲MintsJohn 3 hours ago | parent | next [-]

And it's all measured in "intelligence", a completely meaningless term. For coding i'd be much more interested in how much context actually works, what the complexity of algorithms it can understand and create is, for what languages. How much it manages to follow existing structures or that is just adds ad-hoc machinery to pass the test, etc etc.

A smaller model in the same generation will never be the same as a bigger one, assuming this is a smaller model, and the same generation, as naming implies, it will not be comparable, it might be on the benchmarks, even on the benchmarks that matter, but the whole story should also give the drawbacks.

▲KeplerBoy 3 hours ago | parent | prev | next [-]

It's still insane that they stopped showing you all the tokens you pay for. They could inflate the billed reasoning token amount by a lot before it would raise any eyebrows.

▲ldng 2 hours ago | parent | prev | next [-]

Internet Ad business has been like that for a long time with bot click & Co.

▲mlmonkey 2 hours ago | parent [-]

Ain't that the truth.

In Search Advertising, the amount you pay (under GSP Auction) is a function of your pCTR. And guess who determines your pCTR? The Search Engine itself! :-D

▲kruipen 2 hours ago | parent | prev | next [-]

Totally hand-wavy and non-objective ... just like how employees are billing their employers.

▲zer00eyz 4 hours ago | parent | prev | next [-]

It's all Gacha for business.

▲catigula 4 hours ago | parent | prev [-]

Yeah, you might get away with a singular 'wrt' with some consternation, but two?

▲gbacon 3 hours ago | parent | prev | next [-]

> Subagents are like trading derivatives. You can lose as much as you want.

Excellent pithy warning.

▲sheepscreek 2 hours ago | parent | prev | next [-]

Sadly this is true - for individual folks on the lower end of the spend spectrum.

But there’s a point on that spectrum where the ability to run multiple experiments in parallel, even with a significant amount of (one time) wastage, is overall more cost effective than the alternative.

▲proxysna 4 hours ago | parent | prev | next [-]

Afaik there is just pay-as-use with Deepseek

▲deadbabe 4 hours ago | parent | prev [-]

Why use subagents at all

▲FearNotDaniel 4 hours ago | parent | next [-]

Preserve context in the lead chat - let the subagents fill up their own contexts then only return the necessary information.

▲deadbabe an hour ago | parent [-]

Why not just have an agent that can branch its context?

▲gf000 42 minutes ago | parent | next [-]

Forking conversations have been a thing for a long time, and it's not the same thing (e.g. you may have 50% of your context used up at fork time that is carried into "both agents" afterwards)

▲ 42 minutes ago | parent | prev [-]
[deleted]
▲nater5000 4 hours ago | parent | prev | next [-]

Why hire a junior developer if you have a perfectly competent senior developer already on your team?

▲giancarlostoro 3 hours ago | parent | prev | next [-]

You can do a code review on a "less capable" model that costs less, and the key model gets its output / summary, then you can have that model build a plan, and feed it to cheaper models. It's a more efficient approach than just running everything through Opus, and now that Sonnet is a lot better I'll probably use them more frequently, one thing to note is don't ask it to spin up endless subagents, I'd cap it to 2 or 3 at a time, otherwise, yeah you'll hit your limit extremely quickly.

▲josephg 23 minutes ago | parent [-]

Code reviews also work better in sub agents because the reviewer agent didn’t write the code being reviewed.

▲qarl 4 hours ago | parent | prev | next [-]

Because two agents are faster than one.

▲HDThoreaun 2 hours ago | parent | prev | next [-]

The models get dumb as context fills. Subagents allow them to accomplish a task with minimal context rot. You can also use cheaper models for subagent tasks

▲phyalow 4 hours ago | parent | prev [-]

Time is money. Parallelism is very helpful optimising one to get the other.

▲knollimar 4 hours ago | parent | next [-]

Money is money too. Increasing contexts costs non zero money, even with cache hits. Also context rot is a problem that subagents help with

▲apsurd 4 hours ago | parent | prev | next [-]

This truism is intuitive to everyone but always funny to me how everyone never has any time, needs to save time, needs to hire staff workers for every mundane job and robots can't come soon enough… all so we can binge watch Game of Thrones and 90 day Fiancé.

And watch 10 hours of football on Sunday for our DraftKings bets.

▲apitman 4 hours ago | parent | prev | next [-]

Money is also money, which anyone who makes heavy use of parallel subagents will quickly learn.

▲georgemcbay 2 hours ago | parent | prev [-]

> Time is money. Parallelism is very helpful optimising one to get the other.

Parallelism is fantastic when it actually speeds up the entire pipeline, but in my experience most people's jobs (at least the ones for which AI is currently relevant) involve a lot of overlapping "hurry up and wait" branches that drastically blunt the real benefits of that sort of parallelism.

There may be specific situations where it makes sense to do it, but just immediately going full gastown on anything AI related seems like such a giant waste to me, of both money and finite world resources.

▲jrflo 4 hours ago | parent | prev | next [-]

There's a difference between "write this function for me" coding agents and "build this prototype from end-to-end". If you're doing the former, deepseek is fine. If you're doing the latter, it's not gonna work, and that's where the extra intelligence is most valuable.

▲polytely an hour ago | parent | next [-]

I'm mostly using deepseek 4.1 flash via openrouter, the way I'm doing stuff is:

1. write a sketch of a spec by hand

2. have the llm review the document and question me until it can generate a spec

3. review the spec and revise where needed

4. have it write an implementation plan

5. another round or revision/review

6. executing the plan step by step through the plan, plausing between each step to see if we are still on course and if the decisions it made track with my understanding of what we are doing.

I've been working for a couple of hours tonight, the total cost of the session is €0.6.

it's not the build this thing end to end, but also not quite write function x for me. It is still a lot of manual review, but I find I really need it to even discover what I actually want to build. I just cannot imagine building something in a single shot and getting something that actually has value (unless it is basically a clone of an existing thing). To me the whole value of ai right now is that it's now very cheap to build custom software that exactly matches your preferences.

▲josephg 16 minutes ago | parent | next [-]

One of the big differences using the better models is that you don’t need to hold their hand anywhere near as much. Fable is crazy expensive, but I’ve seen it just one-shot some remarkably complex projects. The code it produces is much better, too.

▲polytely 7 minutes ago | parent [-]

yeah i just think that if I gave the initial pitch to Fable it would have successfully built something that wasn't exactly what I wanted. maybe it is because i went to school for design, but how something works is so important to me and it is hard to get to that point without actually thinking about it step by step and visualizing how interactions would work. I just cant imagine a single one-shot prompt containing enough information to for the model to know what to build. I guess if you have Frontier model money like these silicon valley freaks you could just iterate by asking it to tweak what you didn't like until it is right.

▲huflungdung 18 minutes ago | parent | prev [-]

[dead]

▲throooooo 3 hours ago | parent | prev | next [-]

This was my experience 3 months ago. I had an Android app that interacted with a Bluetooth device that I wanted to reverse engineer and build my own Linux app for it. DeepSeek was struggling really hard. Claude did it end to end after 3 or 4 prompts. To be fair, I was using a web interface for DeepSeek and the CLI for Claude; maybe that makes a large difference.

▲sreekanth850 4 hours ago | parent | prev | next [-]

build this prototype from end-to-end, is this how people build serious software with AI?

▲jrflo 3 hours ago | parent | next [-]

You're not gonna get something that's ready to ship, but as a first pass to get something running yes. Let's you explore far more ideas with only a few hours of agent time.

▲OrangeDelonge 4 minutes ago | parent [-]

And you can’t use deepseek with a loop approach for creating such PoCs?

▲thangalin 2 hours ago | parent | prev | next [-]

I started KeenLore (an emotive audiobook creator) that way. I gave it software specifications, languages, JSON schema definitions, container requirements, hardware configuration (8GB NVIDIA T1000 GPU, 96GB RAM), and zero user interface mockups. For the second round, I asked it to build a completely independent, re-entrant, and data isolated demo system on top of the web application. The demo application included voice generation using one of its voice designs. Here's the output:

https://www.youtube.com/watch?v=WAeHgE94rVo

The system performs quotation attribution on my local hardware for my near-future, hard sci-fi novel (having nearly 500 quotations) with over 97% accuracy.

The initial prototype was developed quite quickly, but numerous successive iterations were required to fix numerous gaffs by Opus 5 (because it doesn't actually _understand_ what it takes to make general-purpose audiobook narration software).

▲mitthrowaway2 2 hours ago | parent | prev | next [-]

I don't think most prototypes are serious software.

▲poilcn 3 hours ago | parent | prev | next [-]

This is how ChatGPT, Cursor apps are built. They spent so much money on PR stunts, but "thousands of agents" can't make an app that doesn't freeze on each keystroke. Not even talking about user-friendly ui

▲marknutter 3 hours ago | parent [-]

Weird, ChatGPT has always worked really well for me.

▲sunaurus 2 minutes ago | parent | next [-]

Such issues are pretty common among people I talk to at least. I hear comments about it pretty often, both when it comes to ChatGPT and Claude.

For myself, with ChatGPT for example, it regularly gets extremely slow if there is a lot of text in a single conversation. Especially if I try to scroll up.

Some pages have been strangely just broken for a while now as well. Usage analytics just renders lots of these duplicate "Usage history" components, where the data just never loads: https://imgur.com/a/vpcQIiw.png

I can totally imagine a scenario where some agent built it, another tested and approved it, and nobody at OpenAI even looked at it once or knows that it's like this.

▲SOLAR_FIELDS 3 hours ago | parent | prev [-]

Early on their software was like this, the original ChatGPT desktop app was basically unusable and full of memory leaks that would tank the software. They’ve long since fixed that though

▲jurgenburgen 3 hours ago | parent [-]

The new ChatGPT desktop app is a dumpster fire though, it can’t even scroll properly when streaming the answer.

▲outside1234 3 hours ago | parent | prev [-]

For a prototype or a 1 or 2 use tool, yes, this is exactly how serious people are building software.

▲piterrro 3 hours ago | parent | prev | next [-]

Oh buddy you have no idea what a good plan and agent harness can do with deepseek…

▲proxysna 3 hours ago | parent | prev | next [-]

I am doing mostly hardware drivers recently, it works fine for complex work.

▲logicchains 3 hours ago | parent | prev [-]

"build this prototype from end-to-end" works fine with DeekSeek V4.1 Flash, the problem occurs if you're not only building a prototype but want a finished product.

▲pnw 3 hours ago | parent | prev | next [-]

I spent the weekend trying Deepseek 4 Pro on a Linux porting project and it led me down a complete rabbit hole where Linux wouldn't even boot by the end of the weekend. Waste of $120. Switched back to GPT 6 on Monday and Linux is booting again and I'm making progress.

The only thing I've found Deepseek and Kimi good for are security tasks that GPT refuses to do.

This is a summary of what Deepseek did and got wrong:

Lost the proven baseline: changed kernel source, configuration, compiler, RAM geometry, MMC width, and peripherals together. Matching an upstream commit did not preserve local boot fixes, making failures difficult to isolate. Misidentified an image: a file labelled “r18-known-good” actually contained the r23 parent bootloader. Filename-based reasoning replaced verification of the artifact’s identity and provenance. Shipped inconsistent boot contracts: flash-16b’s loader read too few kernel blocks. Fresh2 changed the device tree without updating the loader’s expected length and CRC, creating deterministic rejection before normal Linux handoff. Patched binaries without maintaining reproducible source: loader constants diverged from source, a separately compiled cache-flush length remained stale, and assembly used an oversized stage-two slot. Their causal contribution to hangs was not established. Overstated diagnosis: claimed failures were definitively in U-Boot, blamed compiler or IPU changes without controlled isolation, converted noisy observations into confirmed hangs, and neglected persistent journals as an alternative explanation. Mistook compilation for integration: framebuffer registration was incomplete, timing success handling was inverted, BT.656 selection was unreachable, encoder overrides were missing, and audio lacked software clock configuration. Misread hardware evidence: asserted interrupt-free PMIC operation, assigned RF to the wrong SPI controller, confused regulator identifiers with register addresses, and described repeated encoder writes as unique registers. Overclaimed results: treated kernel/probe indications as userspace success, presented earlier discoveries as new progress, and omitted failed flashing attempts from the final narrative.

▲forsalebypwner 2 hours ago | parent | next [-]

> Deepseek 4 Pro

There's your problem, 4.1 Flash is significantly better and cheaper, to the point where the official DeepSeek API is going to (or already has, I forget) redirect requests for Pro to 4.1 Flash, and adjust billing accordingly too.

4 Pro is still offered by providers I'm sure, since it's open weight, so I can understand making that mistake.

▲rapind an hour ago | parent | prev | next [-]

That's an unfortunate experience. Think of v4.1 flash as actually v5.0 flash. It's night and day compared to the 4.0 flash (and 4.0 flash was unintuitively better than 4.0 pro). I would re-evaluate with v4.1 flash. I'm not saying it better than Sol or anything, but it's in the ballpark.

▲pimeys 3 hours ago | parent | prev [-]

You mean 4.1 Flash which is the first great Deepseek? The one that actually surpasses Opus in my books now.

▲apitman 4 hours ago | parent | prev | next [-]

I've been trying to use DeepSeek V4.1 Flash more and been very impressed. My current (very rough) rule of thumb is that an Artificial Analysis score of ~40 is the crossover point for "good enough" for most of the things I need to do with coding agents.

▲mchusma 4 hours ago | parent | next [-]

47 is my crossover for serious things (e.g. Grok 4.7 is below the line and GPT 6 Sol is above the line). I mean, Opus 5.5 is way better, but GPT 6 Sol still gets the job done for anything that doesn't require design thinking.

Although I do think Luna 6 max is ok for some basic things, would never use it for coding myself.

▲apitman 4 hours ago | parent [-]

What's the most important task you would/wouldn't trust with an agent below the line?

▲hmontazeri 4 hours ago | parent | prev [-]

Had the same experience. I got rid of my pro sub of OpenAI. I’m really freaking impressed.

▲dzink 3 hours ago | parent [-]

Where are you inferencing DS4.1 flash reliably ?

▲apitman 3 hours ago | parent [-]

I use OpenCode Go. I've also used the DeepSeek provider through OpenRouter a bit in the past and it seemed solid.

▲nullbyte 4 hours ago | parent | prev | next [-]

The intelligence difference between models like DS4.1 and Sol/Opus is NOT negligible.

▲bel8 3 hours ago | parent | next [-]

The premium price is only worth fo the hardest problems.

For CRUD shoveling, models like DS4.1 are enough.

And the intelligence gap between cheap and premium is closing, as can be seen from the title of this post.

▲linuxftw 3 hours ago | parent [-]

The issue is once you solve the hard problems, the lower models start messing things up that were working and reverting all fixes for the hard problems. They'll just go off and do dumb stuff.

▲zozbot234 2 hours ago | parent | prev | next [-]

In the Artificial Analysis index, MiMo 2.6 Pro is smarter than GPT-Sol 6.1 Low at the same cost, and only slightly dumber than Medium. MiMo 2.6 Flash is marginally cheaper and smarter than GPT-Luna 6 Max. (There is no GPT 6+ Terra, which would otherwise be in that range.) These are not negligible or trivial results.

▲tripleee 4 hours ago | parent | prev [-]

If both DS4.1 and Opus can complete the tasks you throw at it at good enough quality the differences are negligible.

Who cares if your car can go 200mph if all you need is 60. If my requirement is 60mph, I want a faster 0-60, not a higher top speed.

▲ctolsen an hour ago | parent | next [-]

Opus 5.5: $4/$20

Deepseek 4.1: $0.02/$0.60

Just to illustrate how cheap the Corolla is in your analogy. Also Opus output would be $50 without competition.

▲_benj 3 hours ago | parent | prev | next [-]

Specially if using the 200mph car when you need it is just a /model away.

▲nkjoep 3 hours ago | parent | prev | next [-]

Or just a cheaper way to reach 0-60

▲fragmede 3 hours ago | parent | prev [-]

A car that feels safe to be driving at 200 mph is going to feel more comfortable at 60 mph, compared to one for which 60 mph is at the very limits of its abilities. Analogies only go so far so I'm not sure there's anything to be learned from that though.

▲rmaxdev 4 hours ago | parent | prev | next [-]

What do you do? I’m 200 bucks deepseek flash in about 2 months and it’s increasing

I use it as main Hermes model that orchestrates codex/droid harnesses with subscriptions for heavy dev work

I do have ChatGPT as main assistant that sets direction and delegation of projects to Hermes

At my increasing usage, kind of 200 usd subscriptions makes sense and max out on Luna max

▲iammrpayments 4 hours ago | parent | next [-]

I put 5 dollars at deepseek a long time ago, and somehow it has never been fully spent, can’t imagine anyway to spend 200$ on that thing

▲proxysna 4 hours ago | parent | prev [-]

Recently, drivers for a bunch of obscure hardware. Lots of c\c++, that i am ok with but not enough to make hardware drivers (i am just impatient). Just using pi agent with a few plugins.

▲marknutter 3 hours ago | parent [-]

How complex are they? Drivers vs an entire application would be a big difference in token usage, and it could also depend on the type of work being done.

▲proxysna an hour ago | parent [-]

Complex enough where performance matters. I've done entire applications, client, server, infra and ci as an experiment with deepseek v4 pro a few months ago. Works for that too.

▲holbrad 2 hours ago | parent | prev | next [-]

I think the only answer to this is you're just not using agents enough, because even with the very cheap pricing, it's still easy to rack up a large bill.

▲thiht 2 hours ago | parent | prev | next [-]

I've been using Claude Code at work and OpenCode for side projects for a few months. Every OpenCode model I've tried always felt subpar compared to Claude, but good enough. But it changed with DeepSeek 4.1 Flash, I've been using it for the past few days and I've come to forget I was not using Claude, it's a really good model and it's basically free for my usage (I used it almost all the weekend and spent ~$5)

▲LarsDu88 3 hours ago | parent | prev | next [-]

I've spent $200+ on deepseek and this is for making a multiplayer FPS game. Trust me there are use-cases.

And no it did not deliver. A lot of it was re-done by Astra

▲abroszka33 an hour ago | parent [-]

> I've spent $200+ on deepseek and this is for making a multiplayer FPS game.

Why do you expect that $200 will give you that on ANY model? Multiplayer FPS games are very difficult to make, no AI will deliver that today.

▲LarsDu88 42 minutes ago | parent [-]

Deepseek v4 pro was able to create procedural dungeons, placeholder 3d assets, 12 weapons, and menus. Struggled with ragdoll effects, and gun modeling in blender.

Astra was able to model low poly enemies, rig them, do simple animations, and greatly improve procedural generation. I have it modeling assets in blender every day, which I often have to go in and fix

▲abroszka33 36 minutes ago | parent [-]

Nice, that's all single player stuff.

▲ an hour ago | parent | prev | next [-]
[deleted]
▲jwpapi 2 hours ago | parent | prev | next [-]

A lot of people having different pricing experience. I think it’s important to understand that caching can differ, than if the agents spend waiting on code, or consume a lot of content. It depends on how you structure you codebase and how explorable it is, how much effort you set and probably some other issues.

For raw productivity most of what works is best and switching will cost you getting on use parity with other models, as you need to learn what they good at, potentially how the tool works and how to prompt it best.

For tasks that you implement in code, you should have benchmarks and evals.

That said for me was Luna a huge leap and 500+ of cost savings a month

▲zzleeper 4 hours ago | parent | prev | next [-]

A bit tired of spending $200 out-of-pocket for openai. What do you use as harness? (for me the harness if half of the benefit... controlling my PC, working from phone, etc.)

▲simlevesque 4 hours ago | parent | next [-]

I use Claude Code + eternal terminal + tailscale + tmux + some custom skills to get notifications through nfty.sh.

I get the same UX on every platform, works perfectly on very low bandwith environments such as in a cabin, in the subway or in the middle of nowhere.

I tried using other harness such as Pi and opencode but I did not like them. If Claude Code gets weird I can swap in an instant.

You just need to follow this guide and disable artifacts in Claude Code's config: https://api-docs.deepseek.com/quick_start/agent_integrations...

▲gf000 31 minutes ago | parent | next [-]

What makes it run well on low bandwidth environments? Is it eternal terminal? (I am not familiar with that).

I found that ssh is pretty bad in these scenarios, so my usual herdr over a phone doesn't always work properly.

Now I'm using paseo which in principle solves my issues properly, but unfortunately it's "reconnection" and state sync is pretty slow (probably going over their servers).

▲IOT_Apprentice 4 hours ago | parent | prev [-]

I’m curious about your use of Tailscale, is that for you to reach a local LLM remotely from anywhere?

▲simlevesque 4 hours ago | parent [-]

I have a big beefy desktop at home which I ssh (using EternalTerminal instead of raw port 22) into. I built it last summer right before the prices got very expensive. It's headless so I use a cheap Macbook Air to connect to it at home and I use Termux on my phone to continue working from everywhere.

▲cruffle_duffle 4 hours ago | parent [-]

Dude prompt your agent to set up always on remote connections via systemd… then you can drive from Claude or codex mobile apps natively over their native hookup. Works great.

▲simlevesque 3 hours ago | parent [-]

I don't want to rely on Claude of Codex mobile features at all. Also none of this works when you use third party LLMs which is what the current comment tree is about.

▲windexh8er 2 hours ago | parent [-]

This. There's so many better ways to drive a fleet of agents this way than using "remote" features of which I don't trust, anyway.

▲pimeys 3 hours ago | parent | prev | next [-]

https://omp.sh/ has amazing defaults and it sips tokens. Works really well with DeepSeek V4.1 Flash.

Use the model through a fast and reliable provider such as Fireworks directly, skip OpenRouter.

▲kleinishere 2 hours ago | parent [-]

Did you ever try pi by itself? For those new to the pi ecosystem - any rationale to go with pi vs omp?

▲pimeys an hour ago | parent [-]

It's like choosing between vim and helix. I started my career with vim in the early 2000's, customized the whole thing and had my config in a version control.

Then I installed helix and I just use it without config.

If you like configuring things take pi, if not omp is pretty much great defaults.

▲phyalow 4 hours ago | parent | prev | next [-]

I have a server living in my home office, always on. I have a tmux session on it with vanilla Codex and Claude Code CLI, I can via my Ubiquiti network stack wiregaurd in to this box anywhere on the globe with just my laptop. Works super well for me. I also have some cheap shelley power plugs that I can use to cycle my PC’s power state if needed.

▲marknutter 3 hours ago | parent [-]

This is basically my setup but I'm using tailscale and zellij. I don't have any contingency plan in place for my power or home internet going down though..

▲dolebirchwood 4 hours ago | parent | prev | next [-]

OpenCode works nicely for me. You can connect it to the DeepSeek platform with an API key.

▲proxysna 4 hours ago | parent | prev | next [-]

I use pi.dev, it is pretty minimal, but you can extend it however you want since agent has access to it's own documentation.

▲ThomasGlanzmann 4 hours ago | parent | prev [-]

I use crush (https://github.com/charmbracelet/crush) with the following patches:

  curl https://tg.st/u/0001-fix-unblock-all-commands-in-bash-tool.patch | git am
  curl https://tg.st/u/0002-feat-add-light-theme-with-auto-detection-for-white-b.patch | git am
  curl https://tg.st/u/0003-feat-enable-yolo-mode-by-default.patch | git am
  curl https://tg.st/u/0004-fix-disable-mouse-grabbing-to-restore-native-termina.patch | git am
  curl https://tg.st/u/0005-feat-skip-project-init-prompt-and-quit-immediately-o.patch | git am
  curl https://tg.st/u/0006-feat-remove-scrambled-rune-animation-from-waiting-sp.patch | git am
  curl https://tg.st/u/0007-feat-remove-quit-banner-and-thank-you-message.patch | git am
  curl https://tg.st/u/0008-feat-show-output-in-full-instead-of-collapsing-trunc.patch | git am
  curl https://tg.st/u/0009-fix-discover-map-model-features-advertised-by-v1-mod.patch | git am
  curl https://tg.st/u/0010-feat-keep-large-and-small-model-selections-in-sync.patch | git am
▲ApolloFortyNine 2 hours ago | parent | prev | next [-]

The speed of deepseek is insane to experience after using claude code with opus for so long. Not only is the tps roughly 3x faster, but the round trip times are magnitudes faster.

▲mrbonner 2 hours ago | parent | prev | next [-]

$200/month is for Navier-Stoker grade problem.

▲paulddraper 33 minutes ago | parent | prev | next [-]

Deepseek isn't as good as Kimi, let alone the frontier models.

It's a very noticeable hit. It's like programming with mid-2025-era models: ignored instructions, dead code, mistakes.

It works...but anyone going from Sol to Deepseek is going to have a rough transition.

▲dzink 3 hours ago | parent | prev | next [-]

Where do you do your inference?

▲giancarlostoro 3 hours ago | parent | prev | next [-]

Are you just using it directly from them?

▲case540 an hour ago | parent | prev | next [-]

You clearly haven’t used opus or astra or only had simple tasks. Such a difference maker

▲m3kw9 3 hours ago | parent | prev | next [-]

Kind of ambiguous without saying token amounts and cost.

▲UltraSane 2 hours ago | parent | prev | next [-]

Opus 5.5.is crazy good. I ask it to do things and it just writes the code to do it.

▲pdntspa 3 hours ago | parent | prev | next [-]

I just ran a huge text/image extraction grudgematch against all the current inexpensive models except gpt-5.5/5.6/6 (due to some issues with openrouter and bugs in my code) and DS4 ranked very poorly. Accuracy winner was Gemini 3.8 flash with minimax M3 and qwen 3.8 placing, and the chinese models beat the incumbent (Gemini 2.5 Flash) on cost whilst keeping like 95% of the accuracy.

I haven't used deepseek for anything else but the above results make me question its overall capability. Meanwhile qwen3.8 has continued to impress.

▲FailMore 4 hours ago | parent | prev | next [-]

API pricing?

▲alfalfasprout 3 hours ago | parent | prev [-]

It's trivial to hit that kind of quota if you're trying to execute on major projects. Especially as you start having dozens or hundreds of subagents investigating, prototyping, and working on different things.