Remix.run Logo
onlyrealcuzzo a day ago

Isn't this quite a bit behind Sol and Fable and even ChatGPT 5.5 xhigh and Opus 5 max?

In terms of what you get for what you pay for, it's incredible - probably by far the best.

But unless I'm reading things wrong, it does not appear to be top-of-the-line.

dhx 2 hours ago | parent | next [-]

Now that people have played with the model for a bit, there is one reported real world win for DeepSeek v4 Pro 0813 that is perhaps quite consequential given the drama in the US about Mythos.

In a benchmark, DeepSeek v4 Pro 0813 found 87.5% of selected real world software vulnerabilities publicly reported and with CVEs assigned, which is above runner ups Opus 5 and Qwen 3.8 which both found only 81.3%. However there is a downside to this--DeepSeek v4 Pro 0813 is less accurate with a 35% false positive rate versus GPT-5.6-Sol's 15% false positive rate. For vulnerability analysis though, it's probably worth finding that one extra vulnerability no other model has found even if requires significantly more triage to remove false positives, or additional cost to run every vulnerability detection through other models to verify.

[1] https://nitter.net/pilvar222/status/2087691659953815783#m

neosat a day ago | parent | prev | next [-]

This may not 'quite a bit behind' those at all. If you look at the benchmark numbers they are very comparable to Fable, but beyond a certain point the benchmark numbers don't tell you much. Opus #5 beats Fable on some benchmarks but given similar cost almost everyone who has used those two models will prefer to use Fable.

At this price range $0.87per 1M they will get a lot of usage of people trying it out. Given the benchmark numbers, for many people and many use cases this will become their primary driver. There are people and use cases where Fable, Sol will work better but those are likely not the target of DeepSeek anyway.

In terms of performance and price pareto curve I don't think any model can beat this today (though openAI is doing some exciting recent work in efficiency) - which is a remarkable feat for the DeepSeek team.

Either way, what a time for consumers of these models :)

SwellJoe 14 hours ago | parent [-]

I stopped picking Fable because it refuses based on guardrails so often. I do a lot of security related work, and Fable just won't do any of it. So, I don't bother. Unfortunately, Opus 5 also refuses quite a bit of security work, now, as well, so my Anthropic subscription becomes less useful by the day. Fable may be better, but if it won't do the work...

DeepSeek and Kimi K3 will happily do security work, and they do it pretty well.

vinnymac 5 hours ago | parent [-]

Same, I have gotten good results out of Opus 4.6,4.7,4.8 for my security work though. So I continue to use them for this.

Curious if you’ve find yourself enjoying DS or Kimi more than Opus 4?

SwellJoe an hour ago | parent [-]

DeepSeek is more fun, because it's cheap-as-free, Good Enough, quite fast. I use it for all API stuff, automated runs, testing of the security auditing harness and benchmarks I'm working on, etc.

Kimi K3 is smarter, though. At least smarter than DeepSeek V4 Flash 0731. I haven't tried the new Pro version, but will this weekend when I'm working on my personal projects But, K3 has been what I've been using for the actual coding of the harness and such (after Claude models, and then OpenAI models, began refusing to do that work). K3 is very expensive, though. Much more expensive than pretty much everything except Anthropic, and their subscription plans are stingy.

I also like Reasonix quite a bit, as an agent harness, though Kimi Code is also very good. I guess I'll try out the new DeepSeek official harness, as well.

a day ago | parent | prev [-]
[deleted]