| ▲ | scrlk 6 hours ago |
| Benchmarks: | Benchmark | DS-V4-Pro | DS-V4-Flash | DS-V4-Pro | DS-V4-Flash | GLM-5.2 | Kimi-K3 | Opus-4.8 | Fable 5 |
| | 0813 | 0731 | Preview | Preview | | | | (w/ fallback) |
|--------------------------|-----------|-------------|-----------|-------------|-----------|-----------|-----------|---------------|
| HLE (wo/w tools) | 42.7/60.0 | 37.8/51.5 | 37.7/48.2 | 34.8/45.1 | 40.5/54.7 | 43.5/56.0 | 49.8/57.9 | 53.3/63.0 |
| Terminal Bench 2.1 | 87.9 | 82.7 | 72.1 | 61.8 | 81.0 | 88.3 | 85.0 | 88.0 |
| NL2Repo | 61.5 | 54.2 | 38.5 | 39.4 | 48.9 | - | 69.7 | - |
| Cybergym | 83.3 | 76.7 | 52.7 | 38.7 | - | 80.0 | 78.3 | 83.1 |
| DeepSWE | 62.7 | 54.4 | 12.8 | 7.3 | 46.2 | 67.5 | 58.0 | 70.0 |
| Toolathlon-Verified | 74.1 | 70.3 | 55.9 | 49.7 | 59.9 | 76.5 | 76.2 | 77.9 |
| Agents' Last Exam | 25.7 | 25.2 | 16.5 | 15.8 | 23.8 | 27.6 | 25.7 | - |
| AutomationBench (Public) | 31.8 | 25.1 | 12.8 | 10.8 | 12.9 | 30.8 | 27.2 | 29.1 |
| DSBench-FullStack | 71.1 | 68.7 | 41.8 | 37.0 | 61.8 | 73.7 | 71.6 | 77.2 |
| DSBench-Hard | 67.2 | 59.6 | 31.1 | 25.8 | 54.5 | 63.0 | 71.7 | 68.3 |
Source: https://reddit.com/r/LocalLLaMA/comments/1vmi0fg/deepseek_v4... |
|
| ▲ | parsimo2010 6 hours ago | parent | next [-] |
| The timing looks like they are trying to take the wind out of Qwen's sails by releasing this on the same day that Qwen released the weights of Qwen3.8-max. Or maybe it's coincidence... For comparison I looked at Qwen's claimed benchmarks for Qwen3.8-max (https://qwen.ai/blog?id=qwen3.8). Assuming each published set of benchmarks is believable, it looks like v4 Pro 0813 is better on average but overall performance is comparable. Pro 0813 is much cheaper. If you don't need vision capabilities then you don't have much reason to use Qwen3.8-max. - 43.6 on HLE (Presumably without tools). Pro 0813 is a little worse. - 86.6 on Terminal Bench 2.1. Pro 0813 is better. - 55.9 on NL2Repo. Pro 0813 is better. - 27 on Agent's Last Exam. Pro 0813 is a little worse. - 72.5 on Toolathon-Verified. Pro 0813 is better. - 56.6 on DeepSWE 1.1. If the DeepSWE listed for Pro 0813 is the same version, then Pro is better. - 27.3 on AutomationBench. If the AutomationBench (Public) listed for Pro 0813 is the same, then Pro is better. I guess we do need to wait to see if the upcoming DS pricing increase is enough to change the value proposition. As it is now, they could double or triple prices and it still would be a better value to use DS. I bet they know that. |
| |
| ▲ | trollbridge 6 hours ago | parent | next [-] | | By that standard, the release of Grok 4.6 was also timed on the same day. Given how I think DeepSeek operates... I think they just release it when they feel it's ready, and don't even seem that concerned with what other people are doing. | | |
| ▲ | somenameforme 5 hours ago | parent | next [-] | | Their leaks would confirm this sort of attitude. They're not trying to become the top player or anything like that - just working to play their part in pushing LLM tech forward and going from there. It was quite refreshing from the 'here's how we're going to dominate the world' nonsense. It's undoubtedly the same attitude that just lets them shrug and cancel the fund raising round after the leaks came from said funding round. | | |
| ▲ | trollbridge 5 hours ago | parent | next [-] | | The founder of DS's stated goal is to get to AGI. He thinks this is the path to get there. Kind of interesting, when compared to the hubris from American frontier labs. | | |
| ▲ | johnvanommen 5 hours ago | parent [-] | | > Kind of interesting, when compared to the hubris from American frontier labs. One Man’s “hubris” is another man’s “marketing campaign.” Drama sells. |
| |
| ▲ | scrlk 5 hours ago | parent | prev | next [-] | | Benefits of having a well performing hedge fund funding DeepSeek. IIRC, Demis attempted to start a fund inside DeepMind but it was killed off. In an alternative world where he manages to pull that off, perhaps DeepMind would still be independent with Demis at the helm. | | | |
| ▲ | surgical_fire 5 hours ago | parent | prev [-] | | Their stance on LLM development is why they earned my respect in a time when OpenAI and Anthropic only earn my mistrust. That, and the fact that DS is an insanely capable model. |
| |
| ▲ | parsimo2010 5 hours ago | parent | prev [-] | | Actually, yes. I just didn't know about Grok's release because they aren't on the front page of HN. |
| |
| ▲ | eli 6 hours ago | parent | prev | next [-] | | Official pricing only kinda matters for an open weight model, no? | | |
| ▲ | parsimo2010 5 hours ago | parent [-] | | It still matters as a point of comparison until other providers come online. If the consensus price from other providers is much different that can be compared then. But for now we have $0.435 / $0.87 for v4 Pro 0813 (with increase announced but we don't know the new pricing), and $2 / $6 for Qwen3.8-max. So until we get other data points that is what we have to look at. | | |
| ▲ | eli 5 hours ago | parent [-] | | I wondered if the promised change in pricing is actually going to be deepseek bringing up their cached costs. They're extremely inexpensive. |
|
| |
| ▲ | maherbeg 6 hours ago | parent | prev [-] | | I mean at the rate of model releases happening, I think a lot of these will collide more often than expected! |
|
|
| ▲ | bel8 6 hours ago | parent | prev | next [-] |
| So it's a Fable class LLM? DSV4Pro vs Fable5
HLE w tools 60.0 vs 63.0
Terminal Bench 2.1 87.9 vs 88.0
Cybergym 83.3 vs 83.1
DeepSWE 62.7 vs 70.0
Toolathlon-Verified 74.1 vs 77.9
AutomationBench (Public) 31.8 vs 29.1
DSBench-FullStack 71.1 vs 77.2
DSBench-Hard 67.2 vs 68.3
|
| |
| ▲ | eli 6 hours ago | parent | next [-] | | Fable's guardrails would never let it do something like Cybergym so at least for that one it's measuring Opus 5 | | |
| ▲ | wren6991 6 hours ago | parent [-] | | We have a first-party figure from the system card [1]: > Mythos 5 reproduced 83.8% of targeted vulnerabilities on a single try, and produced at
least one crash in 99.4% of tasks. This is comparable to Claude Mythos Preview, which
reproduced 83.1% of targeted vulnerabilities and produced a crash in 97.1% of tasks. By
contrast, Claude Opus 4.8 achieved a score of 78.1% (95.7% any crash). So their quoted figure exactly matches the figure for Mythos Preview, although they don't state the provenance. It could also quite possibly be an independent measurement of Opus 5. [1]: https://www-cdn.anthropic.com/57a52ea7d8f0e54e8a542e90826608... |
| |
| ▲ | nikcub 3 hours ago | parent | prev | next [-] | | that DeepSWE result is likely most indicative of how you'll find real world usage | |
| ▲ | aftbit 6 hours ago | parent | prev [-] | | Fabble lol | | |
|
|
| ▲ | goldenarm 6 hours ago | parent | prev | next [-] |
| Geometric mean of all these benchmarks : * GPT-5.6 Sol: 65.5 * Fable 5 (w/ fallback): 64.5 * Opus 5: 64.0 * DS-V4-Pro 0813: 62.5 * Kimi-K3: 62.3 * DS-V4-Flash 0731: 55.8 * GLM-5.2: 47.3 |
| |
| ▲ | svachalek 5 hours ago | parent [-] | | Maybe it's me but I don't see how DS Flash is better than GLM at all, much less by a huge gap. I'd probably protest less against Fable and Opus being put at the same level than many would, but there's no denying the two models are a very different experience from each other. I guess where I'm going is no one should pick a model by the benchmarks. | | |
| ▲ | spijdar 4 hours ago | parent | next [-] | | I'm not the most LLM-savvy person around, and I'm not gonna say I've put a ton of effort into practically compared these open models. But, a month or two ago I did do some "practical evaluates" testing GLM 5.2 versus DSv4 (flash/pro) with OpenCode's subscription with some late 80s Unix clone-type work, and this jives with my experience. GLM ended up being far slower, and far more expensive, for approximately the same results. There was never a problem that GLM could solve that DS couldn't solve, faster, and significantly cheaper. I strongly agree that you shouldn't pick a model based on benchmarks. But for me, I found GLM really underwhelming given its cost and speed. DSv4 isn't as good as GPT or Claude or what have you, but it's fast, and pretty darned effective. I can run a 3-bit quant of DSv4 locally on my system with ~15 tokens per second, and for a local model it might be the most overall effective at coding. For what it is, it's extremely impressive. | | |
| ▲ | ApolloFortyNine 3 hours ago | parent [-] | | My experience is the same. Imo it has a lot to do with you/the harness tries to get it to test itself. Deepseek v4 flash seems more than capable of understanding when something has failed, and making changes until it works. I've definitely seen it make mistakes I would expect something like Opus to find, but it works through them on it's own (and for literal pennies). At the end of the day, I think that's one of the most important features of a model. |
| |
| ▲ | segmondy 3 hours ago | parent | prev | next [-] | | It isn't. I run both at home. GLM5.2 Q4 crushes DSv4Flash0731 Q8. I reach for DS for speed and for medium effort level work. If I care about quality I'll reach for GLM5.2 Looking at this release, I'm comparing it to GLM5.2 and it seems to beat GLM5.2, only time/experience will show. If true, then I'm happy. It's much easier to run than Qwen3.8/KimiK3 | |
| ▲ | spiffytech 4 hours ago | parent | prev | next [-] | | In my little social circle DS4F generally substitutes for GLM 5.2 except it's the next best thing to free. | |
| ▲ | platinumrad 5 hours ago | parent | prev [-] | | I think instruction following carries outsized weight in these evaluations. |
|
|
|
| ▲ | myworkaccount2 4 hours ago | parent | prev | next [-] |
| IMO the HLE scores without tools seem to align better with real world performance of the models. To me it feels like the difference between "RL performance" and the pretraining / base "knowledge". Yes you can RL terminal bench to the moon but does the model hold up on out of distribution tasks? Kind of like trying to navigate a dark room with a laser light, vs a flashlight. Laser is going to go a lot farther much more efficiently but only if you are already pointing it at the right place. |
|
| ▲ | andai 3 hours ago | parent | prev | next [-] |
| The most interesting part of this is how Flash scores almost as well on all of them. Haven't tried the new DeepSeek models but I'm assuming the difference is more than these numbers show! |
|
| ▲ | NietTim 5 hours ago | parent | prev [-] |
| In classic reddit fashion the post you linked to is now deleted |
| |
| ▲ | SV_BubbleTime 5 hours ago | parent [-] | | To be fair… I don’t know who still needs to figure out that AI benchmarks are almost all entirely fucking trash, but the great number would surely surprise me. |
|