| ▲ | dhx 2 hours ago | |
Now that people have played with the model for a bit, there is one reported real world win for DeepSeek v4 Pro 0813 that is perhaps quite consequential given the drama in the US about Mythos. In a benchmark, DeepSeek v4 Pro 0813 found 87.5% of selected real world software vulnerabilities publicly reported and with CVEs assigned, which is above runner ups Opus 5 and Qwen 3.8 which both found only 81.3%. However there is a downside to this--DeepSeek v4 Pro 0813 is less accurate with a 35% false positive rate versus GPT-5.6-Sol's 15% false positive rate. For vulnerability analysis though, it's probably worth finding that one extra vulnerability no other model has found even if requires significantly more triage to remove false positives, or additional cost to run every vulnerability detection through other models to verify. [1] https://nitter.net/pilvar222/status/2087691659953815783#m | ||