| ▲ | dnhkng 12 hours ago | |||||||||||||||||||||||||||||||
DeepSeek V4 Flash (Preview → 2026-07-31) • Terminal Bench: 56.9 → 82.7 (+25.8) • Toolathlon: 51.8 → 70.3 (+18.5) Compared to GPT-5.6 Terra: • Terminal Bench: Flash 82.7 vs Terra 78.4 • Toolathlon: Flash 70.3 vs Terra 53.1 • DeepSWE: Flash 54.4 vs Terra 69.6 • Agents' Last Exam: Flash 25.2 vs Terra 50.4 Trading blows with Terra, which is pretty interesting. No clear winner on these benchmarks, and wildy differeing scores. Very interesting! | ||||||||||||||||||||||||||||||||
| ▲ | villish 11 hours ago | parent | next [-] | |||||||||||||||||||||||||||||||
> Terminal Bench: Flash 82.7 vs Terra 78.4 Terra 87.4 | ||||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||
| ▲ | throwaw12 12 hours ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||
Open flash model is competing against OpenAI's 'Sonnet' model at the price of GPT 3, I am really excited about this release, hopefully it holds up in real work as well | ||||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||
| ▲ | Iolaum 12 hours ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||
Since they did this with their own harness I m not sure it's apples to apples comparison. | ||||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||
| ▲ | yms_hi 12 hours ago | parent | prev [-] | |||||||||||||||||||||||||||||||
I think it's better than GPT Luna. | ||||||||||||||||||||||||||||||||