| ▲ | simjnd a day ago | |
It's a bit below DeepSeek 4.1 Flash at about twice the size. For a model that was supposed to come out a few months ago this is pretty good. Mistral catching up to the chinese open-weight models is great news. Excited to see how they will build on that! | ||
| ▲ | pimeys a day ago | parent | next [-] | |
They will RL it to be better. What I'm looking for with this is of it hallucinates less that DeepSeek with its sparse attention, and can be used in better summarization and content generation in multiple languages. DeepSeek sucks with languages other than English and China. The price is really good with ML4. | ||
| ▲ | kristianp 18 hours ago | parent | prev [-] | |
I noticed they compared with the previous best deepseek model, 4 pro 0813 for the two coding benchmarks (DeepSWE and Terminal Bench) and not Deepseek 4.1 flash. A glaring omission really. | ||