| ▲ | adev_ a day ago | ||||||||||||||||
> Curious, why do you say the model is mediocre? I haven't tried it, so I can't pass any judgement. It's nowhere mediocre. It's toes-to-toes with GLM-5.3 which is one of the best Open Weight model available (With Kimi K3) for general reasoning. I just runned it on code reviews right now and it was able to catch some thread safety issue than DeepSeek-4.1 didn't. And DeepSeek-4.1 is by no means a bad model. | |||||||||||||||||
| ▲ | user43928 14 hours ago | parent [-] | ||||||||||||||||
It's better than I expected for Mistral. That said, it hardly seems like the innovative underdog some comments here make it out to be. GLM-5.3 and Kimi K3 released 2 and 3 months ago respectively. For all I know anyone could train a model of similar performance by just applying already published research to a new large training run. | |||||||||||||||||
| |||||||||||||||||