Remix.run Logo
▲ porridgeraisin 5 hours ago

First few models will always be slow improving and worse. The way to improvement is working your way through a gajillion evals [1], finding bugs, gaps, and curating training data (this part involves human design as well as raw inference compute) to fix it. This is very time intensive and can't easily be "done once and then everyone has lesser work to do" since every model is different. Well, one way to accelerate it is to simply have more compute, which mostly openai and anthropic have[2].

This is mistrals first 1T-scale model and I expect the 4th or 5th generation to be close to the best for many purposes.

[1] These evals differ from the public ones like terminal-bench, are sometimes model-specific, need real, diverse usage to actually create, and are held secretly since quality of eval is the first driver behind the next step improvement of a model.

[2] It is not close. This model was trained on less than 4k GPUs, whereas astra used north of 100k GPUs.

▲manlymuppet 5 hours ago | parent [-]

Mistral is not that new a player though. How can we give them this much grace when other players like xAI have done more in even less time? I don't think coddling Mistral helps them.

And to the point of scale and training cluster, so what? Not only do Chinese labs have smaller clusters with less empowered GPUs, compute is Mistral's responsibility. You can't take away from other labs just because they fulfill that responsibility better.

▲ 3 hours ago | parent | next [-]
[deleted]
▲porridgeraisin 3 hours ago | parent | prev [-]

xAI has a lot of compute. Deepseek also has a lot, not as much though. But this is changing with their new 160k huawei ascend datacenter in inner mongolia.

The lack of compute is not really attributable in that sense to mistral. First of all it needs general investor and government willingness, which is easier in a larger economy like the US or China.

Second, you need widespread usage of your paid inference service for two reasons: one it pays off your compute cost, and two it speeds up the improvement process.

The vast majority of deepseeks paid customers are within china itself (since openai and anthropic services are not reachable from china) which gives it a market. But for someone in france, there is no reason to use a structurally slower developing model from mistral compared to using one from openai...except when data guarantees are needed, hence the landing page focus on sovereignty. As far as the dual use aspect goes, a model like this is more than enough, so the government will be happy.