Remix.run Logo
▲ amelius 7 hours ago

> have not been a winner-take-all runaway acceleration game where catchup is impossible

From the Mistral site:

> ML4 was trained from scratch on 3,800 NVIDIA Grace Blackwell GPUs in Mistral’s own datacenters in Europe.

It is pretty capital intensive!

▲eigenspace 6 hours ago | parent | next [-]

That cluster is literally orders of magnitude smaller than the compute pools used by Anthropic or OpenAI.

▲amelius 4 hours ago | parent [-]

For training or for inference?

▲ricardobeat 4 hours ago | parent | next [-]

They don't publish numbers, but Anthropic has a single DC with 200k+ GPUs for inference, GPT-6 Astra is said to have trained on 100k+ GPUs.

▲anvuong 3 hours ago | parent | prev [-]

Both, especially for training. Astra and Fable were presumably trained on cluster of 100,000k GPUs, or at least a couple of 10Ks.

3,800 GPUs is nothing in the frontier side.

▲bjenkins358 6 hours ago | parent | prev | next [-]

I’m pretty impressed that they managed to get that close to the frontier with such a small cluster!

▲locknitpicker 5 hours ago | parent [-]

> I’m pretty impressed that they managed to get that close to the frontier with such a small cluster

Chinese companies also managed to put together their models with relatively small clusters.

Perhaps US companies are desperately trying to brute force their way into workable models?

▲everfrustrated 6 hours ago | parent | prev | next [-]

According to Grok thats 7-10 MW. Tiny numbers.

To put that into context, the last wave of capacity SpaceXAI added 400-450 MW.

▲amelius 6 hours ago | parent [-]

But how much of that are they using for training versus inference? They're serving quite a large user base.

▲jayd16 5 hours ago | parent | prev | next [-]

These cards are like $3k each? That's, what, $12M and you keep the hardware? Honestly doesn't seem too bad.

▲amelius 5 hours ago | parent [-]

More like $30k each.

▲jayd16 5 hours ago | parent [-]

Oh the server chip is 10x. That makes a lot more sense.

▲dannyw 5 hours ago | parent | prev [-]

That’s kinda very small and light for modern trillion-param LLMs.