Remix.run Logo
▲ idbnstra 6 hours ago

just curious, how do we know that le chonk isn't just a fine-tuned chinese model? and/or distilled from US models?

▲rahen 6 hours ago | parent | next [-]

It's in the announcement: "ML4 was trained from scratch on 3,800 NVIDIA Grace Blackwell GPUs in Mistral’s own datacenters in Europe".

https://mistral.ai/news/mistral-large-4/

▲karp773 an hour ago | parent | prev | next [-]

It would not take them so long to train it. Their pace would be closer to the Chinese models.

▲Iolaum 6 hours ago | parent | prev [-]

Once it's open weight people will be able to inspect and compare it's tokenizer, architecture etc and tell.

▲ismailmaj 4 hours ago | parent [-]

it's extremely unlikely that they re-use anything from a chinese model, that would be obvious quickly, what's more likely is using documents produced by a better model to create synthetic data.