| ▲ | idbnstra 6 hours ago | |||||||
just curious, how do we know that le chonk isn't just a fine-tuned chinese model? and/or distilled from US models? | ||||||||
| ▲ | rahen 6 hours ago | parent | next [-] | |||||||
It's in the announcement: "ML4 was trained from scratch on 3,800 NVIDIA Grace Blackwell GPUs in Mistral’s own datacenters in Europe". | ||||||||
| ▲ | karp773 an hour ago | parent | prev | next [-] | |||||||
It would not take them so long to train it. Their pace would be closer to the Chinese models. | ||||||||
| ▲ | Iolaum 6 hours ago | parent | prev [-] | |||||||
Once it's open weight people will be able to inspect and compare it's tokenizer, architecture etc and tell. | ||||||||
| ||||||||