All benchmarks I could find are Microsoft-reported. I don’t see it listed on any leaderboards. Still, looks interesting enough.
Also, I wonder what this means in practice, especially in terms of GitHub:
> MAI-Thinking-1 was trained on clean and appropriately licensed data, with AI-generated content excluded from pre-training.
Aaaand the question answers itself, because the above sentence is now gone from the page, replaced with:
> We trained it from the ground up on clean, traceable and enterprise-grade data, without distillation from third-party models.