Remix.run Logo
LaurensBER a day ago

> With a massive 2.4T parameters, this model is continuously evolving. We believe it’s one of the most powerful model available today, compatible to leading frontier AI models , second only to Fable 5.

That's a massive model!

The shift from "value" models to "intelligent, huge and slow" models coming from China is an interesting change in strategy.

My main issue with GLM 5.2 and Kimi 3 is that they're extremely token hungry and thus feel slow(er) to use.

hodgehog11 a day ago | parent | next [-]

Value models are always going to be there; you can always distill from a larger model. Having a really intelligent model, regardless of the size, is much better for building confidence in your brand. That is a big reason why the US companies are still hanging in there.

charcircuit a day ago | parent | prev [-]

The shift isn't new. Kimi K2, a 1T model came out July last year. I am happy that more labs are following the trend as its important for competitive open models to exist.

selcuka a day ago | parent | next [-]

Also DeepSeek R1 was announced 1.5 years ago with ~0.7T parameters, which was a huge model back then.

chronogram a day ago | parent | prev [-]

And DeepSeek has been making huge progress on efficiency, and publishing about, so they came with a 1.6T model that is both fast and cheap to run.