Remix.run Logo
dada216 7 hours ago

We built a complete production-grade inference service from scratch on a cluster of more than 100,000 Chinese-made AI accelerators. All production inference for GLM-5.3-Flash runs on this system.

dzonga 2 hours ago | parent | next [-]

very few people comprehend - how much of an asteroid level event for western AI labs this is.

china has cheap abundant power, now they can make their own inference chips (which was supposed to be a chokepoint), their models yeah can be 6 months behind the frontier - but most people don't need frontier models - small models r more than enough.

my only wish was labs like Mistral would make their own inference chips or partner up eg with established / new chip makers or companies like Oxide.

vatsachak an hour ago | parent [-]

After AI agents get good enough the real bottleneck will be power generation and political systems.

freakynit 6 hours ago | parent | prev | next [-]

Most of the people had kinda guessed this when they decided to provide 100 trillion tokens for free.

gpugreg 4 hours ago | parent [-]

It wasn't a secret either. They blogged about it last month: https://z.ai/blog/glm-5.3-flash#:~:text=Serving%20at%20Scale...

axiosgunnar 7 hours ago | parent | prev [-]

[dead]