Remix.run Logo
dzonga 3 days ago

to me the biggest event more than deepseek launch was when zAI served their latest model on all Chinese chips.

genxy 3 days ago | parent [-]

Model serving is trivial, and inference is just memory bandwidth. The cost of serving will be asymptotic to flash read energy.

Having trained on your own chips, that is the impressive part.

g8oz 3 days ago | parent [-]

Model serving and inference will be trivial in the long term, right now it's a real choke point. Being able to do that with domestic chips is an important win.