Remix.run Logo
porphyra 7 hours ago

Why do they only host small models rather than the 2.4T version? Is the I/O and interconnect between the wafers bad due to the limited beachfront relative to the massive size of the chip?

gardnr 7 hours ago | parent | next [-]

They make a giant inference chip. Their inference service is basically just advertising for their core value prop: hardware.

The CEO was on Gradient Dissent a couple years ago: https://www.youtube.com/watch?v=qNXebAQ6igs

codexon 6 hours ago | parent | prev | next [-]

The wafer only has space for 44 gb of sram. If they offload ram they lose the speedup of having everything on 1 chip (the whole point of cerebras).

porphyra 6 hours ago | parent | next [-]

They can host larger models by pipelining it on multiple wafers. Each wafer stores one layer and N layers can serve an N * 44 gb model with N concurrency. The limitation would of course be inter-wafer I/O, which my comment was getting at. That's probably how they can serve bigger models like GPT 5.6 Sol [1].

[1] https://www.cerebras.ai/blog/accelerating-gpt-5-6-sol-ultraf...

codexon 6 hours ago | parent [-]

I never said offloading was impossible. It will result in a large slowdown.

It would look bad for cerebras if other people are hosting the 27b version and show a higher TPS than cerebras.

minimaltom 6 hours ago | parent | prev [-]

[dead]

altertable 7 hours ago | parent | prev [-]

Mostly economics I'm sure