Remix.run Logo
wyre 18 hours ago

6x. Taalas has Llama3.1 8B running at 18000 tok/s. Cerebras advertised that model at 3000 tok/s.

quotemstr 18 hours ago | parent | next [-]

Interesting. Is the speedup from specializing for the shape of Llama 3.1 or are they (contra my mental model) actually winning on burning in the weights?

akiselev 17 hours ago | parent [-]

The weights are in SRAM, so the LLM architecture is burned in but the weights can be updated.

Tuna-Fish 13 hours ago | parent | next [-]

Taalas only uses SRAM for the KV cache and the activations, the weights are in mask rom in the metal layers.

If they designed this right, it means that once they have a model, so long as they keep the hyperparameters fixed they can change the weights much faster than it takes to spin up a completely new chip, essentially at a cost of doing a minor revision.

quotemstr 12 hours ago | parent [-]

Is the mask ROM really going to be worth it over Carmack's high-bandwidth-flash concept? I mean, sure, I could be convinced I guess, but it's not obvious.

hedgehog 16 hours ago | parent | prev [-]

No the weights are in the metal layers, they cannot be updated.

jfim 12 hours ago | parent [-]

The base weights can't be updated but from what I recall it allows adding a low rank adapter to customize the model a little bit.

HNisCIS 18 hours ago | parent | prev [-]

8B is still three orders smaller than current frontier though