| ▲ | sixtyj 5 hours ago |
| Output from Cerebras with GPT model is 750 tokens per second. Don’t blink. (Chatjimmy has 14,200 TPS.) |
|
| ▲ | tomrod 5 hours ago | parent | next [-] |
| ChatJimmy is a much smaller model and, AFAIK, has no reasoning capability. Absolutely insane raw speed, like a supercar, while Sol is more like a freight truck. |
| |
| ▲ | sixtyj 5 hours ago | parent | next [-] | | At such output speed, I wouldn’t expect reasoning. (But I didn’t know it, thanks.) 700 TPS with reasoning is awesome and it speeds things up. Cerebras as public traded company is worth keeping an eye what they produce. | | |
| ▲ | msdz 5 hours ago | parent [-] | | > At such output speed, I wouldn’t expect reasoning. As the sibling comment to yours mentioned, if they had a reasoning model “hardware-ified” onto a custom chip (as is their plan for IIRC this or next year, a new ASIC), it’d output fast decode speeds for the regular output as well as reasoning sections. Both would be ≈equally fast. | | |
| ▲ | senordevnyc 2 hours ago | parent [-] | | Yeah, I thought reasoning was literally just chain of thought in the output token stream, with the model itself adding delimiters to indicate what part of the output is internal reasoning, and what part is an answer to the user. Is that wrong? | | |
|
| |
| ▲ | notfromhere 5 hours ago | parent | prev | next [-] | | Anything will be fast if you etch it straight to silicon | |
| ▲ | dzhiurgis 4 hours ago | parent | prev [-] | | The knowledge of ChatJimmy is terrible. Even Qwen on my iPhone is better. | | |
| ▲ | walrus01 an hour ago | parent | next [-] | | well, yeah, it's based on a 2+ year old tiny model. It's very much an alpha proof of concept that they can perma-bake an LLM into silicon. https://huggingface.co/meta-llama/Llama-3.1-8B | |
| ▲ | headPoet 3 hours ago | parent | prev | next [-] | | ChatJimmy isn't a model, it's Llama 3.1 8B hardwired into silicon. The point isn't to be a good llm, but to showcase the speedup that's possible | |
| ▲ | senordevnyc an hour ago | parent | prev [-] | | Haha, at first I thought you meant that the knowledge of the existence of an LLM that’s so fast is terrible because it’s ruined every other LLM for you! |
|
|
|
| ▲ | perching_aix 3 hours ago | parent | prev | next [-] |
| Never heard of it before, that's fucking insane. Apparently they baked the Llama 3.1 8B model weights [0] into silicon (the actual hardware is called Taalas HC1). I guess for the trillion parameter models this would not scale due to cost? Imagine buying GPT 6 in the form of a PCI-E card, pulling these speeds, with up to 120 cct agent sessions. It'd be beyond wild. [0] the weights are also using some cut down small format, but HC2 will have regular FP4 supposedly, and support for 20B params on one die |
|
| ▲ | mips_avatar 4 hours ago | parent | prev [-] |
| Unfortunately AMD bought them, so I don't think we will get to see another release from them. |