| ▲ | perching_aix 3 hours ago | |
Never heard of it before, that's fucking insane. Apparently they baked the Llama 3.1 8B model weights [0] into silicon (the actual hardware is called Taalas HC1). I guess for the trillion parameter models this would not scale due to cost? Imagine buying GPT 6 in the form of a PCI-E card, pulling these speeds, with up to 120 cct agent sessions. It'd be beyond wild. [0] the weights are also using some cut down small format, but HC2 will have regular FP4 supposedly, and support for 20B params on one die | ||