Remix.run Logo
▲ guyomes 3 hours ago

If we throw in hardware dedicated to a specific LLM, it seems to be a rather low hanging fruit. Especially considering that this is already happening for vision models [1].

[1]: "FPGA-based CNN Acceleration using Pattern-Aware Pruning" https://inria.hal.science/hal-04689673/document

▲mdp2021 35 minutes ago | parent [-]

> hardware dedicated to a specific LLM

That wording screams "Taalas". Which, importantly, is not the only player trying to abate the distance between data and arithmetics...