| ▲ | guyomes 3 hours ago | |
If we throw in hardware dedicated to a specific LLM, it seems to be a rather low hanging fruit. Especially considering that this is already happening for vision models [1]. [1]: "FPGA-based CNN Acceleration using Pattern-Aware Pruning" https://inria.hal.science/hal-04689673/document | ||
| ▲ | mdp2021 35 minutes ago | parent [-] | |
> hardware dedicated to a specific LLM That wording screams "Taalas". Which, importantly, is not the only player trying to abate the distance between data and arithmetics... | ||