| ▲ | softwarewright 2 days ago | |||||||
Tang nano FPGAs are around $10, the idea being to couple with a $25 ARM or RISC-V dev board and offload the math from a streaming store. Coupling this with an older 8GB VRAM board can theoretically run a model whose weights do not fit. The weights get repeatedly run through the MCU/FPGA and the VRAM is used for KV Cache and context. If you connected several via USB and streamed a MoE model expert per FPGA, you could achieve parallelism. Won't be fast, but, many requests could be processed in parallel, as each request shares the large static weights. The reason for this is, model weights do not need to be randomly accessed. So why store them in expensive RAM. Cerebras and Qrok seem to be using a very different approach than NVIDIA to get orders of magnitudes speed ups. I'm trying to explore other alternative approaches. | ||||||||
| ▲ | wmf 2 days ago | parent [-] | |||||||
I predict that the FPGA adds no value in this scenario. Just process inference on the CPU. | ||||||||
| ||||||||