| ▲ | LoganDark 4 hours ago | ||||||||||||||||
1. No 2. They don't have enough capacity either The current largest FPGA, the AMD Versal Premium VP1902 has 18.5 million logic cells. That's not even enough for the smallest whisper.cpp model (75M). You'd have to order hundreds of thousands of them (or millions) to serve even a single copy of a frontier model, and at that scale inference quickly becomes starved by the speed of light. | |||||||||||||||||
| ▲ | CamperBob2 3 hours ago | parent [-] | ||||||||||||||||
Well, you'd use BRAM to store model weights, not fabric. But still, you only get a couple hundred MB for probably close to US $100k per chip. It's likely that the major FPGA vendors will soon announce parts specifically architected to support LLMs and similar models. But the current generation isn't suitable for that at all. | |||||||||||||||||
| |||||||||||||||||