Remix.run Logo
▲ LoganDark 4 hours ago

1. No

2. They don't have enough capacity either

The current largest FPGA, the AMD Versal Premium VP1902 has 18.5 million logic cells. That's not even enough for the smallest whisper.cpp model (75M).

You'd have to order hundreds of thousands of them (or millions) to serve even a single copy of a frontier model, and at that scale inference quickly becomes starved by the speed of light.

▲CamperBob2 3 hours ago | parent [-]

Well, you'd use BRAM to store model weights, not fabric. But still, you only get a couple hundred MB for probably close to US $100k per chip.

It's likely that the major FPGA vendors will soon announce parts specifically architected to support LLMs and similar models. But the current generation isn't suitable for that at all.

▲LoganDark an hour ago | parent [-]

Would BRAM even have enough bandwidth? The reason I quoted logic cells is because that's the way to get instant throughput, which is practically the only reason to use an FPGA over something like a TPU.

▲CamperBob2 an hour ago | parent [-]

I think so, because memory bandwidth really comes from bus width more than clock speed. You can construct 36-bit wide BRAM arrays with bus width comparable to HBM, just by specifying multiple BRAM arrays in parallel.

Never tried anything like that, though.