Remix.run Logo
▲ fhdkweig 4 hours ago

I know FPGAs are more expensive than GPUs, but are they fast enough to justify the extra cost?

▲zdragnar 4 hours ago | parent | next [-]

It isn't just a matter of speed, it's also a matter of model quality. If they take 6 months to burn Fable to chips, and it takes 2 years to break even between design, custom fab, energy savings, etc, are those chips even worth running when the new models that are running on GPUs at that point are producing 10x better quality results?

Sure, your 2.5 year old models are running faster, but you can't drop prices on them without pushing the break even point further out.

If the cost difference isn't incredibly significant, will people even want to pay for the 2.5 year old model, or will they get more value for their money paying more to get better results from the newer model?

There's a lot of open ended questions that I don't have the insiders knowledge for to suggest whether or not such a capital outlay would be a worthy investment.

My guess is that state of the art stuff will stay on GPUs and models burned into chips will be for "good enough" applications that people are still teasing out. Probably highly specialized models in automated sensor units and such.

▲jerf 4 hours ago | parent | prev | next [-]

FPGAs are FPGAs by virtue of putting on the chips vast, vast arrays of wiring that can be controlled by software. Any given utilization of the FPGA will leave large fractions of the chip resources unused. If you've got a highly stereotypical use case FPGAs will have a "highly stereotypical" set of components being unused, where it would be better to use that space instead to do real work. A lot of people only see the "pro" side of the FPGA proposition without realizing they come with some very substantial "cons" that are intrinsic to the way they work.

▲rjh29 3 hours ago | parent [-]

I guess that's why they work in particular niche spaces like a synthesizer where you have a max of 8 voices and every voice goes through the same pipeline (osc / filter / env / amp) and everything is necessarily running all the time. In that sense I suppose they're very good for modelling any kind of analog circuitry?

Even then, while there are some amazing FPGA-based synths available, companies like Korg just put their code on a raspberry pi and call it a day. The same is true for emulators (SNES Mini etc. are also just raspberry pis under the hood iirc)

▲exmadscientist 3 hours ago | parent | next [-]

You don't really get an FPGA for capability. CPUs are much more capable, and they're general-purpose so they can do absolutely anything with about the same efficiency and just a little more code.

You get an FPGA for timing. They're less capable, but (in many common design architectures), they output their results once per clock, every clock, on time, every time. If you can hit a fabric clock of say 100MHz, clocking all the weird logic you can stuff in there, it gives 100 million outputs per second, never skipping a single one for any reason (short of total failure). The penalty is that making a small change to your desired "program" can be very expensive, and many things won't be realistically possible at all. Or at least won't fit into a part that you can buy. But things like audio, video, and high-frequency trading love being able to guarantee timing.

(Of course there are other ways to write your FPGA HDL, but that's one of the more common ones. And you do see DDR-style clocking, and similar, every now and then.)

▲dist1ll 3 hours ago | parent [-]

> with about the same efficiency and just a little more code.

It depends. For some things, CPUs don't even come close. An XCVU13P FPGA can handle 1.2Tbps of full-duplex Ethernet @ 1 billion pps. And that part costs less than a grand at moderate qty, and with significantly less power consumption than a CPU that'd be capable of operating a dataplane at these speeds.

▲exmadscientist 43 minutes ago | parent [-]

Sorry, I spoke pretty sloppily there.

The point I was trying to make is that the CPU is a general-purpose creature and doesn't really care what you want it to do. If you had a CPU that could handle 1.2Tbps of Ethernet packets at 1Gpps, it could do a whole lot of other things involving 1.2Tbps of data flow too, very easily, if someone wrote the software. And more. (But you're probably not getting 2.4Tbps out of it, no matter what you do.)

An FPGA can not. There's plenty of things that those XCVU13Ps just can't do, or would do worse than a $1 microcontroller. (Setting aside for a moment implementing a CPU inside the FPGA... which does actually happen in just about every large-enough FPGA design, which is its own discussion....)

▲LoganDark 3 hours ago | parent | prev [-]

> In that sense I suppose they're very good for modelling any kind of analog circuitry?

That would be better suited to FPAAs (field programmable analog arrays). FPGAs can usually only work with clocked digital signals.

▲monocasa 4 hours ago | parent | prev | next [-]

They're not magical go faster juice. I don't know of a microarch where they're faster than modern GPUs at ML training or inference.

▲Marha01 44 minutes ago | parent [-]

Having the model weights in static mask ROM would massively improve power efficiency for inference (see what Taalas is doing).

▲monocasa 10 minutes ago | parent [-]

But that's not an efficiency an FPGA provides in the first place.

Additionally, it's not clear how well large mask roms scale. For instance the Nintendo switch cartridges were expected to be mask roms, but instead are Macronix's XtraRom technology, which is essentially a flash cell array made denser by removing the erase functionality. So basically the die gets manufactured with all bits at the same state, a late manufacturing step either empties or fills the floating gate of the bits you want different, and then it's treated as pretty close to a mask rom. It's not even clear if the bits can be changed without a bare die, a floating probe array, and specialized hardware. Though, like flash the electrons in the floating gates will eventually tunnel and cause the data to bitrot.

So from that it appears that even at the tens of millions of chips volumes that would make sense for essentially whatever size of maskrom, the memory manufacturers tap out at 128megabit for a mask rom chip, and push you towards something flash esque.

And at the end of the day, flash without the erase functionality is pretty damn close to a mask ROM, and lets you write it near the end of manufacturing rather than at the something close to the metal 1 layer.

▲fsbonetto 4 hours ago | parent | prev | next [-]

They are more like a way to proving the architecture of the accelerator before committing 100's of millions into a custom ASIC with TSMC

▲ 4 hours ago | parent | prev | next [-]
[deleted]
▲LoganDark 4 hours ago | parent | prev [-]

1. No

2. They don't have enough capacity either

The current largest FPGA, the AMD Versal Premium VP1902 has 18.5 million logic cells. That's not even enough for the smallest whisper.cpp model (75M).

You'd have to order hundreds of thousands of them (or millions) to serve even a single copy of a frontier model, and at that scale inference quickly becomes starved by the speed of light.

▲CamperBob2 3 hours ago | parent [-]

Well, you'd use BRAM to store model weights, not fabric. But still, you only get a couple hundred MB for probably close to US $100k per chip.

It's likely that the major FPGA vendors will soon announce parts specifically architected to support LLMs and similar models. But the current generation isn't suitable for that at all.

▲LoganDark an hour ago | parent [-]

Would BRAM even have enough bandwidth? The reason I quoted logic cells is because that's the way to get instant throughput, which is practically the only reason to use an FPGA over something like a TPU.

▲CamperBob2 an hour ago | parent [-]

I think so, because memory bandwidth really comes from bus width more than clock speed. You can construct 36-bit wide BRAM arrays with bus width comparable to HBM, just by specifying multiple BRAM arrays in parallel.

Never tried anything like that, though.