Remix.run Logo
harrouet 3 hours ago

I could definitely image Apple embedding a kind of LLM-optimized FPGA: slow to load (update) an LLM, but blazing fast at computing tokens.

Who needs memory when your model is set in silicon ?

KeplerBoy an hour ago | parent [-]

You don't an FPGA if you're taping out your own chips. But that is just a MMA accelerator with decent memory bandwidth. No secret sauce here.