Remix.run Logo
▲ chris_money202 a day ago

This is the smallest unit of a typical AI ASIC, for example Google's TPU would have several dozen more compute units inside of it per chip.

In essence this is the simplest unit of an entire AI chip. The more complicated units of AI ASICS are actually the periphery, especially around PCIe and Ethernet and the sub-systems that link many AI ASICs together to move huge amounts of data around ultimately to each TPU.

So its missing ALOT

▲fsbonetto 8 hours ago | parent | next [-]

Yeah thats true, chip to chip communication and KV cache management is make or break for an AI ASIC.

▲AnimalMuppet a day ago | parent | prev [-]

Thanks. But that wasn't my question. For this part, how is the performance? State of the art? Better? Or worse?

▲fsbonetto 21 hours ago | parent | next [-]

It's at 80~90% the max perf it can achieve on this hardware... Which is DDR3 speeds. I'm working on getting FPGA's with DDR4 and HBM2 next.

▲chris_money202 8 hours ago | parent [-]

Why is there any point in trying with faster memory if you are only at 80% perf on DDR3? You'll just go to 40% perf on the next iteration?

You should only upgrade when the bottleneck becomes memory, right now the bottleneck is still on compute until you are at 100% max perf

▲chris_money202 a day ago | parent | prev [-]

Its a SYSTEM on Chip, evaluating 1 function on performance is superficial. This could have the best performance in the world, and it doesn't matter if the bottleneck is upstream