Remix.run Logo
throwawayffffas 12 hours ago

> whoever burns their models to ASICs fastest.

There is already custom hardware see cerebras.

GPUs have a lot of slack there is at least one lab that had a (small 8b) model generate almost 3000 tokens per second on a MI300X for a talk, instead of the typical software stack that did maybe 100ish tokens per second.

High bandwidth flash storage is in the works, i.e hard drives with TBs of storage and over 1 TB per second of read speeds. Meaning that in a couple of years you may be able to buy a card with 40-90GBs of HBM and 4TB of HBF and run a 3T model locally at a reasonable speed for 10-20k as opposed to a cool mil.

topspin 7 hours ago | parent [-]

> Meaning that in a couple of years you may

There is no "may" here. You will see this.

It's always difficult to see it from the present, but we're not at some end stage in hardware development; we're still on the same curve our predecessors also couldn't see: they couldn't imagine that there would be high performance computers carried in our pockets, with staggering amounts of storage and compute, putting to shame the machines they filled rooms with.

throwawayffffas 3 hours ago | parent [-]

Oh yeah the may is on the 2 year time horizon. It could be 3 or 4. Or next year.