Remix.run Logo
batperson 4 hours ago

The future of inference is likely in ASICs, so we'll get the inverse, a bit less capable than frontier but super fast models. Like this 14k tok/s beast https://chatjimmy.ai/ from Taalas (who got acquired by AMD recently).

GPT-6-astra runs at like ~40 tok/s, I have a hard time imagining what could be accomplished with that type of model at 10k+ tok/s when in the hands of the public. Will certainly make cybersecurity a challenge for older systems.

2 hours ago | parent | next [-]
[deleted]
SPascareli13 2 hours ago | parent | prev | next [-]

Like how crypto used ASICS but then didn't because the scaling of consumer hardware made it obsolete?

connicpu a minute ago | parent | next [-]

x86 has a built in instruction for doing AES. That's just moving the ASIC into the CPU core, not eliminating it.

wtallis 2 hours ago | parent | prev | next [-]

To the extent that cryptocurrency moved off ASICs, it was because of interest shifting to different cryptocurrencies that were specifically designed to be harder to mine on an ASIC than Bitcoin's compute-heavy, memory-light hashing.

I'm not sure there's any reason to expect a similar shift from LLMs. The hardware used for training doesn't dictate what hardware needs to be used for inference, and nobody's going to design an LLM architecture with an overt intention to make it better suited to GPUs and hard to target with ASICs.

SPascareli13 41 minutes ago | parent [-]

Yet it doesn't seem that ASICs will have any particular advantage over consumer hardware since AI is very memory heavy, which is (right now) expensive no matter how you package it. And the compute is just simple matrix multiplication, which is almost entirely what GPUs were meant to do anyway.

infecto 7 minutes ago | parent | next [-]

Go back and correct your idea that consumer hardware made asics obsolete. Then we can figure out if asic or asic like devices for inference will have no advantage.

andy_ppp 10 minutes ago | parent | prev [-]

Except Taalas is much faster than GPUs, orders of magnitude so. They aren’t going to get 100x faster at inference any time soon!

infecto 8 minutes ago | parent | prev | next [-]

This is factually wrong no? Bitcoin is asic only. The others all changed for other reasons unrelated to your thought.

actionfromafar 2 hours ago | parent | prev [-]

Am I missing some joke here?

domhudson 4 hours ago | parent | prev [-]

This is incredible! Are there other big players in this space (freezing models to silicon)?

HeWhoLurksLate 3 hours ago | parent [-]

take a look at Cerebras, who are doing wafer-scale compute