Remix.run Logo
Phemist 6 hours ago

The 20W number includes EVERYTHING else the brain does. The chips/models are literally only producing tokens. Let's see an LLM drive a robot harness and have the robot produce speech, as well as move through 3D space, keep track of metabolic needs, etc. etc. etc. before we compare efficiencies. That is even assuming the tokens are of equal quality. This comparison is currently Apples and Oranges.

phoghed 6 hours ago | parent | next [-]

kind of a moot point if you can't get your brain to not do everything else. I think it's a fun comparison, even if it's not a 100% equivalence.

cmrdporcupine 4 hours ago | parent | prev | next [-]

Right, I can do the talked about ~3 tok/sec output and drive a car, hold my bladder, and eat chips at the same time.

Take that, Jalapeno!

falcor84 2 hours ago | parent [-]

For what it's worth, LLMs don't really suffer from incontinence, so at least that part is pretty much a solved problem.

pantalaimon an hour ago | parent | next [-]

They sometimes leak their system prompt

undersuit 9 minutes ago | parent | prev [-]

So why are their water cooling systems filled with leak detectors? /s

CooCooCaCha 5 hours ago | parent | prev [-]

And the brain is literally only producing electrochemical signals.

I don’t see how tokens can’t produce speech or track metabolic needs. You can talk to chatgpt can’t you? Or do you mean literally talking? Because that’s not a brain function, that’s the mouth, vocal chords, and lungs.

Phemist 5 hours ago | parent [-]

> I don’t see how tokens can’t produce speech or track metabolic needs.

It probably could, but the point is this would require additional tokens, blowing up the comparison. The token output of LLMs and "token output" of speech are simply at different abstraction levels. Hence my comparison to the LLM brain driving the robot harness to produce speech etc. This would be more comparable, and also look significantly worse than "only" the 22x less efficient number.