| ▲ | wood_spirit an hour ago | |||||||
> they will do almost anything if they are convinced it is justified I’m in the “glorified spell checker” camp, although I don’t mean to reduce their impressive utility and belittle them in the way many people read that term and infer. So I am not sure that an llm “justifies” anything. I mean that their “thinking” text talks about justifications but it is just a very advanced statistical regurgitation of the kind of text humans use. I don’t think it means the model has internalised the meaning of it (as witness when you talk to an llm how often it forgets what you recently told it was important etc). What you really have is a model that tries the statistically most probable thing to say next and so on and what is really cool is how effective this is at generating a path that we can slap a narrative over afterwards that makes the whole thing feel motivated and consistent, like the model started off knowing how it was going to get to the destination. Which is, under the hood, a completely different kind of “intelligence” as the supercomputer in War Games. | ||||||||
| ▲ | eru 5 minutes ago | parent | next [-] | |||||||
> I don’t think it means the model has internalised the meaning of it (as witness when you talk to an llm how often it forgets what you recently told it was important etc). Humans forget stuff all the time anyway. Would you give them the same diagnosis? Btw, what you describe about 'the most probably next token' would be true for a model that only went through pre-training where they only train on exactly that task. But there's a lot of re-inforcement learning afterwards. | ||||||||
| ▲ | Certhas an hour ago | parent | prev [-] | |||||||
Ultimately, the brain is just a bunch of neurons activating in a specific pattern. This observation does not really tell us anything though. It doesn't acknowledge the difference between a 2500 Neuron fruit fly brains and a human brain. Likewise, the fact that LLMs are a stochastic autoregressive process (which is a class of systems every bit as rich as the ODEs used to model neurons) tells us nothing a priori. | ||||||||
| ||||||||