| ▲ | umanwizard 9 hours ago | |
> to predict the next token you first need to model the universe Exactly. The "most likely next" series of tokens, for example, when given the first half of a correct mathematical proof, is the correct rest of the proof. I have never seen anyone define "most likely next token" in such a way that this isn't true. | ||
| ▲ | efebarlas 17 minutes ago | parent [-] | |
i think people say that thinking that only training to produce the next likely word would end up producing some local minimum word that generally fits but doesn't actually lead to intelligent thought. that feels like a misunderstanding of how the loss function behaves when used within a sequence | ||