| ▲ | grebc 4 days ago |
| The blog post itself says one thing, but then demonstrates the exact thing they’re arguing against. If you can’t grasp that logic gap then there’s no point discussing further. |
|
| ▲ | garrinm 4 days ago | parent | next [-] |
| I try to make 3 claims in the post, it was a bit clumsy I'll admit that. 1. At inference time, LLMs emit one token at a time given the prior tokens. This looks like prediction and I concede that. 2. During pre-training, LLMs predict the next token and compare to the actual next token in the training data. This is the classic setting for ML predictions. And I think its meaningful, the model really is predicting what the ground truth next token will be in the data. 3. During post-training, in the case of RLVR, there is no ground truth next token. In pretraining, the question is "what token actually came next?". In RLVR, the question is "what sequence of actions gets a high reward?" And the whole point is that thinking about the RLVR is important. A mental model that stops at 1 or 2 is incomplete and doesn't capture what drives LLM tokens. |
| |
| ▲ | grebc 4 days ago | parent | next [-] | | My understanding about your third point is the LLM generates lots of different answers, then they’re ranked according to some computation the creators came up with. I’m still not sure what doesn’t qualify any of that as a prediction, and I’ll be more blunt: a guess. | | |
| ▲ | danielmarkbruce 4 days ago | parent [-] | | A guess at what though? One guesses at truths they don't know, or events that haven't happened yet. What is the model guessing? | | |
| |
| ▲ | danielmarkbruce 4 days ago | parent | prev | next [-] | | Probably the easiest way to describe an LLM that it's a policy. There is a reason that word has stuck in RL. And it's not just RLVR. RLHF has been going on for years and years. LLMs have not been "next token predictors" for probably 5-6 years. | |
| ▲ | what 3 days ago | parent | prev [-] | | It’s still just predicting the next token though just with a different reward between 2 and 3. |
|
|
| ▲ | danielmarkbruce 4 days ago | parent | prev [-] |
| Nope, it doesn't. No logic required, you can just build an LLM yourself, including post training. You'll see that predicting the next token isn't something the model does or is optimized for in RLHF or RLVR. You can hand wave all you like, but you have never done it. |
| |
| ▲ | grebc 4 days ago | parent [-] | | Yes, no logic is necessary for LLM adherents we're all finding out. Carry on good soldier. | | |
| ▲ | danielmarkbruce 4 days ago | parent [-] | | If you haven't built one, and don't understand how they work, why comment? | | |
| ▲ | grebc 4 days ago | parent [-] | | You don't need to build a car to understand one. That you tie yourself up in knots of fancy acronyms instead of plain words and that your argument boils down to semantics of the word prediction, it's pretty clear what is up brother. | | |
| ▲ | danielmarkbruce 4 days ago | parent [-] | | Lol, sure, just read a blog post and you'll understand how a car works....It's very simple.... | | |
| ▲ | doc_ick 4 days ago | parent | next [-] | | Just like how reading a math book doesn’t teach you math, why do they make us read anyway? (Sarcasm) if reading a blog post didn’t teach someone how a car works how come it “can” work for next token predictors | | |
| ▲ | danielmarkbruce 3 days ago | parent [-] | | Mine was sarcasm. People who actually understand cars have built them. Until you build something, you don't understand it. | | |
| ▲ | grebc 3 days ago | parent [-] | | Now you’re claiming people don’t understand unless they build something. Boy, oh boy, do you keep digging your logic hole that much deeper. As mentioned earlier, Sam thanks you for your obfuscation efforts while his equity keeps going up. The swindle continues. | | |
| ▲ | danielmarkbruce 3 days ago | parent [-] | | I'm not the one hiding behind a fake name. If you want to understand how this stuff works, there are totally decent books about building them from scratch. It's not that hard, and you'll likely find it interesting. Sebastian Raschka and Nathan Lambert have good books out, and the Allen Institute has available all the code and data they have used for several projects. | | |
| ▲ | grebc 3 days ago | parent [-] | | Now a fake name accusation is thrown by someone with three first names. Keep digging that hole, I’m sure you’ll surface somewhere with some sunshine. |
|
|
|
| |
| ▲ | _superposition_ 2 days ago | parent | prev [-] | | Lol isn't this what llms do?
Did you not just undermine your entire argument? |
|
|
|
|
|