| ▲ | danielmarkbruce 4 days ago |
| The word "predict" has a meaning. I don't "predict" my next move in chess. I might predict what someone elses first move is. |
|
| ▲ | stanleykm 4 days ago | parent | next [-] |
| In any case this is all very pedantic. In the process of selecting a move to make there is a prediction. Whether that prediction is the opponent’s next move or what your next move should be based on the game’s existing state, there is a prediction that the next move you make will improve your chance to win. Maybe the probability in that selection is 100%. You have no other possible move. It doesn’t matter. All we are doing here as far as I can tell is arguing over where the prediction happens and whether that counts as predicting something. |
| |
| ▲ | danielmarkbruce 4 days ago | parent [-] | | There is no truth for RLHF or RLVR. You can't predict against something if you can't check against the truth. It's not pedantry. The objective function changes. The optimization changes. THese are real things when training a model, not hand wavy philosophical ideas. |
|
|
| ▲ | ordersofmag 4 days ago | parent | prev | next [-] |
| The LLM does not determine the next token. It generate odds for all of the tokens it knows as to their likelihood of being 'next'. It's up to the harness running the LLM (and in most cases the a temperature setting) to actually decide on a particular next token. I think it's more accurate to call the thing the LLM actually generates (an ensemble of probabilities) a 'prediction'. It might be accurate to say the harness decides on the next token based on the prediction from the LLM. The role of the LLM is much more akin to predicting your opponents move than deciding your own. |
| |
| ▲ | danielmarkbruce 4 days ago | parent [-] | | Respectfully, go build one, including doing RLHF and RLVR. Those phases generate lots of tokens, then get scored on the entirety of the output, then optimize based on a scoring of that output. It doesn't check a "prediction" against what was actually "next" in data, because there isn't any "next token" data it's training on. | | |
| ▲ | angoragoats 4 days ago | parent [-] | | > It doesn't check a "prediction" against what was actually "next" in data Literally no one here is claiming that it does. This is one of the many flaws in the article. | | |
| ▲ | garrinm 4 days ago | parent [-] | | It does in pre training, but not in RL post training. And not at inference time. Reading over all these comments I get the feeling my mistake was not clearly delineating inference time and train time. | | |
| ▲ | danielmarkbruce 4 days ago | parent | next [-] | | Your mistake was assuming people would be bothered to understand the details of how things work. Most people are lazy and don't know the details of how anything works. | |
| ▲ | angoragoats 3 days ago | parent | prev [-] | | My point is that “it doesn’t check the accuracy of the prediction against the data” is a non-response, because no one calling it a “next-token predictor” is making the claim that it does do that or that they’re calling it a next-token predictor because it does that. | | |
| ▲ | danielmarkbruce 3 days ago | parent [-] | | Many people are in fact claiming the thing you are saying they are not - even if you are not. The reason they are claiming it is that it was true at one point, and most intro courses/blog posts/videos still describe them that way and then hand wave some "other stuff at the end". You can even see a comment here that refers to the gpt-2 paper. LLMs were trained to predict the next token, produced a distribution to do so, were scored against their prediction v the truth, and the weights updated so that the probability distribution made it more likely to predict the truth from that sample next time. They were, in every sense of the word, a next token predictor. They are no longer that thing due to post training. They simply aren't making a prediction, and they aren't even optimized for the next token. If I give a distribution of the heights of the population, I'm not giving a prediction either. Distributions don't imply predictions. Why the desperation to hang onto the word "prediction"? | | |
| ▲ | angoragoats 3 days ago | parent [-] | | > Many people are in fact claiming the thing you are saying they are not - even if you are not. My original comment said “no one here.” Please show me where someone in the comments here is claiming that. > Why the desperation to hang onto the word "prediction"? No desperation here. It’s just a word that conveniently describes (especially to laypeople) what’s going on, even if it may not be the most mathematically correct or rigorous word to describe what’s going on. I think you’re being needlessly pedantic. Why the desperation to refute it? |
|
|
|
|
|
|
|
| ▲ | 4 days ago | parent | prev [-] |
| [deleted] |