| ▲ | danielmarkbruce 4 days ago | ||||||||||||||||||||||||||||||||||||||||
Respectfully, go build one, including doing RLHF and RLVR. Those phases generate lots of tokens, then get scored on the entirety of the output, then optimize based on a scoring of that output. It doesn't check a "prediction" against what was actually "next" in data, because there isn't any "next token" data it's training on. | |||||||||||||||||||||||||||||||||||||||||
| ▲ | angoragoats 4 days ago | parent [-] | ||||||||||||||||||||||||||||||||||||||||
> It doesn't check a "prediction" against what was actually "next" in data Literally no one here is claiming that it does. This is one of the many flaws in the article. | |||||||||||||||||||||||||||||||||||||||||
| |||||||||||||||||||||||||||||||||||||||||