| ▲ | FergusArgyll 2 hours ago | |||||||
More RLVR. Give it verifiable problems, if it doesn't find a solution move on, if it does, use that as a reward signal. | ||||||||
| ▲ | Gecko4072 2 hours ago | parent [-] | |||||||
Can’t this be extended quite far? Use a cerebras-served model, use verification techniques to generate and solve millions of problems and then use that as training? | ||||||||
| ||||||||