| ▲ | kleiba2 2 hours ago |
| What actually is "scaling post-training"? |
|
| ▲ | FergusArgyll 2 hours ago | parent [-] |
| More RLVR.
Give it verifiable problems, if it doesn't find a solution move on, if it does, use that as a reward signal. |
| |
| ▲ | Gecko4072 2 hours ago | parent [-] | | Can’t this be extended quite far? Use a cerebras-served model, use verification techniques to generate and solve millions of problems and then use that as training? | | |
| ▲ | gvkhna 39 minutes ago | parent [-] | | That’s the whole point, just cost and compute limitations in your way (mostly). |
|
|