| ▲ | danielmarkbruce 3 days ago |
| Nathan Lambert wrote a good book recently, and he and his team wrote the paper below about Tulu 3 (Allen Institute). Both are good reads. https://arxiv.org/pdf/2411.15124 |
|
| ▲ | doc_ick 3 days ago | parent [-] |
| Thank you for providing an arxiv! An aside, I finally do appreciate single column format now, makes it easier to convert to epub. |
| |
| ▲ | danielmarkbruce 3 days ago | parent [-] | | When you are done with the section on RLVR, consider whether the model is predicting tokens, or making moves. There is a reason the word "policy" is used in RL. | | |
| ▲ | doc_ick 2 days ago | parent [-] | | Would still say it’s a token predictor, a fancy one though. I suppose we can agree to disagree. |
|
|