They were RL trained on verifiable rewards. It's not purely learning to predict the next token of a human produced stream.