It's not unexpected. Current model gains are mainly from RLing a pretrained model on lots and lots of scenarios. They have the models run scenarios, and RL on successful runs.