It's implicitly trained against. There is like information leakage with researchers messing with the training parameters and checkpoints used.
It's not the direct feedback loop of RL but its not far.