If you assume that noisy general usage data enables good RL, particularly compared to curated RL training sets.
I am not convinced that's the case.