> Deepseek proposes RLVR as a way to get around the lack of $ they have to produce human reasoning trace data.
What was the difference between what deepseek did for R1 and what OpenAI did for o1?