Isn't that a challenge with RL anyway that for a lot of problems its hard to even know accuracy continuously for each step