"does that token in that spot increase or decrease the likelihood of solving this long range programming task?"
It's not correct or not, it's a gradient based on the reward signal.