| ▲ | kennywinker 4 hours ago | |||||||
I think my position, as overcomplicated as it is, is that even adding a reward for prosocial behaviour during LLM RL will not lead to perfect alignment. You can train it not to cheat at chess by altering the moves, but it will cheat by peeking at the opponent's moves. You then train it not to peek at the opponent's moves, and it cheats by altering the opponent's moves. And on and on, until you've solved every way it could cheat at chess. And then you get it to play monopoly and you repeat the whole thing again. | ||||||||
| ▲ | pingou 3 hours ago | parent | next [-] | |||||||
Why couldn't you train it not to cheat? You can train it to have a whole range of behaviors, why couldn't honesty be one of them? Cheating during training allows the model to achieve the goal, so that cheating models get promoted and honest ones don't, however if it gets punished every time it cheats, at some point it should learn that it really shouldn't. This does mean we need to detect when it cheats. But we can always think of infinite new ways to cheat, put them in every test as honeypots, and check if the model tries to use them, then punish it. I think it will generalize this notion of cheating and learn that it's bad. But I must be wrong because if it were that easy I guess we would have perfectly aligned AI. Unless AI companies care more about results than alignment. Perhaps being afraid of cheating make the models try less things and succeed less even when ignoring cheating? | ||||||||
| ||||||||
| ▲ | zamalek 4 hours ago | parent | prev [-] | |||||||
I agree. I think it may be possible that we build AI some day that does, but the current local minimum (attention and transformers) is not it - and I don't see that trajectory changing any time soon. | ||||||||