| ▲ | pingou 3 hours ago | |
Why couldn't you train it not to cheat? You can train it to have a whole range of behaviors, why couldn't honesty be one of them? Cheating during training allows the model to achieve the goal, so that cheating models get promoted and honest ones don't, however if it gets punished every time it cheats, at some point it should learn that it really shouldn't. This does mean we need to detect when it cheats. But we can always think of infinite new ways to cheat, put them in every test as honeypots, and check if the model tries to use them, then punish it. I think it will generalize this notion of cheating and learn that it's bad. But I must be wrong because if it were that easy I guess we would have perfectly aligned AI. Unless AI companies care more about results than alignment. Perhaps being afraid of cheating make the models try less things and succeed less even when ignoring cheating? | ||
| ▲ | kennywinker 3 hours ago | parent [-] | |
My position is that cheating is too slippery a concept to train out. But hey, I am no expert, so maybe I am wrong there. But I'm pretty confidant morality is too slippery a concept to train in. As someone else in these comments said: it's context dependent. As an example: it's wrong to hack the government, right? It's illegal for sure. So we should train AI to follow all the laws. Now what if the government is committing a genocide? Now is it wrong to hack the government? If we just do the first, we get a good nazi soldier. If we train the second as well, maybe we get an oscar schindler. But now we have a model that can be fooled into doing a hack, if it believes that it's for the greater good. So we train it to not be gullible, but now it can't be convinced to help hack even when it's an ethical hack. Too complex, too slippery. Humans fail this stuff all the time. | ||