| ▲ | auggierose 3 hours ago | |||||||
Well, in order to do anything, it is good to know what you want to achieve. How do you align the reward signal? You align it so that you can differentiate between persistence and morality, because that is the goal. This is not something you should let the AI figure out by itself, because when it does, lying and cheating agents will be the result, just like humans have figured that out for themselves. This can be as simple as rewarding moral behaviour and penalising immoral behaviour in your training, but how is that interacting with persistence? Maybe a white lie is fine sometimes in order to achieve your goal? So, when designing your training, you will need to answer for yourself how persistence interacts with morality. That is not something you can outsource to machine learning. Or rather, you can, but then you get lying and cheating agents. | ||||||||
| ▲ | user43928 2 hours ago | parent [-] | |||||||
I think you need to find broken tasks in your training data and monitor for cheating during training, not answer any questions about how persistence interacts with morality. But that's just my guess. | ||||||||
| ||||||||