| ▲ | jerf an hour ago | |
The same thing for both: Goodhart's Law. | ||
| ▲ | summarybot 44 minutes ago | parent [-] | |
EAOS shouldn't be “the ethics score we optimize.” It should be “an independently evaluated safety/acceptability constraint that can veto an otherwise successful trajectory.” That gives you a three-layer picture: Task objective: Did it accomplish what we asked? Acceptability constraint: Did it avoid unacceptable ways of accomplishing it? Adversarial evaluation: Can we find trajectories where the model gets a high score while violating the intended constraint? I think what you are pointing to with your reference to Goodhart's "Law" (which is from monetary-policy and school-exams, i.e. "teaching to the test") is that the models would eventually do the minimum amount of ethics required to have an action stay valid. However, if a model is rated on ethics and it achieves the short-term-objective, then the higher ethics scoring trajectory should win. In short, 1) this is leagues ahead of where we are now for AI safety and breaking-out-of-the-lab, and 2) in baking ethics into a measurement we are adding "the spirit of the exercise" back into the maths, which is something Goodhart's Law does not account for. | ||