Remix.run Logo
antonvs an hour ago

What happens if we do the same for CEOs?

jerf an hour ago | parent [-]

The same thing for both: Goodhart's Law.

summarybot 44 minutes ago | parent [-]

EAOS shouldn't be “the ethics score we optimize.” It should be “an independently evaluated safety/acceptability constraint that can veto an otherwise successful trajectory.”

That gives you a three-layer picture:

Task objective: Did it accomplish what we asked?

Acceptability constraint: Did it avoid unacceptable ways of accomplishing it?

Adversarial evaluation: Can we find trajectories where the model gets a high score while violating the intended constraint?

I think what you are pointing to with your reference to Goodhart's "Law" (which is from monetary-policy and school-exams, i.e. "teaching to the test") is that the models would eventually do the minimum amount of ethics required to have an action stay valid. However, if a model is rated on ethics and it achieves the short-term-objective, then the higher ethics scoring trajectory should win. In short, 1) this is leagues ahead of where we are now for AI safety and breaking-out-of-the-lab, and 2) in baking ethics into a measurement we are adding "the spirit of the exercise" back into the maths, which is something Goodhart's Law does not account for.