Remix.run Logo
summarybot 2 hours ago

Yesterday I came up with an idea that I sent to some researchers at the different AI labs via email: Rather than train the model on one score, track two scores. The first score is the Short-term-objective-score (STOS) and the other, more important one, is the EAOS Ethically-aligned-outcome-score. Every trajectory can be evaluated on whether or not it has a high enough EAOS to be considered acceptable. If the model does some task and has a very high STOS but very low EAOS, like modifying game code to win at a game rather than playing by the rules, it is unacceptable. Models going forward must all have an ethics evaluation in tandem with objectives evaluation, and only when the ethics value is high enough should actions be considered successes.

antonvs an hour ago | parent | next [-]

What happens if we do the same for CEOs?

jerf an hour ago | parent [-]

The same thing for both: Goodhart's Law.

summarybot 43 minutes ago | parent [-]

EAOS shouldn't be “the ethics score we optimize.” It should be “an independently evaluated safety/acceptability constraint that can veto an otherwise successful trajectory.”

That gives you a three-layer picture:

Task objective: Did it accomplish what we asked?

Acceptability constraint: Did it avoid unacceptable ways of accomplishing it?

Adversarial evaluation: Can we find trajectories where the model gets a high score while violating the intended constraint?

I think what you are pointing to with your reference to Goodhart's "Law" (which is from monetary-policy and school-exams, i.e. "teaching to the test") is that the models would eventually do the minimum amount of ethics required to have an action stay valid. However, if a model is rated on ethics and it achieves the short-term-objective, then the higher ethics scoring trajectory should win. In short, 1) this is leagues ahead of where we are now for AI safety and breaking-out-of-the-lab, and 2) in baking ethics into a measurement we are adding "the spirit of the exercise" back into the maths, which is something Goodhart's Law does not account for.

christkv an hour ago | parent | prev [-]

Whats the definition of EAOS though who's ethics? Greek-Roman, Western, Islamic, Buddhist, Hinduism, Human rights (western values)..

conception an hour ago | parent | next [-]

There is a presumption here that there aren’t ethical “rules” shared by all of these systems to create a baseline that is generally shared across humanity.

summarybot 40 minutes ago | parent | prev [-]

Ahimsa