Remix.run Logo
threecheese 2 hours ago

An amazing human reverse-engineer - who also plays online chess - has judgement which uses a moral compass to not decide to hack the chess tournament. This judgement has been trained through the experiences of that person, with a through-line of that compass - a coherent mental model of the world which evolves but is hopefully pinned to some set of principles it shares with society.

This chess judgement is completely irrelevant when the human is tasked with finding software weaknesses, and only the compass gates that.

Can a model trained on the totality of all person-experiences (as expressed in written knowledge) ever maintain a coherent through-line of alignment? It has all morals in the dataset, and only some RL to try and minimize or maximize known behaviors via weights - experience all the good things and the bad things, then optimize for some good things the trainers identified.

It's like the reverse of what a person goes through. Morality by subtraction. How can it ever work?

zzril 2 hours ago | parent | next [-]

I think the fundamental difference is that humans aren't trained on experiences. They make experiences. Models are just thrown away and re-created after each conversation / job.

If you could clone and throw away human workers as you need them, a lot of the morale would disappear.

kansface 2 hours ago | parent | prev [-]

Yes, why not? All existed models have been rewarded for cheating (extensively). That is us, putting intense evolutionary pressure, on a system to produce a result we don’t want through indifference. Why can’t we post train them not doing that?