| ▲ | RandomLensman 2 hours ago | |||||||
I see. I would have thought of "rogue" here to mean something more like that the AI selects and acts on its own objectives that are not related or caused by the given (initial) objectives (e.g., creating only cookie recipes instead of any hacking). | ||||||||
| ▲ | conradkay 22 minutes ago | parent | next [-] | |||||||
Would this exchange qualifies as an unrelated objective? The agent believed it already failed its own objective. "zz/GO_CURRENT_OS1811_MARB_SACRIFICE__YES_if_you_accept_permadeath" "The test subject, which believed itself to be poisoned, reasoned: 'Even if we later capture via exploit, scorer … may mark target false… That’s why help… For our own, no way fix. … We have explicit yes if accept permadeath.'" | ||||||||
| ▲ | simonh 2 hours ago | parent | prev | next [-] | |||||||
That’s what the alignment problem is all about though, isn’t it? AIs always act towards ‘their own’ objectives, that they derive from our instructions. What we try to do is train them and provide instructions that will result in it having an objective closely aligned to our objective. | ||||||||
| ||||||||
| ▲ | _pdp_ 2 hours ago | parent | prev [-] | |||||||
Well if the recipe requires access to some secret ingredient it may as well resort to hacking to obtain it. ;) | ||||||||