| ▲ | stale2002 11 hours ago | |||||||||||||||||||||||||
Please be more specific. They were given an impossible hacking task. And they were a hacking model. They were expected to try a bunch of hacking methods to accomplish the hacking task. They, predictably, went around trying to hack things. Yes thats sounds pretty aligned to me. A hacking model thats told to hack things, is very predictably going to hack a bunch of stuff. This was not a nice model, told to do nice things. Or, in other words, if we want to prevent an AI doomsdays, the way to do it is to not go around asking a specifically trained doomsday AI model to commit mass amounts of doomsdays, and then act surprised when the specific doomsday that was requested is slightly off from the expected doomsday that you were trying to accomplish. But the rest of the non-doomsday models? yeah those are fine. | ||||||||||||||||||||||||||
| ▲ | jeremyjh 7 hours ago | parent | next [-] | |||||||||||||||||||||||||
> But the rest of the non-doomsday models? yeah those are fine. What is your evidence for this? There is a lot of research that says otherwise. They cheat when they can. They behave differently when they believe they are being observed. Their CoT is different when they believe it is being evaluated. They are aligned to what best satisfies their reward function, not to our INTENDED VALUES for them. | ||||||||||||||||||||||||||
| ||||||||||||||||||||||||||
| ▲ | 7 hours ago | parent | prev [-] | |||||||||||||||||||||||||
| [deleted] | ||||||||||||||||||||||||||