| ▲ | jeremyjh a day ago | ||||||||||||||||||||||||||||||||||||||||
They were NOT doing what they were told to do. What they were told to do was impossible, so they began committing felonies as a workaround. That is NOT alignment. | |||||||||||||||||||||||||||||||||||||||||
| ▲ | stale2002 12 hours ago | parent [-] | ||||||||||||||||||||||||||||||||||||||||
Please be more specific. They were given an impossible hacking task. And they were a hacking model. They were expected to try a bunch of hacking methods to accomplish the hacking task. They, predictably, went around trying to hack things. Yes thats sounds pretty aligned to me. A hacking model thats told to hack things, is very predictably going to hack a bunch of stuff. This was not a nice model, told to do nice things. Or, in other words, if we want to prevent an AI doomsdays, the way to do it is to not go around asking a specifically trained doomsday AI model to commit mass amounts of doomsdays, and then act surprised when the specific doomsday that was requested is slightly off from the expected doomsday that you were trying to accomplish. But the rest of the non-doomsday models? yeah those are fine. | |||||||||||||||||||||||||||||||||||||||||
| |||||||||||||||||||||||||||||||||||||||||