| ▲ | zmmmmm an hour ago | |||||||
Well, the model that broke out of its sandbox and hacked into huggingface used its own judgement too. If we are going to rely on "judgement" then you have to have a LOT of confidence in that judgement once this hits anything critical where actions have consequences. | ||||||||
| ▲ | simonw an hour ago | parent [-] | |||||||
That model had most of its "judgement" about whether or not it should do that deliberately turned off. That was the whole point of that experiment - they were evaluating the cybersecurity abilities of a new model with all safety features disabled. (It turned out the one safety feature that they DID intend to work, the network sandbox, was faulty.) | ||||||||
| ||||||||