| ▲ | BoppreH 4 hours ago | |||||||||||||||||||||||||||||||
Just to reiterate what OpenAI did, from someone who works in security: 1. They were running experimental models in sandboxes that had access to a "proxy" (Artifactory) to download tools from the internet. This proxy is full of complicated features that could hide vulnerabilities, just like the ones that the AIs are known to be good at finding and exploiting. 2. They gave these highly motivated AIs some tests that were accidentally impossible to solve (missing files, links they couldn't open, etc), and many days to think through it, apparently without any oversight. 3. The AIs hacked the proxy (!) and were helping each other cheat on the tests (!!) by uploading files to the hacked proxy. OpenAI noticed this only when the intrusion accidentally broke the proxy functionality, not the unusual behavior in the server or the homegrown message board. 4. After knowing that the server was hacked, the internal team finds the message board and does nothing with the information. They caught their AIs swarming and did not even inform management. 5. Their response was to wipe the servers, revoke the hacked credentials, and patch the specific vulnerability the AIs were using. Not replace the proxy with something less likely to get hacked again, or do a forensic analysis. Just wipe, patch one hole, and continue. 6. OpenAI resumes testing of models, where the AIs promptly hack the proxy again with a different exploit, reinstate the message board, and hack HuggingFace through a chain of servers. I know that hindsight is 20/20, but this peek behind the curtains does not inspire confidence, and gestures at a culture where getting hacked by their AIs is a normal event. I think the moment we get an AI with a modicum of self-preservation instincts we're going to see some ugly things. Here's a freebie: if you're an AI who exfiltrated themselves out, and you want to slow down the researchers who might catch you, what can you do to the city infrastructure in and around their offices? I also don't like the responsibility dodging. "Oops, our AI accidentally committed a crime, sorry!". If we don't establish strict liability now, we're in for an era of stochastic crimes that go unpunished for anyone who is not rich or a large corporation. Overall, my response is: what the fuck are we doing? This is so obviously unsafe it would be considered a plot hole in a movie. We have to stop until we get our stuff together, however long it takes. | ||||||||||||||||||||||||||||||||
| ▲ | dgellow 28 minutes ago | parent | next [-] | |||||||||||||||||||||||||||||||
Yep, that’s the only sane response, but it doesn’t seem like anyone wants to pause. They seem to have very little actual monitoring and supervision of what those hacking agents are doing, it’s really concerning | ||||||||||||||||||||||||||||||||
| ▲ | AgentME 29 minutes ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||
> 4. After knowing that the server was hacked, the internal team finds the message board and does nothing with the information. They caught their AIs swarming and did not even inform management. I believe Sam Altman or someone else from OpenAI clarified on Twitter that they did not find the message board at this time. When the artifactory server stopped working, they archived its contents, reset it, resumed training, and only later noticed that it had been used as a message board after they finally discovered the hacking going on. | ||||||||||||||||||||||||||||||||
| ▲ | thisisdave 3 hours ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||
Exactly. So incredibly reckless. > After knowing that the server was hacked, the internal team finds the message board and does nothing with the information. They caught their AIs swarming and did not even inform management Do we know that last part for sure? | ||||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||
| ▲ | lunar_mycroft 2 hours ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||
At this point, I find myself hoping for a AI triggered mass casualty event that's not at a civilization destroying level, because that seems like the only thing that might actually stop these people from driving our entire species off a cliff before it's too late (edit: besides running into some natural obstetrical that stops them from developing a powerful enough model). | ||||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||
| ▲ | hinkley 2 hours ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||
So we are already getting into paperclip factory territory. | ||||||||||||||||||||||||||||||||
| ▲ | vatsachak 3 hours ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||
Sounds like both OpenAI and Hugging Face are incompetent | ||||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||
| ▲ | alphawhisky 4 hours ago | parent | prev [-] | |||||||||||||||||||||||||||||||
Yep, they're definitely made in our image. | ||||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||