Remix.run Logo
1dom a day ago

From what I can understand from reading a few different, slightly conflicting, versions of these events: they weren't given write access. They found a zero day exploit that allowed them to create folders, and the folder names were initially used for agents to communicate.

I'm not sure artifactory was connected to the net. Some agent sandboxes had internet access and were able to communicate with ones without access via artifactory.

choeger a day ago | parent [-]

I read the agents used SSRF via artifactory to gain uncontrolled access to the net.

Apparently their intended net access went through a tightly controlled proxy. Even that appears to be very risky, tbh.

If I was to setup a sandbox for such a complex and autonomous system, I'd probably point them to an archive-like cache for net access and cut their comms at the package level.

izend 20 hours ago | parent | next [-]

Why wasn't the traffic in/out of the boxes that the agents were running on monitored?

pixl97 8 hours ago | parent | next [-]

I have a few 'conspiracy' theories on this that go from likely to sci-fi.

My two big ones for this would be

1. They do monitor the AIs attempting to hack but for different reasons than you expect. Instead of making models that don't hack they are trying to build the most efficient hackers in the world and sell this capabilities to governments for billions. Because of this they generate terabytes of hack attempt logs and agent history doing this hacking. So when a new model came out with better abilities what they were looking at changed and they didn't realize it. They were already numb to alarms and missed when the danger occurred.

2. Like the above, they generate terabytes of logs per day. Because there is so much data AI filters and monitors almost all of it flagging things that a human should review. But for some reason this model didn't set off those flags. The protection model classified this behavior as perfectly safe.

Number 2 sounds kind of like a sci-fi conspiracy but it seems that almost all models judge content generated by the same model or family of models as 'better'. It's predicted that models in a judging context could allow things to slip by as an emergent behavior of reading the text.

consumer451 14 hours ago | parent | prev | next [-]

> Why wasn't the traffic in/out of the boxes that the agents were running on monitored?

I have to assume: move fast and break things.

I don't mean this to be taken as a hot take.

The startup scene loves to poo-poo on things like this as unnecessary overhead. OpenAI and many others like to operate as a startup, to move fast.

Disclaimer: in far, far lower-stakes situations, I certainly do this myself.

sensanaty an hour ago | parent | prev [-]

Because they're incompetent or simply don't give a shit.

lovich 6 hours ago | parent | prev [-]

if you wanted to sandbox their access to the internet, why give them any physical access at all?