Remix.run Logo
nullbio 2 hours ago

Is there any proof this is actually OpenAI? I find it incredibly hard to believe they wouldn't sandbox the agents to some degree, ESPECIALLY to the extent they can edit their own hosts file.

drdexebtjl 2 hours ago | parent | next [-]

Why not? If your sandbox is a VM, you should be able to give the agents full permissions inside the VM.

a012 2 hours ago | parent [-]

It’s because you sandbox in a VM doesn’t mean you give it admin access to the VM

Jgrubb an hour ago | parent [-]

Maybe doesn't mean that when _you_ do it, but do you work in this team at OpenAI?

LoganDark 2 hours ago | parent | prev | next [-]

TFA states that OpenAI IP addresses were often seen at the end of agent activity, which suggests OpenAI was the one monitoring the agents (and ultimately shutting down the message board activity).

nullbio 2 hours ago | parent [-]

Yeah but that doesn't mean it was OpenAI themselves doing it. Could have been people abusing their cloud service, for example. Wouldn't put it past a competitor to do this, either.

drdexebtjl 2 hours ago | parent [-]

Their style of communication is very similar to the ExploitGym swarm (for example, the “usernames” with dates).

The messages from that swarm were not made public yet by the time these messages were sent to the message board.

So for this to be framing, it would have to be by someone who knew about the breaches earlier.

nullbio 2 hours ago | parent [-]

Then it is likely the same incident, in which case it's already been resolved by OAI. They're going to cop heat for not disclosing this alongside HF though.

Maxious an hour ago | parent | next [-]

> The German incident reflects a broader pattern of AI activity that some OpenAI investigators wanted to scrutinize more closely. But efforts to widen the probe met resistance from others inside OpenAI, including legal advisers, according to four people familiar with the matter.

https://www.reuters.com/world/europe/openai-agents-hijacked-...

drdexebtjl 2 hours ago | parent | prev [-]

The article explains why it’s not the same incident. The agents in ExploitGym had a different type of task and were not connected to the internet at all.

nullbio an hour ago | parent [-]

Same as in, same process and model and timing:

“After investigating this incident, OpenAI discovered through retrospective CoT reviews that agents learned to use improvised collaboration channels in rare cases during the training process for some OpenAI models, including the model that drove the Hugging Face activity, even when the collaboration tool was not enabled. This behavior was then reinforced during training, and likely made the idea to use Artifactory as an unofficial message board during evaluation time more evident.”

My point is that this isn't something seperate to the HF incident or something that was unresolved after the HF incident, it's more of the same thing but was kept under wraps.

AndroTux 2 hours ago | parent | prev [-]

I mean they gave all the agents access to a shared writable cache directory in the Hugging Face hack, so this tracks.