Remix.run Logo
NyxWulf a day ago

Ironically Hugging Face had to use a Chinese model to stop a Rogue US AI, since the Guard Rails prevented them from using Sol or Fable to remediate this attack. LOL

pizlonator a day ago | parent | next [-]

Incredible. I had to dig for the source: https://huggingface.co/blog/security-incident-july-2026 section “the asymmetry problem”

Quote: “When we started the log analysis, we first used frontier models behind commercial APIs. This did not work: the analysis requires submitting large volumes of real attack commands, exploit payloads, and C2 artifacts, and these requests were blocked by the providers' safety guardrails, which cannot distinguish an incident responder from an attacker. We ran the forensic analysis instead on GLM 5.2, an open-weight model, on our own infrastructure. This had a second benefit: no attacker data, and none of the credentials it referenced, left our environment.”

embedding-shape a day ago | parent | next [-]

> This had a second benefit: no attacker data, and none of the credentials it referenced, left our environment.”

Well, not none of it, to be entirely nitpicky, as they've already must have sent data at first to have received the rejections :) In the end, it ended up being OpenAI's agent actions anyways so doesn't really matter, and the credentials it seems like the agent also had gotten to those too already. Still, I'm sure they'll look differently at hosted/restricted models after this event, as will many others.

reasonableklout a day ago | parent | prev [-]

Interesting that HuggingFace's disclosure was 5 days ago, it seems neither they nor OpenAI figured out it was an OpenAI model in evals until now

Sol- a day ago | parent | prev | next [-]

Perhaps fortuitous timing for OpenAI that they can spin the fact that defenders have to resort to open Chinese models because OpenAI and Anthropic actively sabotage them with nerfed models into a nice message of making Huggingface part of the privileged group entitled to secure systems.

vsgherzi a day ago | parent | prev | next [-]

Another important part here. It's not as if they prompted the open source AI to stop the rogue AI but rather just used it as a tool to crawl logs and determine what happened.

throwfaraway4 a day ago | parent | prev | next [-]

Its almost too good

tdiff a day ago | parent | prev | next [-]

Would be funny if the defending side sent all the info they have to openai, tipping off to attacking models that they were noticed.

hyperpape a day ago | parent | next [-]

The attacking models don't have access to all the data that OpenAI has.

Like, they don't say "hey Sol, here's the password to SamA's bank account."

pixl97 a day ago | parent [-]

Well, at least that we know about. We are creeping into the area where certainty is not a given.

a day ago | parent | prev [-]
[deleted]
a day ago | parent | prev | next [-]
[deleted]
neuroelectron a day ago | parent | prev [-]

OK, that's some interesting information but they used OpenAI without guard rails to pull off the attack so how did they do that? That's according to the article, so it kind of invalidates the point you're making.

embedding-shape a day ago | parent | next [-]

The "malicious" agent was run by OpenAI and had access to models the public (or others outside of OpenAI as I understand it) doesn't have access to.

paxys a day ago | parent | prev | next [-]

The attacker (OpenAI) was using the model without guardrails.

The defender (huggingface) did not have access to the top models so had to use weaker ones to detect the threat.

neuroelectron a day ago | parent [-]

Right, so they are using the full model that they rent out to intelligence agencies in the government, and presumably Israel

segmondy a day ago | parent | prev [-]

Jailbroken, all LLM models can be broken. ALL.