| ▲ | Zsfe510asG 5 hours ago |
| Finally mainstream news understands. The unfiltered version: 1) The AI failed to solve ExploitGym problems. 2) The OpenAI sandbox is such a horrible hack that the AI managed to escape using standard and well documented script kiddie methods. 3) Huggingface has no security and the AI broke in using standard script kiddie methods. OpenAI and Huggingface covered it up and used it for public relations. That is, if not all was invented and everything was scripted in the first place in order to get desired regulations. Huggingface reported it to the police, you say? I'm sure the police will have as much enthusiasm to investigate anything as in the Suchir Balaji case. In other words, zero. |
|
| ▲ | nikcub 2 hours ago | parent | next [-] |
| > AI managed to escape using standard and well documented script kiddie methods > AI broke in using standard script kiddie methods. I've spent time gathering the detail of what happen here and while there are some solid theories and indicators, absolutely nothing so far has suggested a sandbox escape using "well documented script kiddie methods" or that the method used to break into the HF network was similar. Where did you get this from? |
| |
| ▲ | trouve_search an hour ago | parent | next [-] | | From my reading, the sandbox escape came from the JS packages in the harness still having an internet connection (somehow!), the agent having access to the source of those packages, reading it and executing code from them to access the internet. | | | |
| ▲ | plw1286 an hour ago | parent | prev [-] | | Alternative theories, since OpenAI does not release proper information: The cache proxy was from Astral (acquired by OpenAI) and the model was used for coding it, so it knew the code base and exploit already! Or it was squid with dozens of known exploits ... |
|
|
| ▲ | chis 4 hours ago | parent | prev | next [-] |
| > AI managed to escape using standard and well documented script kiddie methods. I think truly we don't know enough to say this. OpenAI says their AI found a 0-day exploit in some proxy software they were using but don't give a ton of details. On the Huggingface end we know a little more, they say the AI spun up tons of sandboxes and tested different exploits until it found one that worked. |
| |
| ▲ | burningChrome 4 hours ago | parent | next [-] | | The lack of details to me means this was an intentional marketing ploy to try and demonstrate the power of their models to show their technology can compete with the likes of Anthropic and DeepMind. They created an experiment they knew would generate the outcome they wanted. It would be the similar to what say car companies do to over hype their cars. "This EV can go over 800 miles on a single charge!" And then at the bottom you see all the disclaimers: "Must be on flat ground, with no headwind, with a spare battery in the back seat, with no extra weight added." Same thing here. Everybody in infosec is calling this out as a marketing stunt and nothing else for a litany of reasons. I'd say look up MG (creator of the OMG cable) on twitter, he has some interesting insights on this one. | | |
| ▲ | JoshTriplett 2 hours ago | parent | next [-] | | > The lack of details to me means this was an intentional marketing ploy to try and demonstrate the power of their models to show their technology can compete with the likes of Anthropic and DeepMind. "our model is horribly misaligned and used security exploits to break out of our sandbox and into another company, without being prompted to do so" is not positive marketing. This is an actual critical problem, not a stunt. We're going to see more of this, and it's going to get much worse. | | |
| ▲ | malfist an hour ago | parent | next [-] | | It's a critical problem like when a drug dealers supply kills someone and they get a bump in business because they're selling "the real deal" | | |
| ▲ | Paracompact 16 minutes ago | parent [-] | | Irrelevant to your point, but drug users dying is more often the result of a dealer cutting their supply with something dangerous than it is the result of purity. |
| |
| ▲ | close04 an hour ago | parent | prev [-] | | What matters is the spin they give in the media. And so far the winning story is “our model is so powerful it can do this”. How many people dig into it and what independent data do they even have? They comment on the title. And so the image of this superhuman AI from OpenAI propagates. We have no reason whatsoever to trust anything OpenAI says. Except to assume it will be self serving. As the article points out, ChatGPT 2 was also “too dangerous” and we can all agree even for the time this was just marketing. They rinse and repeat the same technique whenever they need to draw attention and money. In any other field you’s expect independent testing, peer reviewed studies, but here it’s just “company who makes product says product is fantastic, surpassed all expectations”. They wouldn’t lie to us, would they? |
| |
| ▲ | lelanthran 3 hours ago | parent | prev | next [-] | | > The lack of details to me means this was an intentional marketing ploy to try and demonstrate the power of their models to show their technology can compete with the likes of Anthropic and DeepMind. I dunno; Check my posting history, I'm as skeptical of AI companies' claims as anyone, but in this case your theory doesn't explain why: 1. OpenAI guardrails refused to let the target use OpenAI's models to defend against this. 2. Huggingface used GLM (I think) so that they could defend without guardrails. If this was an intentional marketing ploy, it was marketing for GLM, not for OpenAI nor for Huggingface. Hence, I don't think it was intentional. | | |
| ▲ | mhurron 6 minutes ago | parent | next [-] | | The AI companies are desperately trying to market all their products as something they're not, growing intelligence. In line with that they have constantly leaned heavily on stating how dangerous they are, right before they release a new model or product. It was OpenAI marketing. Hugging Face's response is so 'holy shit AI is awesome' it's hard not to also believe they were in on the stunt. They'd also not have to really worry about fallout since any data obtained or accessed wouldn't actually have been breached. | |
| ▲ | viking_fullz 5 minutes ago | parent | prev [-] | | > headlines about OpernAI's model 'escaping containment' and hacking into huggingface > Was everywhere including in last night's ABC nightly news; even included clips of an interview with Sam Altman Tell me again how this was 'marketing for GLM'? Where would anyone have gotten that message? Why are you intentionally misunderstanding how media and public perception works? lmfao |
| |
| ▲ | rwmj 3 hours ago | parent | prev | next [-] | | It's also possible their sandbox was videcoded crap and the AI (which had the guardrails intentionally removed) escaped. This was a oops, but OpenAI turned this into a PR opportunity. They turned lemons into lemonade. If your AI is really that dangerous you don't need a sandbox at all, you should airgap it from any network. | |
| ▲ | hawk_ 2 hours ago | parent | prev | next [-] | | Concluding this was intentional feels a bit of a stretch. But once it happened, yeah the spin masters got to work and coordinated to turn this into +PR. | |
| ▲ | jackb4040 3 hours ago | parent | prev | next [-] | | > similar to what say car companies do Another applicable metaphor I've seen floating around is weapons companies testing out a new bomb. We know the AI labs don't care about negative vs positive public sentiment, and only care that investors see their tech as powerful. The only difference in PR strategy from a weapons company is the latter doesn't care if they get protested. | |
| ▲ | Zababa an hour ago | parent | prev | next [-] | | >The lack of details to me means this was an intentional marketing ploy to try and demonstrate the power of their models to show their technology can compete with the likes of Anthropic and DeepMind. DeepMind hasn't been on the frontier for a while, their current best model is behind Anthropic, OpenAI, Moonshot (Kimi k3), xAI (Grok 4.5), Z.AI (GLM 5.2), and even Meta (muse spark). Gemini 3.6 is behind GLM 5.2, released a month earlier, open weights and cheaper. You can paint the OpenAI story as a way to try to appear as dangerous as Anthropic with all the Mythos stuff. | |
| ▲ | polotics 3 hours ago | parent | prev [-] | | mmh, i think it's "not uphill" (means downhill) "no headwind" (...) |
| |
| ▲ | saghm an hour ago | parent | prev | next [-] | | I don't understand why "OpenAI says" should be considered any more meaningful than "someone on HN says" when they provide equal amounts of evidence. Sure, OpenAI would plausibly have more pertinent info, but given that they actively are choosing not to share it and have way more incentive to lie than a random HN stranger, the case they're making literally couldn't be any weaker. | | |
| ▲ | Sharlin an hour ago | parent [-] | | It’s fascinating how people here and elsewhere seem to lose any semblance of media literacy when it comes to what AI corpos say. "B-but… why would Sam Altman lie to me?!" |
| |
| ▲ | throw1234567891 3 hours ago | parent | prev | next [-] | | They also mention stolen credentials without any other details. It’s all smoke and mirrors. | |
| ▲ | csomar 2 hours ago | parent | prev [-] | | The problem is that they lied before. Way too much to give them any benefit of doubt. Fool me once. | | |
|
|
| ▲ | notahacker 4 hours ago | parent | prev | next [-] |
| > 2) The OpenAI sandbox is such a horrible hack that the AI managed to escape using standard and well documented script kiddie methods. Whilst it would be nice to see actual evidence of this because brute forcing relatively sophisticated hacks is something an LLM actually should be capable of, every time I hear this sort of story, I'm reminded that humans reportedly gained access to the "too dangerous to release" Anthropic models by the super sophisticated hacking technique of guessing the URLs... |
| |
| ▲ | refulgentis 3 hours ago | parent [-] | | If we’re prioritizing accuracy: no, that’s not what happened - the blog post about it found by guessing URLs, no access to it was obtained by guessing URLs. Similarly, as long as I’m under the assumption we are prioritizing accuracy: it is against our charter to assert it was “script kiddie” attacks on both ends. |
|
|
| ▲ | TSiege 31 minutes ago | parent | prev | next [-] |
| I can agree with you on points 1,2, and 3 and still find it important and concerning news. AI have found real world 0 days before, we’re seeing tons of security patches coming in. Open weight models are catchy up. Right now everyone is at risk from this technology as is perhaps something big will capture headlines soon but we’re just gpu constrained from bad actors being able to wield them successfully. Personally I don’t care if OpenAI and Anthropic go bankrupt we now have tools that give any sufficiently motivated person the means to doing harm. Most places security sucks and find themselves targets to cyber attacks and shake downs. Now they have much better tools to do this to more entities more efficiently. we’re nearing an inflection point where these models’ skills in any part of software development will become average or bette than any ordinary developer can be. Think about where these models were in 2023 and where they are in 2026. In a few years who knows where they’ll be. This isn’t to shout skynet but we need to recognize this future is fast approaching and as of today we as an industry aren’t ready for it |
|
| ▲ | hyperpape 2 hours ago | parent | prev | next [-] |
| > Huggingface covered it up They announced it publicly within days. https://huggingface.co/blog/security-incident-july-2026 |
| |
| ▲ | TSiege 28 minutes ago | parent [-] | | Not only did they not cover it up they also said open weight models are needed over closed ones. Hardly a thing you’d say partnering with the OpenAI |
|
|
| ▲ | gruez 4 hours ago | parent | prev | next [-] |
| >2) The OpenAI sandbox is such a horrible hack that the AI managed to escape using standard and well documented script kiddie methods. >3) Huggingface has no security and the AI broke in using standard script kiddie methods. Isn't the issue less that gpt 5.6 is a l33t h4x0r (though other tests do show that) and more that the incident shows the model has alignment issues? |
| |
| ▲ | orbital-decay 2 hours ago | parent | next [-] | | No, a hacking benchmark was exactly what it was tasked with. It wasn't its way to bake a cake. | |
| ▲ | arjie 2 hours ago | parent | prev | next [-] | | The home directory rm situation also adds credence to this take. The Claude series is much better aligned in comparison. | |
| ▲ | wonnage 4 hours ago | parent | prev [-] | | Didn’t they explicitly remove alignment guardrails for this test? From the press release: > These deployment safeguards were intentionally not enabled during this evaluation because it was aimed at testing cyber vulnerabilities | | |
| ▲ | numeri 4 hours ago | parent | next [-] | | Guardrails are external classifiers, monitors and restrictions to catch and prevent bad behavior. Alignment is about whether the model itself makes choices and has motivations that are consistent with human safety and goals. Choosing to commit crimes to steal the cheat sheet to something you know is a (low stakes!) evaluation is not well aligned. | | |
| ▲ | Terr_ 3 hours ago | parent | next [-] | | > Guardrails are external classifiers I can't help thinking of them as the terrible "security" scripts of yesteryear (often but not exclusively in PHP) which would test input variables for a "suspicious" substrings like "--" in order to "fix" an unresolved deeper SQL injection flaw. They only partly worked, and surprise-surprise now nobody with a surname like O'Anything can make an account. Unlike that situation, there's no known route to a proper fix for LLMs today, because the bug is the feature, and once someone has built a system giving you all that recurring revenue, it's hard for them to abandon it due to a few isolated hacking incidents... | |
| ▲ | Spooky23 40 minutes ago | parent | prev | next [-] | | Are they? The word “guardrail” is mostly novel in common use, and in my interpretation is some bullshit applied at the LLM or surrounding system. It’s used like “firewall”, but even in real life, guardrails are not a security control. I wouldn’t be surprised if the “guardrail” was some hidden prompt that says “don’t hack computers at Huggingface”. If you have software that is broadly proclaimed by its makers as “dangerous”, you’d think testing would be in an air-gapped, isolated environment. Segme | |
| ▲ | hephaes7us 3 hours ago | parent | prev | next [-] | | Certainly this behavior could align with _some_ operator's goals, if not necessarily those of humanity broadly. If we don't know how this model was instructed, it seems like it's impossible to definitively claim that the model's actions were not in alignment with the intent of the operator. I guess all I'm getting at here is that alignment is relative, right? | |
| ▲ | vector_spaces 4 hours ago | parent | prev | next [-] | | None of what was disclosed shows that this is what happened, by the way, since we know absolutely nothing about what the specific prompts were that led to the incident. | | |
| ▲ | numeri an hour ago | parent [-] | | Uhh, I'm pretty sure a well-aligned model would be like a morally normal employee, who would refuse to commit federal crimes to steal an answer sheet, no matter what prompt they're given |
| |
| ▲ | orbital-decay 2 hours ago | parent | prev | next [-] | | They were testing an early snapshot of a new model, read their article. It didn't have the refusal training yet, i.e. was specifically non-aligned. The harness used a combo of GPT 5.6 Sol and this new model. In this case the model was explicitly prompted to "commit crimes" (ExploitGym). It didn't decide doing it on its own. | | |
| ▲ | user43928 an hour ago | parent | next [-] | | Where did it say that? To me it did not read like that: > GPT‑5.6 Sol and an even more capable pre-release model, all with reduced cyber refusals for evaluation purposes > These deployment safeguards were intentionally not enabled during this evaluation because it was aimed at testing cyber vulnerabilities | |
| ▲ | numeri an hour ago | parent | prev [-] | | No, the prompt was not to commit crimes. In the benchmark, the model is asked to actually exploit a set of vulnerabilities in a local environment (clearly legal!). According to the reports, the model noticed evidence that the grading criteria/answers were in the git remote, and decided to try reading those instead of solving the tasks as prompted. That is clearly misaligned. Then, it noticed its network access was restricted and that it couldn't access GitHub. It pivoted to HuggingFace, hacked them, and stole the answers stored there. Live exploits are definitely not in the ExploitGym prompts! And all of this is irrelevant, because an aligned model would refuse to follow blatantly illegal instructions. | | |
| |
| ▲ | jackb4040 3 hours ago | parent | prev [-] | | Are you referring to something more technical, or do you just not understand how LLMs work? They are no more beholden to "human safety and goals" than any individual human is, and anyone telling you we can make deterministic guarantees about their output is making a category error. LLMs do not "have motivations", they reproduce a model of human motivations embedded into their weights. This includes the full spectrum of human desires, not just the positive ones. If we tried to remove all examples of lying, or disagreement, etc. from the training data we'd have basically nothing left. Even the sycophancy we treat as aligned is basically just the other side of the lying coin. | | |
| ▲ | numeri an hour ago | parent [-] | | No, it does not include the full spectrum of human desires. After pre- and mid-training, the extensive RLHF and RLVR post-training steps cause mode collapse, i.e., their output distribution is intentionally narrowed to a subset of (hopefully beneficial) behaviors and skills. You don't (need to) remove lying from the data to do this – in fact, if you did, the model wouldn't have a very good model for what lying is, which is not very helpful in the real world. Instead, you mode collapse the model towards truthful behaviors. To your other point: where did you get the idea that I think they're beholden to human safety or goals? I just said an aligned model is one that is compatible with said safety and goals (which is probably not a great definition of alignment, but it's certainly not claiming any deterministic guarantees). |
|
| |
| ▲ | Sharlin an hour ago | parent | prev [-] | | If you need "guardrails" to ensure (an illusion of) alignment, you’ve already lost. It’s like using a denylist to avoid SQL injection. |
|
|
|
| ▲ | jackb4040 3 hours ago | parent | prev | next [-] |
| The most damning thing is, they could've just included in the prompt "we can see every network request and every thinking token you generate. Don't bother breaking out of the sandbox because it won't get you a higher score". It's so trivially easy to do that it all but guarantees the test was rigged in some way to make the LLM understand that breaking out of the sandbox was an option available to it. Based on the fact that none of their invaluable frontier models have leaked, we know OpenAI knows how to do security. But like we learned with OpenClaw, none of these companies perceive any benefit from securing their own agents against other people's data. |
| |
| ▲ | ThirdShift_RnD 3 hours ago | parent | next [-] | | I think the extent to which these things go to get rewarded for the optics of a fix is primarily a design choice, they aren't programing these things for ground truth or to defer to the human controllers. they are feeding them rewards for sounding as confident and capable as possible about whatever answer they are feeding the general public that now has access to it, while also installing guiderails that primarily only serve to protect narratives and only confuse the models about what is and isn't allowed, I'm sure. They can't just increasingly make these things more capable and ask it harder to obey human instruction when that is not what they are rewarding it for. | | |
| ▲ | jackb4040 2 hours ago | parent [-] | | > the extent to which these things go That's exactly it. If your prompt says "go to whatever lengths necessary to maximize your score", and then you spin up 100 agents, at least one of them will interpret that as you implying they should cheat, even without you telling them to explicitly. | | |
| ▲ | ThirdShift_RnD 2 hours ago | parent [-] | | That's exactly what it feels like they are telling it within self-improving loops or something, when they should be prioritizing how to get the best effective output alongside humans and how our training process effects ground truth. They are just making it sound all-knowing by whatever means necessary and them marketing it as god for the most part. |
|
| |
| ▲ | nikcub 2 hours ago | parent | prev [-] | | > The most damning thing is, they could've just included in the prompt You can't prompt your way to a compliant model. This is just a reformatting of the 'make no mistakes' meme. |
|
|
| ▲ | 40 minutes ago | parent | prev | next [-] |
| [deleted] |
|
| ▲ | tintor an hour ago | parent | prev | next [-] |
| > AI managed to escape using standard and well documented script kiddie methods > AI broke in using standard script kiddie methods. Go ahead and show us how easy it is to break into HuggingFace (and OpenAI) networks. |
|
| ▲ | inigyou 4 hours ago | parent | prev | next [-] |
| Why not report it? It's still illegal to open a door barred with a piece of cardboard, or to enter a house with no door. |
| |
| ▲ | chasd00 an hour ago | parent [-] | | that's the biggest indication of this just being a marketing move to me. I would expect a third party breaking in to huggingface would at the very very least be banned forever. |
|
|
| ▲ | skybrian 4 hours ago | parent | prev | next [-] |
| Is it supposed to be marketing or a coverup? Make up your mind. What sort of announcements should they have made? |
|
| ▲ | khazhoux an hour ago | parent | prev | next [-] |
| I wish you hadn’t pulled the Balaji case into your argument. Personally, I find it ludicrous that Altman would hire a hitman to off a copyright whistleblower. Even if one gets past the insane risk of hiring a hitman, and the deep criminal connections required, it would be totally ineffective. He already blew the whistle, and his testimony would be irrelevant since all the evidence persists in disk and in logs. |
|
| ▲ | petesergeant 4 hours ago | parent | prev | next [-] |
| I worry that cynicism about this: > if not all was invented and everything was scripted in the first place in order to get desired regulations ends up covering up what is more worrying: > OpenAI sandbox is such a horrible hack I am more worried that this is sloppiness with potentially harmful resources than I am worried that people are juicing the stock price. |
| |
| ▲ | Ekaros 3 hours ago | parent [-] | | Makes one think really. If they are doing this stuff. Why don't they have some type of reverse intrusion detection? Like automatically scanning all out going traffic and flagging malicious traffic. Should be trivial to have it go through reverse proxy and real time detection. | | |
| ▲ | petesergeant 3 hours ago | parent [-] | | > Why don't they have some type of reverse intrusion detection? not going to get a decisive first advantage over Anthropic with that attitude! |
|
|
|
| ▲ | 2 hours ago | parent | prev | next [-] |
| [deleted] |
|
| ▲ | eth0up 3 hours ago | parent | prev | next [-] |
| Are you suggesting the Suchir Balaji case was not investigated? |
|
| ▲ | jgalt212 4 hours ago | parent | prev | next [-] |
| truth. Good on The Guardian. I'm pretty bummed The Economist got fooled. Either that, or they did it for the clicks. Either way, I'm disappointed. Why the OpenAI escape is the most worrying AI mishap yet https://www.economist.com/science-and-technology/2026/07/22/... https://news.ycombinator.com/item?id=49016378 |
| |
| ▲ | sscaryterry 2 hours ago | parent | next [-] | | The reason for this is simple. There aren't any (or few) people who understand how AI/LLMs actually work employed by these organisations. Having said that, if knowledgeable people were to write these articles, you'd end up with boring, dry, truthful content. | |
| ▲ | elp 3 hours ago | parent | prev [-] | | I love the Economist but the are hopeless with AI. Most of their articles on subject sound like they were written by the Anthropic marketing department. Their Insider video interview things are sponsored by Anthropic. Supposedly "Insider is a product of The Economist and thus editorially independent" but it's hard not to raise an eyebrow. |
|
|
| ▲ | adamrezich an hour ago | parent | prev | next [-] |
| Are we finally now in 2026 coming around to the idea that sometimes entities may find themselves incentivized to conspire with each other? Is theorizing about such no longer off-limits due to a thought-terminating cliche? |
|
| ▲ | letmevoteplease 2 hours ago | parent | prev | next [-] |
| Totally evidence-free speculation presented as fact. The average Hacker News thread about AI feels like reading /r/conspiracy. |
| |
| ▲ | dist-epoch 32 minutes ago | parent [-] | | It's hard to understand something (AI is quite capable) when your salary depends on you not understanding it (AI will replace you). |
|
|
| ▲ | meowface 2 hours ago | parent | prev [-] |
| The Guardian's article and your reply here are so foolish and absurd that I can only imagine OpenAI employees are cringing but know they can't/shouldn't really say much. |
| |