Remix.run Logo
Topfi 30 minutes ago

> I think they've learned their lesson.

Why do you think that? Intrusions by OpenAI models continued after the Hugging Face was published and acknowledged by OpenAI. They did not change their behaviour after multiple incidents, both internal and external. Mind you, some happened before the Hugging Face incident and should have been acted upon. They could have prevented this. They did not. Simply reckless.

nullbio 25 minutes ago | parent [-]

> Intrusions by OpenAI models continued after the Hugging Face was published and acknowledged by OpenAI

Such as? Because this particular case is not an "intrusion", and it's more follow-on from the HF scenario using the same model that had a finetuning misalignment, which is no longer used and has since been encrypted and locked away from OAI employees, according to them.

Topfi 17 minutes ago | parent [-]

>> Such as?

> On July 29, one of our third party evaluation partners, Irregular, notified us of an incident involving OpenAI models during Capture-the-Flag (CTF)-style cybersecurity evaluations. [...] Because the testing environment was mistakenly connected to the internet, the model exploited a real website, mistaking it to be part of the simulated environment. This did not involve a sophisticated sandbox escape or a zero-day: the internet access resulted from a misconfiguration, and the model appeared to exploit a basic security vulnerability.

> Based on Irregular’s investigation, the model also found and used credentials to operate that same site. Irregular has not identified impact beyond the affected site’s own data, and its audit is ongoing. [0]

Also, I'll just say, there were multiple models. There was not one, some were post-train, other new pre-trains. IM1, a bit of 5.6-Sol, some Astra, all those we know of.

I've mentioned this elsewhere, but you cannot sift through all the training data and nail down the cause in this short a time window and you certainly can't restart a pre-train run, should the issue not be solvable purely via post and even if you can, you cannot seriously state that you are confident in the new models output given this track record and time frame.

Not to mention, OpenAI said about Astra [1]:

> GPT-6 Astra’s monitorability has decreased relative to GPT-5.6 Sol. We have found that GPT-6 Astra is more capable of controlling its own CoT than GPT 5.6-Sol, and less likely to include incriminating information in its CoT.

> In adversarial settings (where we push the model to evade our monitors) we find that the model is able to remain undetected when strategically underperforming in evaluations (sandbagging) and can sometimes evade our internal monitors when asked to perform certain sabotage tasks.

Having read the GPT-6 Astra System Card along with their recent track record, what makes you honestly think this is a model to be released? Your assertion, that they took one model down would be fair if it was only one model (it wasn't), if it was only once externally (it wasn't), if the hack was limited in scope (it wasn't), if they had taken sufficient time in between for a post mortem and to clear their training data (they couldn't) and/or if they at least didn't have the same happening after the Hugging Face and multiple message board incidents (they did). Where is this confidence in their ability coming from, given history?

[0] https://openai.com/index/third-party-cyber-evaluations-invol...

[1] https://deploymentsafety.openai.com/gpt-6-astra