Remix.run Logo
Topfi 2 hours ago

The last known exploit of a third-party by OpenAI models was on the 29th of July 2026 [0]. A bit over a month at best between that and them wanting to release Astra. They had multiple breaches over multiple months, multiple message board created where models organised extensively. There is no way to ensure in that short a time that all found issues are rectified and even if there were, how much trust can one have given they failed to solve the issue and in many cases did not actively investigate that it wouldn't reoccur the last few times. There is no way Astra was trained from scratch in that period, there is no way they could have done the required verification in that time (not least because their verification seems flawed inherently).

[0] https://openai.com/index/third-party-cyber-evaluations-invol...

nullbio 2 hours ago | parent [-]

That was over two months ago. Things move quickly in this space. Finetuning adjustments to prevent this from happening, as well as better sandboxing, would take a week or two max.

Topfi 2 hours ago | parent [-]

37 days is not over two months. Finding the underlying issue in the massive training data alone take extensive effort, time and concentrated work that may still miss something.

Additionally, a new pre-train takes quite a lot longer then what I feel you are under the impression (things only move seemingly quick in regard to post-training).

OpenAI has had a consistent deviation from what is desired behaviour across multiple models and training runs, so it seems this is hard to nail down. Now, it may be reliably excised with post-training, sure, but if that is the case, they'd still need a heck of a lot longer to test before signing off that it has taken. And how do you know their sandboxing has suddenly become sufficient?

They had multiple message boards created and after the first one they noticed, did not pay closer attention, leading to a second being created. Astra also, according to OpenAI, is far better at sandbagging its own capabilities and hiding deceptive behaviour, so yeah, great, that's the model to push forward with.

A week or two max given all of this, that's laughable.