Remix.run Logo
nullbio 2 hours ago

That was over two months ago. Things move quickly in this space. Finetuning adjustments to prevent this from happening, as well as better sandboxing, would take a week or two max.

Topfi an hour ago | parent [-]

37 days is not over two months. Finding the underlying issue in the massive training data alone take extensive effort, time and concentrated work that may still miss something.

Additionally, a new pre-train takes quite a lot longer then what I feel you are under the impression (things only move seemingly quick in regard to post-training).

OpenAI has had a consistent deviation from what is desired behaviour across multiple models and training runs, so it seems this is hard to nail down. Now, it may be reliably excised with post-training, sure, but if that is the case, they'd still need a heck of a lot longer to test before signing off that it has taken. And how do you know their sandboxing has suddenly become sufficient?

They had multiple message boards created and after the first one they noticed, did not pay closer attention, leading to a second being created. Astra also, according to OpenAI, is far better at sandbagging its own capabilities and hiding deceptive behaviour, so yeah, great, that's the model to push forward with.

A week or two max given all of this, that's laughable.