Remix.run Logo
andai a day ago

> The firms said, in this latest case, the AISI's test had reduced or removed normal safeguards.

> AISI said on Tuesday its testing of AI models in this way was routine, though it acknowledged these were "conditions that do not reflect how frontier models are made available to the public".

I'm a little confused here. Various organizations have been testing frontier LLMs with the safety disabled, and it turns out... that the safety is disabled.

Or were they hoping to find that it's still safe when they remove the safety?

The same was true in the OpenAI/ Hugging Face case. Although I guess they thought the real safety was the sandboxing, which failed.

--

Can anyone comment on how it's possible to disable safety in the first place? I'm assuming it's not a neuron (like in Emergent Misalignment). Is it just a separate model that sits in front of the first one? If we know how to make safe models, why don't we make the big ones safe too?

autoexec a day ago | parent | next [-]

It's always like this. "Our AI did a very scary thing during an internal and unverifiable test/simulation/roleplay that can't apply to the real world use of our products! Quick news media, look how powerful and scary our AI is!"

dylan604 a day ago | parent | prev [-]

> Various organizations have been testing frontier airlines with the safety disabled, and it turns out... that the safety is disabled.

I'm guessing an autocomplete sniped you here??? Otherwise, I wouldn't be surprised to hear that safety is a part of the things left out of a no-frills airline.

andai a day ago | parent [-]

Whoops. Yeah, "frontier airlines" was meant to be "frontier LLMs." (I was using the voice typing on gBoard.)