| ▲ | simonw an hour ago | |||||||||||||||||||||||||||||||
I don't think this exposes an alignment failure, because the test here was run with the alignment features deliberately turned off. OpenAI said: > We estimate maximal cyber capabilities by running this evaluation without production classifiers used to prevent models from pursuing high-risk cyber activity. It was a test of raw capabilities of the underlying model. | ||||||||||||||||||||||||||||||||
| ▲ | manux an hour ago | parent | next [-] | |||||||||||||||||||||||||||||||
I guess we could debate what counts as alignment, but I think my initial point remains that if the underlying base model needs these classifier guardrails so badly then the way we train the base models is creating fundamentally misaligned models that are happy to pursue illegal behavior. I'm sure OpenAI would argue that base model + guardrail is aligned, but considering the "relative intelligence" of these two pieces, the fact that guardrails can just be turned off, and these kind of incidents, I am not reassured. We may well get another "oopsie" moment with much more catastrophic consequences even from otherwise well intentioned actors. | ||||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||
| ▲ | mjamesaustin an hour ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||
Alignment isn't alignment if it can be turned on and off at the whim of company employees. This time the damage was minor, relatively speaking. What happens when a model just "testing its capabilities" breaks into banking infrastructure or government military assets? The damage could be catastrophic. | ||||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||
| ▲ | didibus an hour ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||
Classifiers and such are guard-rails, alignment to me and I assume most people, is about the model training, and it's tendency to respond, agree/disagree, push-back or not, be willing to cheat or even deceive the prompter, etc. | ||||||||||||||||||||||||||||||||
| ▲ | joe_the_user an hour ago | parent | prev [-] | |||||||||||||||||||||||||||||||
This is assuming a situation where, A) models "unaligned" by default and B) alignment can be added (though prompts and related things). The point, which isn't very surprising but still notable, is that the models are "amoral" out of the box. And moreover, we know that there is almost always a means to "jailbreak" them into that out-of-the-box capability (or that sometimes just randomly "jailbreak" in various ways). Also, saying the models are amoral doesn't mean they don't know good and evil - once they do acts defined as evil, they know "themselves" through their and so self-define themselves as evils (or objectively predict what a secretly/open evil actor would do based on their data). Which is to say I once laughed at the mis-alignment doomers but I can't see strong barriers against the doom scenario now. | ||||||||||||||||||||||||||||||||