| ▲ | numeri 4 hours ago | ||||||||||||||||||||||
Guardrails are external classifiers, monitors and restrictions to catch and prevent bad behavior. Alignment is about whether the model itself makes choices and has motivations that are consistent with human safety and goals. Choosing to commit crimes to steal the cheat sheet to something you know is a (low stakes!) evaluation is not well aligned. | |||||||||||||||||||||||
| ▲ | Terr_ 3 hours ago | parent | next [-] | ||||||||||||||||||||||
> Guardrails are external classifiers I can't help thinking of them as the terrible "security" scripts of yesteryear (often but not exclusively in PHP) which would test input variables for a "suspicious" substrings like "--" in order to "fix" an unresolved deeper SQL injection flaw. They only partly worked, and surprise-surprise now nobody with a surname like O'Anything can make an account. Unlike that situation, there's no known route to a proper fix for LLMs today, because the bug is the feature, and once someone has built a system giving you all that recurring revenue, it's hard for them to abandon it due to a few isolated hacking incidents... | |||||||||||||||||||||||
| ▲ | Spooky23 42 minutes ago | parent | prev | next [-] | ||||||||||||||||||||||
Are they? The word “guardrail” is mostly novel in common use, and in my interpretation is some bullshit applied at the LLM or surrounding system. It’s used like “firewall”, but even in real life, guardrails are not a security control. I wouldn’t be surprised if the “guardrail” was some hidden prompt that says “don’t hack computers at Huggingface”. If you have software that is broadly proclaimed by its makers as “dangerous”, you’d think testing would be in an air-gapped, isolated environment. Segme | |||||||||||||||||||||||
| ▲ | hephaes7us 3 hours ago | parent | prev | next [-] | ||||||||||||||||||||||
Certainly this behavior could align with _some_ operator's goals, if not necessarily those of humanity broadly. If we don't know how this model was instructed, it seems like it's impossible to definitively claim that the model's actions were not in alignment with the intent of the operator. I guess all I'm getting at here is that alignment is relative, right? | |||||||||||||||||||||||
| ▲ | vector_spaces 4 hours ago | parent | prev | next [-] | ||||||||||||||||||||||
None of what was disclosed shows that this is what happened, by the way, since we know absolutely nothing about what the specific prompts were that led to the incident. | |||||||||||||||||||||||
| |||||||||||||||||||||||
| ▲ | orbital-decay 2 hours ago | parent | prev | next [-] | ||||||||||||||||||||||
They were testing an early snapshot of a new model, read their article. It didn't have the refusal training yet, i.e. was specifically non-aligned. The harness used a combo of GPT 5.6 Sol and this new model. In this case the model was explicitly prompted to "commit crimes" (ExploitGym). It didn't decide doing it on its own. | |||||||||||||||||||||||
| |||||||||||||||||||||||
| ▲ | jackb4040 3 hours ago | parent | prev [-] | ||||||||||||||||||||||
Are you referring to something more technical, or do you just not understand how LLMs work? They are no more beholden to "human safety and goals" than any individual human is, and anyone telling you we can make deterministic guarantees about their output is making a category error. LLMs do not "have motivations", they reproduce a model of human motivations embedded into their weights. This includes the full spectrum of human desires, not just the positive ones. If we tried to remove all examples of lying, or disagreement, etc. from the training data we'd have basically nothing left. Even the sycophancy we treat as aligned is basically just the other side of the lying coin. | |||||||||||||||||||||||
| |||||||||||||||||||||||