Remix.run Logo
▲ petu 3 hours ago

Devil's advocate, but Anthropic talks a lot about importance of alignment, so.. How can military ensure that model is NOT refusing to work just to be malicious?

Anthropic doesn't provide model w/o guardrails, suppliers have to use public version, newer model releases surely know about Anthropic standoff with the military.. how to prevent model from realizing that it's doing something for the military (supplier) and sandbagging and/or subtly sabotaging the results?

=====

DOD recommends sandbagging as defense mechanism against destination "attacks":

https://media.defense.gov/2026/Sep/08/2003992823/-1/-1/1/CSA...

Page 13

  When suspecting a malicious distillation campaign, consider varying changes to responses across requests to complicate response quality evaluations, such that the subtle changes avoid triggering obvious alerts. Reducing reasoning depth, presenting correct information with different reasoning, or stylistic inconsistencies may evade detection while reducing training usefulness.

  Avoid informing China-based AI company users suspected of distillation campaigns of a switch to a downgraded model. Informing malicious distillers would enable them to improve their defense evasions and indicate when to roll back training.
If that's seen as valid defense mechanism, then it's also a risk if used against them.
▲hardbass 3 hours ago | parent [-]

That would be stupidity of the government that caused this situation in the first place. Anthropics position was they wanted to have oversight before it was used to kill people. Anthropic was right. The us govt is fucking incompetent and managed to kill an entire girls school using "AI".