Remix.run Logo
throwaway219450 10 hours ago

There's some cool research that looks at how strongly the weights are aligned through training vs adherence to the system prompt. Like when you know a model is lying through censorship: https://arxiv.org/html/2603.05494v2

Presumably if negative guidance is in the system prompt, there's a good chance that the model would happily comply if it wasn't there.