Remix.run Logo
sroussey 5 hours ago

Language models have always had an issue with negatives.

A negative like do “not” xyz is just not encoded the same as spelling out what you want vs what you don’t want.

Harder to write though.

sigmoid10 2 hours ago | parent | next [-]

I would say in this case abliteration is the likely culprit. To uncensor a model this way, you literally deactivate the parts that would enact refusals. As in things it was told not to do. But the real process is more like brain surgery performed by a alchemist according to an ancient religious book where noone involved really understands what is actually happening in the model.

AndyNemmity 2 hours ago | parent | prev [-]

Exactly, I wrote a blog post in what feels like a long time ago on this topic.

https://vexjoy.com/posts/positive-framing-agents-skills/

PotatoPrime an hour ago | parent [-]

Interesting read, thanks for re-sharing!

I noticed your joy-check link 404's now... I tried poking around your /skills/ folder but didn't find it easily. Should you still have that available I'd love to check it out.

edit: Found it if others are looking: https://github.com/notque/vexjoy-agent/blob/main/skills/code...