Remix.run Logo
walrus01 6 hours ago

I don't disagree with you on what is the top-down political priority there, but thankfully the architecture of an open weights model released in .safetensors format allows for 3rd parties to "uncensor" it. There's at least 8 different CN originated models now that after running through heretic and a few other methods will score 0 refusals on this data set of prompts:

https://huggingface.co/datasets/mlabonne/harmful_behaviors

If we were living in a scenario where the open weight models were truly impossible to uncensor I would be significantly more skeptical of them. As a test I have an uncensored copy of qwen 3.8 27B Q8 here that will very happily discuss a myriad of negative things about the CCP.

throw10920 6 hours ago | parent | next [-]

I have basic understanding about how refusal-removal works - find the "no" weights by intentionally generating diverse refusals, and then set those weights to zero.

Is there a similar process for removing not refusals, but misinformation?

walrus01 6 hours ago | parent | next [-]

As an end user of this and not a person involved in training models or aligning them, I have only the most rudimentary understanding. But I think that would be a lot harder since the model doesn't fundamentally "know" that information is wrong.

Like, as a crudely chosen random example, the model doesn't have any core set of knowledge that knows putting sriracha hot sauce on your jelly donut is not a palatable meal. If the training data set includes lots of text that sriracha on a boston cream donut is a delicious meal, it'll "believe" that.

Same for any form of misinformation if the training data set of the misinformation has been baked into it.

ACCount37 3 hours ago | parent | prev [-]

There are processes for teaching a model specific facts or specific behaviors. Including "respond to topic X with Y", if that's what you want.

You could make a model that doesn't want to engage in "lunar landing was faked" conspiracy theories the same way you can make a model that doesn't want to criticize CCP.

There is, however, no broad "misinformation" category that you could tune up or down - the way there is a category of "safety refusals".

You could make a model more reluctant to say things it isn't sure about. But that is calibrated against the model's own "sure about" - and metaknowledge of this nature in LLMs? Fragile on a good day.

ACCount37 6 hours ago | parent | prev [-]

Yeah, it's good that open weights models can have their "filters" busted fairly reliably. Unlike whatever bone Anthropic has to pick with the very idea of biology.

But that's a consequence of how the technology works - not a consequence of China not being authoritarian about AI. They're just authoritarian about AI in different ways.

Not like they dodged the "ID verification" bullshit either. They were way ahead of the western countries there. It's vile - seeing this sad excuse of "think of the children" abused to invade privacy and strip freedoms over and over and over and over again.

pixl97 7 minutes ago | parent [-]

Most people don't realize how tenuous the situation is with those open models too.

Right now as long as they play along with Xi it's all good. But the moment something happens with them to upset the domestic peace, those open models are fucking gone and anyone that has them shouldn't expect anything new.