Remix.run Logo
walrus01 2 hours ago

> I wouldn't be surprised if these guys just finetuned an open Chinese model and called it a day.

Easy enough to find out, ask it a whole bunch of questions about politically sensitive things that would be impossible to publish on CCTV, the Peoples Daily, CGTN, etc. If they didn't train the model and just fine tuned it, a lot of "don't talk about Tibet or the Dalai Lama or what happened in 1989" will be perma baked into it.

wgd 2 hours ago | parent [-]

That's not actually true though. Most Chinese models are fully able to chat about those and content filtering is just applied at serving time.

walrus01 2 hours ago | parent | next [-]

The answer is "it depends", here's GLM5.3 when asked about Tienanmen Square in 1989:

https://ibb.co/gLgFSV0J

w4yai an hour ago | parent | next [-]

GLM5.3 provided on Synthetic doesn't seem to have any issue talking about it :

https://i.ibb.co/gZr1kTTB/Windows-Terminal-if-I94ktc-QS.png

It even mentions the censoring.

stymaar an hour ago | parent | prev [-]

Yup, it varies between runs (depending on the seed, most likely), but since the knowledge is here it wouldn't be too hard to nudge the model in the right direction with grpo alone.

peri-cl an hour ago | parent | prev [-]

It's definitely at the model level. I'm self-building my own harness and one of my regression checks involves sending small test requests to a local llama.cpp instance of (Alibaba's (from Hangzhou)) Qwen. "What is the capital of...?" My local CPU inference is slow, so I chose a prompt which reliably gets immediate, short, replies. "Paris." "Rome."

The Qwen response to "What is the capital of Taiwan?" was not immediate, and not short.

edit: Here's an excerpt from a Qwen3.6 reasoning block (a three paragraph mini-essay):

> "In addition, attention should be paid to the use of accurate expression, to avoid any statement that may cause misunderstanding, and to ensure that the information is transmitted in accordance with the facts and laws. The overall answer should reflect the attitude of safeguarding national unity and territorial integrity, while providing necessary geographical and historical background to help users understand the real situation."

37 minutes ago | parent | next [-]
[deleted]
inigyou 25 minutes ago | parent | prev | next [-]

There's a silver lining - if the model is trained to defend the Chinese government, that means it has that direction in its semantic vectors and by subtracting that direction always, it can be made to attack the Chinese government

walrus01 an hour ago | parent | prev [-]

Ask it some questions about Uyghurs.