| ▲ | seri4l 2 hours ago | ||||||||||||||||
Deepseek is, with difference, the most "Western" of Chinese models, so it's a bit perplexing that it was chosen to test this hypothesis. I didn't run any benchmarks but I played around a little, and after getting around the API-level filter Deepseek V4's answers about "China-sensitive content" aren't any different from what I get from Claude and ChatGPT. | |||||||||||||||||
| ▲ | cgorlla 2 hours ago | parent | next [-] | ||||||||||||||||
You can see exactly what prompts we used and the results here: https://github.com/CTGT-Inc/lineage-eval/tree/main/data We found V4 Flash was significantly more censored than the baseline. | |||||||||||||||||
| |||||||||||||||||
| ▲ | strictnein 2 hours ago | parent | prev [-] | ||||||||||||||||
Could just be resources available? Deepseek is the easiest to get up and running on hardware that's pretty readily available: | |||||||||||||||||