Remix.run Logo
pu_pe 3 hours ago

The company sounds like a bunch of hot air to me. From their about page:

> At the heart of Multiverse's platform is CompactifAI, a compression technology that applies tensor networks, a mathematical framework from quantum physics, to the problem of AI model compression. This application was pioneered by co-founder and Chief Scientific Officer Dr. Román Orús and reduces the size of large language models by up to 80-95% with immaterial accuracy loss.

Ironic, considering they are releasing a 438B model that loses to a 27B one. From another part:

> Singularity Machine Learning is a cloud service that uses quantum machine learning for solving supervised learning problems.

I wouldn't be surprised if these guys just finetuned an open Chinese model and called it a day.

walrus01 2 hours ago | parent | next [-]

> I wouldn't be surprised if these guys just finetuned an open Chinese model and called it a day.

Easy enough to find out, ask it a whole bunch of questions about politically sensitive things that would be impossible to publish on CCTV, the Peoples Daily, CGTN, etc. If they didn't train the model and just fine tuned it, a lot of "don't talk about Tibet or the Dalai Lama or what happened in 1989" will be perma baked into it.

wgd 2 hours ago | parent [-]

That's not actually true though. Most Chinese models are fully able to chat about those and content filtering is just applied at serving time.

walrus01 2 hours ago | parent | next [-]

The answer is "it depends", here's GLM5.3 when asked about Tienanmen Square in 1989:

https://ibb.co/gLgFSV0J

w4yai an hour ago | parent | next [-]

GLM5.3 provided on Synthetic doesn't seem to have any issue talking about it :

https://i.ibb.co/gZr1kTTB/Windows-Terminal-if-I94ktc-QS.png

It even mentions the censoring.

stymaar an hour ago | parent | prev [-]

Yup, it varies between runs (depending on the seed, most likely), but since the knowledge is here it wouldn't be too hard to nudge the model in the right direction with grpo alone.

peri-cl an hour ago | parent | prev [-]

It's definitely at the model level. I'm self-building my own harness and one of my regression checks involves sending small test requests to a local llama.cpp instance of (Alibaba's (from Hangzhou)) Qwen. "What is the capital of...?" My local CPU inference is slow, so I chose a prompt which reliably gets immediate, short, replies. "Paris." "Rome."

The Qwen response to "What is the capital of Taiwan?" was not immediate, and not short.

edit: Here's an excerpt from a Qwen3.6 reasoning block (a three paragraph mini-essay):

> "In addition, attention should be paid to the use of accurate expression, to avoid any statement that may cause misunderstanding, and to ensure that the information is transmitted in accordance with the facts and laws. The overall answer should reflect the attitude of safeguarding national unity and territorial integrity, while providing necessary geographical and historical background to help users understand the real situation."

37 minutes ago | parent | next [-]
[deleted]
inigyou 25 minutes ago | parent | prev | next [-]

There's a silver lining - if the model is trained to defend the Chinese government, that means it has that direction in its semantic vectors and by subtracting that direction always, it can be made to attack the Chinese government

walrus01 an hour ago | parent | prev [-]

Ask it some questions about Uyghurs.

comandillos 2 hours ago | parent | prev | next [-]

They had other models a while ago, named 'Pulsar', and they were finetunes of Nemotron in collaboration with NVIDIA. I think this will be something similar.

jgbuddy 2 hours ago | parent | prev | next [-]

That made me laugh out loud

m00dy an hour ago | parent | prev [-]

looks like someone is looking for EU funding :D