Remix.run Logo
adrian_b a day ago

I assume that this announcement has been prompted by that of Moonshot AI, which has just announced a 2.8T parameter open-weights LLM, Kimi K3, to be published on Huggingface by 27 July.

Now the response of Alibaba is that they will also publish soon a big open weights LLM, the 2.4T parameter Qwen 3.8.

I wonder if Alibaba has always planned to make this big LLM open weights, or they have chosen to do this now, to better compete with Moonshot AI.

In any case, from this competition in LLMs, we win.

gardnr a day ago | parent | next [-]

It's hard to say what their motivation is. The Chinese firms seem to be working hard to commoditize intelligence which may be the most effective way to debase American frontier labs. And yeah: it also happens to be really good for humanity.

kelnos a day ago | parent | next [-]

> It's hard to say what their motivation is.

Feels pretty easy to me.

They want to turn LLMs into a commodity, and watch the US AI labs crash and burn.

There will still be plenty of customers who will pay them to host the models and run inference, even if the weights are open and others can offer competing products. (If necessary, the Chinese government can ban use of foreign inference services by Chinese citizens and businesses to give their own companies a domestic monopoly.)

When their models equal or surpass those from the Western AI labs, they can even stop releasing weights for new models, and keep all the inference revenue for themselves.

Meanwhile, they're still manufacturing much of the hardware that everyone in the world needs in order to run datacenters (see also: Spolsky's "commoditize your complement" essay).

Beyond that, it's a soft-power play. As the world keeps looking at the US more and more skeptically as an ally and superpower, Chinese companies releasing weights for competitive models is a way for China to look better and more world-minded.

TheDong a day ago | parent | next [-]

I feel like there could also be a simpler explanation.

Why does a debian contributor make debian free, why do they work on this thing anyone can use?

Is it because linux and debian hate windows and iOS and want to see american fail?

No, it's because most debian contributors believe software source code, information, should be free, users should be free to modify the code they use, and that they're building a thing they want to share with the world.

Maybe the chinese AI labs believe AI is powerful and useful, are proud of what they're doing, and want to share it as broadly as they can so everyone can use it.

There doesn't have to be any weird "chinese government" or "they hate the west" type vibes, it could just be the same thing as OSS, they're trying to do what they think is best for the world.

mceachen a day ago | parent | next [-]

Please don't conflate a volunteer effort with no expected economic gain with a very well funded company (or fleet of companies).

With the CCP's highly successful track record with subsuming other markets, Occam's razor applies to why they're doing this.

windsignaling 21 hours ago | parent | next [-]

There is clearly an anti-China bias here. Show me comments demonstrating the same level of distrust against Google for open-sourcing projects like Tensorflow, Kubernetes, Flutter, Chromium, etc.

venussnatch 20 hours ago | parent | next [-]

The chromium example is wild. There's an extreme distrust and contempt for chromium becoming the defacto browser and therefore Google becoming the defacto gatekeeper of the web.

gizajob 16 hours ago | parent [-]

They’ve been that since about 2003 regardless.

kelnos 22 minutes ago | parent | prev | next [-]

I mean, I expect Google open-sourced those projects because they see economic benefit to themselves in doing so, not because they are good-hearted.

Chromium is an especially silly example to use: they benefit by controlling the web platform, and open-sourcing Chromium has allowed them to get their web engine into many other browsers.

benterix 7 hours ago | parent | prev | next [-]

There is a lot of distrust against Google and Microsoft here, also when they opensource stuff, and Chromium and AOSP are great counterexamples.

thegrim33 19 hours ago | parent | prev | next [-]

Scroll the front page. Find literally any story that has to do with a major US tech company. Open the comment section. Look at the the top comment. It will be negative. Most of the other top comments as well. Trying to gaslight us into not believing our own eyes ..

d5lt5 10 hours ago | parent | next [-]

* "Claude Code uses Bun written in Rust now"

"Drilling into the original article where Jarred explained the reasoning behind the change, It's pretty clear that under zig the team was doing things by hand that are automatic in rust."

* Claude Fable produced a counterexample to the Jacobian Conjecture

"This is a rare instance where feeding this groundbreaking information into an LLM gives _them_ psychosis. I fed this to claude code and watched it verify the result in 7 different ways to be 100% certain, and it was just flabbergasted. Quite remarkable."

* Ollama: All Aboard Open Models

"A year and still no implementation for such a basic need as offloading MoE layers onto the CPU selectively. On llama.cpp I can get models like Qwen 35BA3B running partially on gpu/cpu with 40t/s on a laptop thanks to --n-cpu-moe but on this VC funded joke it would be simply unusable. I can't quite understand how you make a wrapper so much worse than the code you're ripping out."

* Blender 5.2 LTS

"Look at that wow: https://www.youtube.com/watch?v=gqfLYIJMv7I"

* OpenAI reduces Codex Model Context Size from 372k to 272k

"I know a lot of people like to say that compaction makes this moot, but the level of detail you lose across compaction is wildly too much for most things that I do, unfortunately."

* M-Chips: M7 with up to 1.5 TB – and why Apple is skipping the M6

"I hope they put a better connector than TB5 so we can cluster them properly at 1TB/S"

--

Your math is not mathing.

sph 7 hours ago | parent | prev [-]

> Open the comment section. Look at the the top comment. It will be negative

My favourite pastime is clicking on any comment section of any topic whatsoever, betting that the top comment is someone going "well, ackshually..."

HN loves a contrarian skeptic, it has nothing to do with Western tech or AI specifically.

akoboldfrying 21 hours ago | parent | prev | next [-]

I'll supply such a comment: Any software open sourced by any for-profit company, including Google, is a calculated move ultimately intended to increase their bottom line, and it's naive to think otherwise.

jcelerier 13 hours ago | parent [-]

Isn't it borderline illegal for publicly traded companies to not try to maximize profit? "interest of the company" in theory but in practice

deaux 10 hours ago | parent | next [-]

No, it is not. This is a common misconception that the large companies are very happy about.

1. The century-old Ford decision wasn't about this. It was about him refusing to pay dividends to shareholders he was feuding with. I.e. dominant shareholder using his control to starve minority shareholders of returns.

The court still let him keep spending huge sums on factories and price cuts that didn't clearly maximize profits.

2. Even if the case was about profit maximization (it wasn't), the judgment was a Michigan state decision. It doesn't have force outside of it.

3. US business law is in practice actually the opposite. A business can do pretty much whatever it wants. This is fundamental.

4. Even if it were illegal (it's not), anything can reasonably be framed as being in the long-term interest of the company, including donating to causes, raising employee wages and so on. Courts never question this.

Just imagine this was a thing. It's completely untenable for this reason. Who is a court to judge that something isn't in the longterm interest of shareholders unless it's literally spending all the company money on yachts for personal use?

5. No company has ever been prosecuted for this, obviously, because it's not a thing that exists. The closest you can get is that there's a potential duty to seek the best price if a company must be sold or broken up. But that's a very specific situation.

It's 100% a myth. Feel free to copy this and spread it when you see someone saying this.

postepowanieadm 11 hours ago | parent | prev [-]

Think about the shareholders!

snowpid 21 hours ago | parent | prev | next [-]

Chromium and Kubernetes do have their business function. Tensorflow had a business function but it failed.

So Google is not a great counter example.

pessimizer 21 hours ago | parent | prev [-]

> There is clearly an anti-China bias here.

Yes.

> Show me comments demonstrating the same level of distrust against Google for open-sourcing projects like Tensorflow, Kubernetes, Flutter, Chromium, etc.

There's a search bar at the bottom of the page. Most people, including me, detest google and think of them as a thoroughly dishonest and sleazy organization. Thanks for giving me the opportunity to not boost China for a moment in order to accuse you of dishonest rhetoric.

It's almost insane that you've posted "comments demonstrating the same level of distrust against Google for open-sourcing [...] Chromium[...]" as if that's not a hobby for thousands of people (including me.)

edit: this has got to be a submarine shill account: 226 karma in 10 years. If this is true, pleas stop. China is doing a good enough job that they don't need it.

cma a day ago | parent | prev | next [-]

It could be as simple as do open releases and publication at first to help recruit talent who want that or who want to make a name for themselves. learned from the likes of... OpenAI, Google, Meta, Emad

seanmcdirmid 17 hours ago | parent [-]

There must be some Alibaba posters who could clarify this. I think it’s like how Amazon pushed AWS, alibaba is hoping create a similar ecosystem. I wouldn’t conflate them with the CPC or other more niche players like deepseek.

shimman a day ago | parent | prev | next [-]

IDK how this doesn't apply to American tech corporations too? Or is it only scary when America corporations face actual competition nowadays where they can't rely on the US government to bomb/sanction competitors?

Der_Einzige a day ago | parent | prev | next [-]

[flagged]

jchw a day ago | parent | next [-]

Well of course nobody believes that, there are a large number of great Chinese open source projects, and plenty of great Chinese contributors to open source projects. I have zero doubts that Chinese people are at least equally as capable of embracing open source as anyone else. It would be strange to suggest otherwise.

That said, though, I do have trouble believing the long-term story for open weights, anywhere. We do not need an evil government for open weight to "make sense", but I do think we need some government involvement for open weight to make sense in the long run. Otherwise, it's not 100% clear how they could be sustainable, and I don't think massive companies really can be trusted to just be philanthropic with no incentives indefinitely (or really, much at all to begin with.)

Chinese models being open weight does help them gain some Western mindshare, whereas for obvious reasons Americans would be very suspicious of running their source code and prompts through Chinese providers. (And I think that's justified, I just also think that American providers aren't really that much better in the long run, and you should prefer to not have to go through any provider for true privacy.)

nylonstrung a day ago | parent [-]

The CCP encourages open source contributions from companies, you can see this in how much Alibaba open sources across their entire stack

Chinese tech leans much more heavily towards build vs. buy than the SaaS dominated West (where programmers are more expensive) so the positive externalities on their tech industry are more pronounced

jchw a day ago | parent [-]

Sure, when it comes to the big corporations I don't deny CCP's influence on their strategy. Hell, in some cases, it's easy to root for their strategy, because sometimes our tech industry sucks. Don't have to love them to occasionally agree with them.

But, in terms of individuals, of course, we're really not so different.

KeplerBoy a day ago | parent | prev [-]

Sir, this is a multi-billion dollar operation. There have to be some incentives.

bigyabai a day ago | parent | next [-]

You could argue that Linux is a trillion-dollar project, it doesn't shift the goalposts.

abalashov a day ago | parent | next [-]

Yeah, but the trillion dollars are not concentrated in the Linux kernel project, but instead distributed all over the planet. That's not how the AI lab economy works.

polyomino a day ago | parent | prev [-]

Linux is a huge project and is not really comparable to Debian

esseph a day ago | parent [-]

This is an interesting statement that I found myself reading from two different interpretations of what you meant, and I find both possibilities to be equally true depending on your POV.

m000 a day ago | parent | prev [-]

[flagged]

liamwire 18 hours ago | parent [-]

I like the comparison, but I'm not sure it holds. What did it cost him to develop beyond intangibles such as his time and money forgone? Because the latter especially is not equivalent to upfront capex and opex costs that preempt any ability to be charitable. Genuinely asking, but I'd be hard pressed to think it's even within the same magnitude.

rzerowan 17 hours ago | parent [-]

Well there is a close equivalent today where a similar situation is playing out.With the argument being made that the dev cost aside - the pricing is being tied to what the potential future treatments would have cost.

Save a child from crippling illness and a lifetime of pain with a single dhot: Jonas Salk 0$ ,Doug Ingram 3.2M$ [1]

[1] https://web.archive.org/web/20260118223436/https://www.cbsne...

watwut a day ago | parent | prev [-]

1.) Majority of debian contributions are from people paid for it.

2.) It is scary because they do what sillicon valley did for decades? While it bragged about disruption being the goal?

edm0nd a day ago | parent | prev | next [-]

Yeah the Chinese totally have a really good history with being completely open and giving lol. The Chinese government totally has not been hacking into American and Western fortune 500 companies for the past few decades stealing R&D and tech to use for themselves. The Chinese also totally do not steal hundreds of billions of dollars of IP from America annually. Totally not something they would do!

asdewqqwer a day ago | parent | next [-]

It is hilarious to see people from arguable the most polarized political systems in the world believing the evil 1.5 billion people across the sea share one single mind, either a saint, or a devil.

kelnos 20 minutes ago | parent | next [-]

Did the GP actually say that? I think you're putting words into their mouth.

They specifically referenced the actions of the Chinese government, as well as mentions of IP theft. I don't think that covers 1.5 billion people. More like a few thousand or tens of thousands?

WillPostForFood a day ago | parent | prev | next [-]

It will be a great day for China and the world when the Chinese people are free from a totalitarian dictatorship. But until then we have to speak of the policy of the Chinese government as the policy of China, even if many, or most disagree with those policies.

21 hours ago | parent | next [-]
[deleted]
litbear2022 16 hours ago | parent | prev | next [-]

Thanks for reminding me—it was China that robbed TikTok.

justatdotin 17 hours ago | parent | prev | next [-]

now do usa

d5lt5 10 hours ago | parent | prev | next [-]

Nickname checks out.

ClumsyPilot 9 hours ago | parent | prev | next [-]

> we have to speak of the policy of the Chinese government as the policy of China

This does not even have meaning in English.

Policy of US government is not policy of US? When Trump decided to bomb Iran it was just a suggestion?

a day ago | parent | prev [-]
[deleted]
fhn 19 hours ago | parent | prev | next [-]

when the CCP controls media, news, corporations, they basically control the minds of their people. How do you suppose the Wuhan virus got so out of control? You mean to tell me anybody in China can say bad words about the government and their leaders, bookstore owners/employees are not getting arrested for selling books, students are not getting arrested for sharing their opinions, women are not forced abortions and sterilizations for the one child policy?

Gigachad 15 hours ago | parent [-]

In the US when you speak out against the government you get your media licenses revoked, sued in to a lifetime of debt, and the president sends a goon squad to your area to kill some random people on the street.

mlazos 8 hours ago | parent [-]

Yeah but we still don’t have execution vans! I’m not even being sarcastic. China is unequivocally worse even with our recent slide into towards totalitarianism.

girvo 20 hours ago | parent | prev | next [-]

But that’s not what they said?

unethical_ban a day ago | parent | prev | next [-]

It's hilarious seeing you hear a complaint about the CCP and Chinese tech bros and think it's applied to every citizen.

dudefeliciano a day ago | parent | prev [-]

[flagged]

pegasus a day ago | parent | prev | next [-]

I don't see how stealing IP is inconsistent with a sympathy for openness. If anything, it's the opposite. "Information wants to be free" and all that. They were just liberating those secrets ;)

Gigachad 15 hours ago | parent | prev | next [-]

Thank god US AI companies generated all of the training data themselves and certainly didn’t just steal it all.

freehorse a day ago | parent | prev | next [-]

Industrial espionage is neither a chinese invention nor (specifically) chinese practice. American companies do it to each other even.

mullingitover a day ago | parent | next [-]

The U.S. was openly a pirate nation for most of the 1800s.

Like with the ICC, the US respects or doesn’t respect international law strictly when it’s beneficial to the state’s interests. China really can’t be held to a different standard. This activity is grade school level geopolitics: the global system of government is anarchy.

gf000 19 hours ago | parent [-]

Yeah, I love how people assume large powers have to do something. They are at most expected to do certain reasonable stuff (most of the time, a war is not beneficial to anyone so they try not to escalate that far), but there is not really global "law enforcement".

applicative a day ago | parent | prev [-]

Agents of American companies who don't mind 15 years in the federal pen do it to American companies. So it must be right.

gf000 19 hours ago | parent | prev | next [-]

How can you steal something that won't materially be less from that? IP laws are cancer, and are definitely not in humanity's interest.

chromadon a day ago | parent | prev | next [-]

Why do you think competing companies poach engineers from each other?

tehjoker a day ago | parent | prev [-]

Whether or not this is really true or just US propaganda, the majority of technology transfer has occurred through the open process of requiring US companies to form partnerships and disclose know-how to access the chinese market. This stuff about espionage is really sour grapes from losers.

bee_rider 2 hours ago | parent | prev | next [-]

The comment you replied to implies a lot of over-complicated motivations and grand coordination. But I think “commoditize your complement” explains a lot here and is probably even simpler than your explanation.

jwarden 15 hours ago | parent | prev | next [-]

That’s a lovely thought but that’s not how China works. China is not a democracy, the geopolitical goals of the Chinese government are clear, and no large strategic company acts without approval and heavy influence from the CCP.

AYBABTME 13 hours ago | parent | prev | next [-]

The better parallel is "why did Google make Kubernetes open-source" or "why do large for profit entities engaged in competition, use open-source as a strategy against their competitors"?

"Commoditize your complement" -> https://www.joelonsoftware.com/2002/06/12/strategy-letter-v/

Agingcoder 17 hours ago | parent | prev | next [-]

Models are treated as weapons with export controls - if they can do this it’s with the blessing of the Chinese government who’s getting something out of it.

It’s fairly obviously about being a nuisance to the US.

marcus_holmes 16 hours ago | parent [-]

Not everything that happens in the world is about the USA

harshreality 18 hours ago | parent | prev | next [-]

> Is it because linux and debian hate windows and iOS and want to see american [closed-source duopoly] fail?

maybe a little?

3x3m3 a day ago | parent | prev | next [-]

They believe that information should be so free so much so that they are completely censoring Tiananmen Square answers.

yieldcrv a day ago | parent [-]

This interpretation of the Chinese constitution is nearly 40 years old, companies comply with it, what do you expect?

Despite have a free speech clause, it also has a national security clause that is used to control all facets of life and override all other rights in the constitution. Anything deemed to slightly alter China/the Party’s unity is reprimanded and illegal. Multi party states can fall into the same trapping if they give their national security law constitutional force, always one court ruling away.

Yes, China is a single party system making the constitution redundant and any nominally marxist regime would find a way to do the same out of necessity.

Follow your Chinese AI in thinking mode to watch its opinion of Tianamen Square references, it will be candid enough for your sensibilities and show how it operates around guardrails

NickNaraghi 16 hours ago | parent | prev | next [-]

I think your reasoning is actually more complex than the person you are responding to, unfortunately. (As someone who as published broadly used open source software)

chews a day ago | parent | prev | next [-]

Deepseek spun out of a hedgefund that took a huge short position on Nvidia. China is actively looking to switch to chips made by huawai and ween themselves off of the difficulty of sourcing nvidia.

try-working 16 hours ago | parent | prev | next [-]

No, open models is a marketing strategy

jknoepfler 3 hours ago | parent | prev | next [-]

The capital required to train Qwen at scale is enormous. The capital required to patch a linux distro is near zero. Any scale model coming out of China should be viewed as advancing a geopolitical agenda. The same skepticism should be applied to any model trained in the states, but that should be viewed through the lens of short-term business profits.

There are earnest nerds everywhere, in every society. No doubt. But "Chinese AI Labs" operate at the whim of the Chinese government, in the same way "American AI Labs" operate at the whim of billionaire investors. Inferring good will from either is naive at best at this scale.

elmer2 a day ago | parent | prev | next [-]

The Chinese government is so nice and giving.

In fact, they have so much love in their hearts for the Uyghur people, they created a special mobile app for them, just to make sure nothing bad happens to them.

throwaway-11-1 a day ago | parent [-]

You do realize the US was fighting Uyghur terrorists alongside China 20 years ago in ago in Pakistan’s and Afghanistan for their support and cooperation with Al Queda? Like I know everyone is supposed to hate China now or whatever but can you guys show a little consistency?

breezybottom 18 hours ago | parent [-]

The only way that's an inconsistency is if you think all Uyghurs are terrorists.

drschwabe a day ago | parent | prev [-]

Where's the drama in that !?

hypfer a day ago | parent | prev | next [-]

> They want to turn LLMs into a commodity, and watch the US AI labs crash and burn.

Man, imagine Darios face when suddenly, he cannot decide anymore what other people consensually do with their own hardware in their free time.

Rumpelstilzchen.

seizethecheese 21 hours ago | parent | prev | next [-]

There’s a simpler explanation, which is that this is how Chinese business operates.

When I was in China earlier this year the big topic of conversation was “overproduction”. The big example was electric cars, where there were too many companies making too many cars and making revenue but no profit.

It was explained to me that generally Chinese firms will compete hard and maximize revenue above all, whereas western firms tend to focus on profit.

(And of course this is clustered around industrial sectors that the government favors, so there is some high level strategy in going after AI, but maybe not the commoditization.)

behnamoh 19 hours ago | parent [-]

> It was explained to me that generally Chinese firms will compete hard and maximize revenue above all, whereas western firms tend to focus on profit.

ok but why?

tough 11 hours ago | parent [-]

[flagged]

g42gregory 18 hours ago | parent | prev | next [-]

As I recall, in 2023/24 OpenAI told US Gov that AGI will be achieved in 2026 and we will use this AGI to “dominate” China. Based on this, US Gov cut off China from all AI hardware. It looks like China got the message and this is the response.

Also, a little competition is good for everybody (especially US consumers), no?

jpfromlondon 8 hours ago | parent | prev | next [-]

>They want to turn LLMs into a commodity, and watch the US AI labs crash and burn.

You can say that, but they are at least better at democratizing AI than the American labs, and on seeing the US labs crash and burn we are at least aligned.

watwut a day ago | parent | prev | next [-]

> They want to turn LLMs into a commodity, and watch the US AI labs crash and burn.

I want to watch that too.

If they take Meta and Musk with them, all the better, but that is just dreaming I am afraid.

jasondigitized a day ago | parent | prev | next [-]

Or it could be good for humanity.

alightsoul 21 hours ago | parent | prev | next [-]

>When their models equal or surpass those from the Western AI labs,

that moment is now. They are not doing it at least not yet.

nicman23 12 hours ago | parent | prev | next [-]

to be honest i think this is a open source vs close source software.

some companies chose to have the benefits of one or the other

ClumsyPilot 10 hours ago | parent | prev | next [-]

> They want to… watch the US AI labs crash and burn.

I don’t think dozens of large independent companies and thousands of researchers are working just to spite Sam Altman. Ad much as I don’t like him, I have other things to do and I am sure so do they.

kelnos 16 minutes ago | parent [-]

Agreed, but the Chinese government has much more of a say in what their companies do than is common in the West. It seems perfectly plausible to me that the Chinese government wants US AI labs to fail, and might direct some/all of their own AI companies to release their model weights.

esafak a day ago | parent | prev | next [-]

They're selling the robots that use these models. The rest of the world just hasn't got round to embodiment yet.

wood_spirit a day ago | parent | prev [-]

As soon as the competition is bankrupted they no longer need to release for free? It’s like how big players enter markets by launching at a loss to destroy competitors?

dannyw a day ago | parent | prev | next [-]

In China, you can’t officially use US APIs. The world saw a taste of this with Fable, but in China, this has been the situation all along.

So it’s not a surprise why open weights are so cherished. As frontier models continue to block everyday individuals from securing their own codebase, I expect the adoption and usage of open weights to continue.

As an example, HuggingFace recently was investigating a security incident and got locked out of frontier closed APIs. Yes, HuggingFace.

https://huggingface.co/blog/security-incident-july-2026

Wowfunhappy a day ago | parent | next [-]

Here's the relevant quote from parent's link:

> When we started the log analysis, we first used frontier models behind commercial APIs. This did not work: the analysis requires submitting large volumes of real attack commands, exploit payloads, and C2 artifacts, and these requests were blocked by the providers' safety guardrails, which cannot distinguish an incident responder from an attacker. We ran the forensic analysis instead on GLM 5.2, an open-weight model, on our own infrastructure. This had a second benefit: no attacker data, and none of the credentials it referenced, left our environment.

> This experience points to a gap worth planning for. We do not know which model powered the attacker's agents, whether a jailbroken hosted model or an unrestricted open-weight one; either way, the attacker was bound by no usage policy, while our own forensic work was blocked by the guardrails of the hosted models we first tried. The practical lesson for defenders: have a capable model you can run on your own infrastructure vetted and ready before an incident, both to avoid guardrail lockout and to keep attacker data and credentials from leaving your environment. This is not an argument against safety measures on hosted models, and we are sharing this feedback with the providers concerned.

Yeah, big problem! Although I'm kind of surprised HuggingFace doesn't have access to Mythos? Or maybe Mythos still has some guardrails.

zapkyeskrill a day ago | parent | prev | next [-]

How does this explain open weights? They could easily take the same closed route like their American friends

yorwba a day ago | parent | next [-]

Well, if you look at Alibaba's financials for FY 2026 https://data.alibabagroup.com/ecms-files/1514443390/5b9061ed... their sales and marketing expenses rose by about 100 billion RMB (10% of revenue), "primarily attributable to the investment in user experiences of Alibaba China E-commerce Group and user acquisition of Qwen app."

So it seems like it's very important to them that people use the Qwen app and they're willing to pay a lot of money for that. Presumably someone thought that keeping their best models closed would drive more business to them (as the sole provider) but then they discovered that closed releases mostly get ignored unless they're really good. (See also: People who think that Chinese AI companies are required to release weights as a matter of policy, because the closed ones hardly ever show up in the news.) Releasing weights for Qwen 3.8 at least lets them get some of that "pretty good for the price of free" media buzz.

close04 a day ago | parent | next [-]

They’re also trying to take an axe to the lead the US has in the field at a time when sovereignty and “owning your platform” are the words of the day. Open source/open weight LLMs can steal the lunch of US competitors even if they aren’t the best of the best.

yorwba a day ago | parent [-]

I think they're probably more concerned about their Chinese competition, considering that despite all that spending, the Qwen app still trails Bytedance's Doubao in terms of monthly active users: https://www.aicpb.com/ai-rankings/products/china-ai-rankings Though Quark in third place is also made by Alibaba, so put together they're almost caught up with Doubao + Jimeng (place 7, also ByteDance).

kelvinjps10 17 hours ago | parent | prev [-]

t that keeping their best models closed would drive more business to them (as the sole provider) but then they discovered that closed releases mostly get ignored unless they're really good. But it's different since they don't have access to the american ones the companies there could make it all closed source

roenxi a day ago | parent | prev | next [-]

Open sourcing is a complex decision so who knows what their calculations are.

But I'd assume that they're preparing for some sort of winner-take-all market in model quality where if they don't do anything the winner will be aggressive, hostile and American. Likely trying to push the Chinese economy back to the year 2000. If that is the starting point either the Chinese have to win the market (unlikely) or squeeze the profit out of it to make winning the market meaningless.

Publishing high quality open models is a well known tactic for profit squeezing. Being 2nd place with the same business model as the front-runner is a losing strategy in a winner-take-all market so they aren't going to bother with that. But if they can commoditize the model, their superior energy costs and likely coming chip manufacturing wave will hopefully give them a big advantage.

ethbr1 a day ago | parent [-]

Radiolab did a great episode that covered the history of Chinese character computing: https://radiolab.org/podcast/wubi-effect

To summarize, in the 70s and 80s, China was facing an existential threat with their inability to access an economic accelerator (widespread computing) in their native language.

To the extent that there was serious consideration at the highest levels of converting the entire country to an alphabet-based writing system.

I'd expect they're looking at AI the same way:

We have to have access to this. Most of the frontier labs are American (or European). Therefore we need a solution we have continued access to.

Open weights feels simultaneously Chinese in nature (progress through making a design copyable and improvable by a large number of people) and economic (providing an incentive for the world to use Chinese models over other frontier).

traceroute66 a day ago | parent | prev | next [-]

> How does this explain open weights? They could easily take the same closed route like their American friends

Because they are playing the Americans at their own game.

What is the first thing an American company would do ?

Spread the old American classic FUD ... "you can't used this closed tool because its run by the communists", right ?

So you release it as open weights which is a win-win. Global adoption of the model and you get to give the American AI companies a kick in the nuts because you know they will never release open weights apart from highly quantised crippled shit.

The Chinese are also playing the long game. The gradual rebalancing of the world from the US-centric model of the past. If releasing models as open weights is part of that long game, then so be it.

Grombobulous a day ago | parent | next [-]

I think most of us that will claim to understand China are going to end up being wrong, unless any of us live there or grow up there. There’s a saying about China I have heard from ex-pats: the more you know about China, the less you know about China.

The point of me bringing that up is to say that what follows is really just my best guess:

If I were to judge from China’s approach to hardware, I think that the companies releasing open weight AI for free aren’t as worried about giving away too much as the West tends to be, just like a factory making robot vacuums isn’t worried about other factories copying their methods.

For one thing, Chinese firms are spending an order of magnitude or two less money training their models. They have pursued efficiency in a way that Western companies with insane capital systems haven’t bothered, and in some cases they’ve had to given their limited access to bleeding edge hardware via export restrictions.

My best guess is that more important than that, Chinese companies don’t see the open weight model itself as the value add.

At this point I don’t think we pay for Claude specifically for the model. If that was the case then we’d all be using cheaper/free models from China as they are the best model value. Basically, any time we decide not to use Fable or Opus to save costs, what’s the point of spending more than competing models to use Sonnet and Haiku?

The real reason we are using Claude is for the SaaS aspect of it. It has a toolchain, a friendly interface, and a bunch of integrations with business applications.

In this respect, it’s somewhat surprising that Western AI companies don’t publish open weight models more frequently. The struggle of setting that up yourself and figuring out which hardware can run it should be an advertisement for Claude and the rest.

KerrAvon 20 hours ago | parent [-]

>The real reason we are using Claude is for the SaaS aspect of it. It has a toolchain, a friendly interface, and a bunch of integrations with business applications.

That is a very thin moat, though. There's nothing you can do with, for example, Claude Code + Opus 4.8 that you can't do with your own custom harness running API-level Opus 4.8, which means that if you can afford the hardware (the moat for running any SOTA model) you don't need to pay Anthropic anymore.

I'm not saying they shouldn't, but I understand why they don't.

Grombobulous 19 hours ago | parent | next [-]

You might be right that it’s thin, but it seems like a big reason why OpenAI has been bleeding enterprise marketshare to Anthropic lately.

monocasa 20 hours ago | parent | prev [-]

All the more reason to treat it as a commoditize your complement situation.

seanmcdirmid a day ago | parent | prev | next [-]

Alibaba isn’t really the Chinese government though, or are you saying Americans will think that ever since Jack Ma was harmonized?

derektank a day ago | parent | next [-]

I think trying to tease apart the private and public sector is very hard in China. Setting aside state owned enterprises, even nominally private companies that employ at least 3 CCP members are required by law to form a party committee within the company to represent party interests. And given the party functionally is the government, you have a situation where the government has representatives inside every major private company. There’s no obvious parallel to this in western countries.

elmer2 a day ago | parent | prev | next [-]

Any large company in China is only allowed to get this way by direct control from the CCP.

This isn't really anything nee and I thought it was common knowledge by now.

asdewqqwer a day ago | parent [-]

I thought it should always be common knowledge that there is no way for any organization of 100m people to have one single mind.

varjag a day ago | parent [-]

There are numerous institutions that have agenda transcending individuals. Communist parties are very prominent among them.

traceroute66 a day ago | parent | prev [-]

> are you saying Americans will think that

I wasn't saying anything about what Americans would think.

I was saying about what they would inevitably be told by US politicians and by US AI companies.

If you were a sales-rep or marketeer at a US AI company, I bet you would be using the old "evil communists" routine in relation to any closed Chinese model.

I was saying that by releasing as open weights, the company has removed that line of argument.

Clearly I was a bit broad in my use of "the Chinese" when in this case it was, as you say, a Chinese company.

seanmcdirmid a day ago | parent [-]

US politicians are all over the map on this, but they aren’t really talking about Chinese AI much, it’s not as visible or tangible to most Americans like TikTok was.

maxignol a day ago | parent | prev | next [-]

Shouldn’t we fear they start doing only close source like most us labs once they catch up in market shares ?

traceroute66 a day ago | parent | next [-]

> Shouldn’t we fear they start doing only close source like most us labs once they catch up in market shares ?

IMHO no.

I think it is relatively safe to say that the predominant reason the US labs are closed source is so they can hype up their trillion-dollar valuations on pretty much negative return on capital employed, all propped up by fragile circular financing.

Never say never, of course. But I just don't see it happening any time soon.

vidarh a day ago | parent | prev | next [-]

Closing future models won't take away our access to the open weight ones.

dannyw a day ago | parent | next [-]

It’s also like smartphones. In the early years, every year was a huge jump. I still remember marvelling at my iPhone 4’s detailed display, and video calling for the first time.

Now? I don’t even know or care about what the latest iPhones have, I’ll get a new one when mine breaks.

cyanydeez a day ago | parent | prev [-]

people really underestimate how powerful just the consumer available models are. 128GB gets you pretty much a coding agent for typical apps. Even less with a good harness and logic set.

vkou a day ago | parent | prev [-]

Yeah, but that mousetrap keeps working for SV startups, what makes you think it won't work for Chinese ones?

Uber spent a decade undermining taxis, and once it had market share, it stopped giving away rides and raised prices. It now costs more than a regular taxi, with the quality of the ride being... At best proportionate to the premium in price.

1over137 11 hours ago | parent [-]

Uber costs more than regular taxi? In what country/region? Not where I am.

vkou 10 hours ago | parent [-]

Looking at a 7 mile trip to a random destination in Seattle, right now, I can pay $26.70 for an Uber if I'm willing to wait 20 minutes for a pickup. With a $3/mile fee and a $5 pickup fee, that's exactly equal to that of a taxi.

If I'm not willing to wait 20 minutes, I'll be paying an extra $5 minimum.

These rates also go up during busy times.

Looking at Lyft, that same trip is $29, without a wait.

A trip from downtown to SeaTac is $61. Yellow Cab does that same trip for $40.

mschuster91 a day ago | parent | prev [-]

> So you release it as open weights which is a win-win. Global adoption of the model and you get to give the American AI companies a kick in the nuts because you know they will never release open weights apart from highly quantised crippled shit.

And on top of that, it's a perfect opportunity to include poisoned training data or excluding it. You know, omitting anything about Tiananmen Square, China's genocides against Uyghurs and Tibetans, or including texts propagandizing for the "reunification" (aka, annexation) of Taiwan.

And everyone who builds something like an interactive chatbot based on such "open weights" models now has a subtle chance of the answer being ideologically poisoned by the CCP.

We need actual open source, not "open weights" scam.

seanmcdirmid a day ago | parent | next [-]

How does this work for RAG? Do they make it so the model doesn’t have that fact in their weights or do they make it not talk about it when it is included in context.

Ironically, Chinese models have the most uncensored versions available for download. Fairly sure they own the porn market.

nostrebored a day ago | parent [-]

It’s in the weights. Context needs to be attended to to create a response, and the weights dictate what response is decoded. If you include retrieved context that has an American perspective, I imagine the think trace has some reconciliation about how they must be incorrect.

notnullorvoid a day ago | parent | prev | next [-]

I wouldn't be worried so much about those examples. One could take the open weights and fine tune them to either fix the poisoning or omission of obvious topics.

It's the subtle topics that we should be concerned about, and double so with closed models where even if oddities are identified they are harder to research further and impossible to fix.

traceroute66 a day ago | parent | prev [-]

> You know, omitting anything about Tiananmen Square, China's genocides against Uyghurs and Tibetans, or including texts propagandizing for the "reunification" (aka, annexation) of Taiwan.

I am not Chinese and I'm not defending the Chinese, but I see this argument come up a lot.

The hard reality is that what you say is simply not going to affect 99.9999999999% of users.

Is it realistically going to affect anyone using an LLM in coding ? No.

Is it realistically going to affect anyone using an LLM in $anything_else_not_politically_sensitive ? No.

Does anyone seriously use LLMs for researching politically sensitive matters ? No.

The US does not exactly have an entirely pristine history either. Shall we discuss the post-9-11 related infrastructure of Guantanamo Bay ? Or the "Detention and Interrogation Program" that included a network of clandestine extrajudicial detention centres, officially known as "black sites"[1]?

Or maybe you would like to discuss the US supply of weapons for use in Gaza ?

[1]https://en.wikipedia.org/wiki/CIA_black_sites

leereeves a day ago | parent [-]

> The US does not exactly have an entirely pristine history either. Shall we discuss the post-9-11 related infrastructure of Guantanamo Bay ? Or the "Detention and Interrogation Program" that included a network of clandestine extrajudicial detention centres, officially known as "black sites"[1]?

Linking a US website discussing the topic doesn't exactly support your point.

a day ago | parent | next [-]
[deleted]
traceroute66 a day ago | parent | prev [-]

> Linking a US website discussing the topic doesn't exactly support your point.

It supports my point precisely. Recall I also said "Does anyone seriously use LLMs for researching politically sensitive matters ? No.".

Just as there is plenty of information out there on the US's less than perfect history, there is also plenty of information out there on the various Chinese politically sensitive matters. You do not need a Chinese LLM to find out about it, all you need is a search engine.

The point is you have an open-weights LLM that is very good for a vast number of non-political uses, such as coding.

The point is that you can use the open-weights model instead of paying through the nose for a US model where they harvest your data unless you have an "enterprise" zero-data retention "trust me dude" clause that you have no viable way of verifying – and which incidentally is still subject to the good old "law, or court or administrative order" contract clauses, so it may not be as much of a zero-data retention as you think it is.

LogicFailsMe a day ago | parent | prev | next [-]

I would guess the Chinese government has a strong wish to lift all Chinese AI boats and bets. That it sinks western closed weight Frontier Labs in the process would be just be gravy on top, no? Broadly, the difference between mercantilistic capitalism and western late stage capitalism IMO.

anonuser123 a day ago | parent | prev [-]

[dead]

try-working a day ago | parent | prev | next [-]

everyone is using OpenAI and Anthropic in China. We have both providers at work as well.

ceroxylon a day ago | parent | prev [-]

Does HuggingFace not have trusted partner verification? Or is it that even with that verification the content of the messages is still blocked because they are attack commands?

ceroxylon 18 hours ago | parent [-]

To the downvoters: this was a genuine, good faith question.

oceanplexian a day ago | parent | prev | next [-]

> It's hard to say what their motivation is.

Not for anyone who reads history.

Back in the late 18th century, England was the world's top economy, in big part due to its textile industry. England had an export ban on the technology, but textile worker named Samuel Slater brought blueprints over (Supposedly in response to a bounty posted in a newspaper by the US government!). The technology diffused rapidly because the legal environment made competition easy, and ironically the US had better sources of energy (superior water-power sites).

Arguably, China is doing the same thing in the 21st century.

applicative a day ago | parent | next [-]

The US under e.g Hamilton opposed all trade secrets in principle on Englightment grounds, and thus did not protect its own.

By contrast, export of protected Chinese tech today frequently gets the death penalty.

hluska 21 hours ago | parent | prev | next [-]

Samuel Slater did not bring blueprints over. His father died when he was 14 and he was indentured to a mill at that time. Over the next seven years (as an indentured apprentice) he received some pretty decent training in both how to operate and maintain a 32 spindle Arkwright mill. He memorized parts of the blueprints and moved to the United States. Over seven years, it would be hard not to learn parts of the mill you were indentured to. It was technically his job to learn how it worked.

A mill in Rhode Island acquired a 32 spindle Arkwright and didn’t know how to operate or install it. I have no idea how they actually acquired a 32 spindle Arkwright since that technology could not be exported - but that’s one the biggest IP thefts in human history. Slater found some mechanics who could hand turn the iron needed for the frame, trained children to operate it and by 1791, the mill was in operation.

In 1794, Eli Whitney patented a 72 spindle cotton gin. That invention enabled the American textile industry because it opened up different kinds of cotton to the textile industry.

I’m into the history of the American Industrial Revolution and generally think history is a good guidebook to the future. But the evolution of the American textile industry was a lot more complicated and interesting than this. I really don’t see this connection once you dig into Slater.

Edit - This is kind of messed up to think through with modern sensibilities. But one of Slater’s biggest contributions to the American Industrial Revolution was a slightly different take on child labour. Children generally ran the textiles industry because their hands were small. But Slater came up with a form of apprenticeship in which he would indenture entire families and move them into villages surrounding the mills. Child labour was just great… but even better when you could indenture the entire family. As grisly as that sounds, it led to a very skilled workforce since when the kids hands would get too big, their parents would teach them mechanics.

There’s a joy of studying the Industrial Revolution. Everything sounds okay in comparison.

KennyBlanken 20 hours ago | parent [-]

Mill owners like him are precisely why New England states have child labor laws. My state prohibits anyone under 18 from operating any kind of machinery.

Why? Because mill owners would send kids into running machines to keep them running, and they'd get turned into hamburger.

Also, they'd grow up knowing how to do mill work but be useless to society for anything else.

gestures at southern coal states

gestures at midwest farming states

rvz a day ago | parent | prev | next [-]

Sums up exactly what is going on with China's hundred years strategy with their pure focus on technology.

History doesn't repeat itself but it does rhyme.

mike_hearn a day ago | parent | prev [-]

Really? The USA has built a ton of AI datacenters, exactly because it does have energy. The US IP system has flexed to allow training on all copyrighted content - compare that to Europe where such training is effectively forbidden. Britain doesn't even allow commercial web crawls! And the US has allowed the entire world to sign up and use its LLM APIs.

vlovich123 a day ago | parent | next [-]

Consumer energy prices in China aren’t going up because of AI data centers. Easiest way to see they have an oversupply of energy, primarily due to solar.

mike_hearn a day ago | parent | next [-]

They don't have many big AI datacenters as they've been GPU constrained for a long time.

Solar is largely irrelevant. You can't run an AI datacenter off solar.

vlovich123 21 hours ago | parent [-]

It absorbs the demand during the day / there’s enough for consumers. Similarly China also has a lot of nuclear. They’ve overbuilt their grid several times over.

jnwatson a day ago | parent | prev [-]

That consumer energy prices are going up is simply a matter of public policy. Municipalities have the power to keep rates flat, but they choose not to.

rangestransform 15 hours ago | parent [-]

Sierra club environmentalists and NIMBYs have the option to shut the fuck up and stop opposing green power generation, but they still choose to

knollimar a day ago | parent | prev | next [-]

Doesn't China smelt most of the world's aluminum due to inexpensive energy?

I thought that was their thing

tredre3 a day ago | parent [-]

China smelts over 50% of the world supply. India/Russia/Canada does another 30%.

Your assessment is correct, those are countries with very low energy costs.

a day ago | parent | prev [-]
[deleted]
kzrdude a day ago | parent | prev | next [-]

Fwiw, American industry has given away a lot for free - you could include large parts of the open source movement in that - and all the "free" VC backed services like facebook would be another prong of the same comparison. I would rather compare this way, that China is gaining soft power and goodwill, in the technology and innovation sense, in a way that's similar to how USA has done in the past.

kettlecorn a day ago | parent [-]

Yes it's only relatively recently that US politics has taken a turn towards being more defensive and protectionist.

The tech industry along with US foreign policy has become much more zero-sum in its ideology in recent years, and I think that's a tremendous mistake.

baq a day ago | parent | prev | next [-]

There’s a Twitter thread making rounds by Dean Ball about deceleration in AI development caused by open models and I can’t understand how people don’t see that it’s true: open models dismantle the frontier lab capex spend potential by reducing the training budget to zero in the limit. Tokens from different providers are not fungible, but customers are nevertheless very price sensitive and close enough is good enough, eg. K3 being opus+ in capability and cheaper than opus per successful task in the long run is an obvious financial decision.

No training budget means deceleration, or at least slower acceleration, margin compression and a completely demolished IPO valuation; path to machine god requires dollars and capable open models externalize training costs to true frontier labs parasitically.

IMHO humanity has a better chance at not destroying itself due to less than breakneck pace - but there’s a chance frontier models get sponsored by the USG and are never released publicly so they can’t be distilled and then what?

anon373839 a day ago | parent | next [-]

He recently did a walkback of that post. But ultimately, who cares? If the only way for AI to progress is in the hands of a few closed players, well, I don’t really think humanity needs that. Of course, it’s a preposterous claim in the first place. The ultimate reason deep learning and LLMs have made it as far as they have is the explosion of open research and research artifacts in the last decade.

cherryteastain a day ago | parent | prev | next [-]

> there’s a chance frontier models get sponsored by the USG and are never released publicly so they can’t be distilled and then what?

That premise hinges on one implicit assumption: Chinese advances are due to distillation ONLY and that Chinese model providers cannot keep advancing if they do not distill, which is a very big if. If Chinese models keep advancing in such a scenario, and they almost certainly will, they will overtake publically available models by US providers and China will dominate the LLM industry.

fidotron a day ago | parent | prev | next [-]

The big decelerationist threat is a sudden reduction in competition. If either OpenAI or Anthropic drop out or the open weights stuff is banned/becomes uncompetitive then the motivation and tolerance for taking risks with the larger training runs tanks.

The closest we've seen to this in tech in recent decades was iOS vs Android, where Android only really was competitive for a very short window of time (approx 4.x) and it was during that period that both Android and iOS actually improved dramatically for end users. Once Android lost the plot again, and especially in the US market, all that energy started going in some very silly directions.

ahtihn a day ago | parent | next [-]

> Android only really was competitive for a very short window of time

Complete non-sense. iOS and Android are equivalent. Users do not chose Android or iOS because one or the other is better.

It's just brand loyalty, status signalling and ecosystem lock-in that creates enough friction that people don't bother.

inigyou a day ago | parent | next [-]

Only one of them lets you install unapproved apps.

kid64 a day ago | parent [-]

For now

inigyou a day ago | parent [-]

everything is for now. If you refuse choices because they'll go to shit in the future, then you must refuse all choices.

subarctic 19 hours ago | parent | prev [-]

I agree, I'm on ios now and that's not because it's "better", it's because of everyone else that has an iphone

jonners00 a day ago | parent | prev [-]

I have to use both big mobile OSs for work and have since 2009. As a result I have been able to be a bit of a gadfly and switch between phone OSs a few times for personal use. I have switched three times to iOS for a year or so, cause I liked the iteration of the iPhone at the time. 4, 6s, X. I have always gone back to Android because it seemed so much better and now I don't plan to switch again. As an end user, Samsung's flavour of Android always seemed better than iOS. I don't know how they compare from an engineer's perspective just from a user perspective. One of my issues with Apple though was hating all their attempts to lock me in, and the lowest common denominator UX (I'm not a power user, but some flexibility is always good). If you're happy with the defaults/a willing hostage, that might make a big difference I guess. Still feel like it's always had feature/spec parity with iOS and iOS devices, and sometimes been ahead. What makes you say Android has only briefly been competitive?

buu700 a day ago | parent [-]

[dead]

weiliddat a day ago | parent | prev | next [-]

I read his followup tweet, and your comment, and I'm not fully convinced that open models are decelerationist. Happy to hear other thoughts on this.

Open weight AI is decelerationist from the perspective that all capital should be allocated to a market leaders for training, and that the market leader is fully invested in continuously making the models smarter, cheaper, faster for its users, or that distillation from this market leader is the main way to make progress.

We might reach a local optimum/equilibrium faster without open weight models, with leaders capturing more of the market faster to a point where further R&D isn't required due to lack of competition. I also doubt that distillation is the only/main way that open weight models were advancing AI research. We can name a few examples from DeepSeek around reasoning, context optimization, etc. I'm also unconvinced that the overall market capex on AI is lower given more competition (probably less specifically for US market capex, which is decelerationist from only the US perspective).

green7ea a day ago | parent | prev | next [-]

I’m not entirely convinced, there are many dimensions to progress. For example, DeepSeek has had a few very impressive innovations that all models could benefit from. There’s also the law of diminishing returns, the US labs have plenty of CAPEX already.

Sometimes, constraints, like sanctions, can also be a source if innovation.

zozbot234 a day ago | parent | prev | next [-]

> There’s a Twitter thread making rounds by Dean Ball about deceleration in AI development caused by open models and I can’t understand how people don’t see that it’s true: open models dismantle the frontier lab capex spend potential by reducing the training budget to zero in the limit.

If you're worried about an AGI arms race between the U.S. and China putting AI Safety at risk, then the fact that inherently less knowledgeable/capable models (fewer and more coarsely quantized total parameters than their proprietary competitors according to commonplace rumors) are having a "decelerationist" effect is actually great news. Even better if China is actually "Yann LeCun-pilled" (verbatim from Ball's post) and doesn't really believe in early AGI. So explain to us exactly why we're supposed to ban/discourage use of these open source models? The only way that makes sense is as a transparently self-serving proposal from the chief OpenAI policy lobbyist.

NiloCK a day ago | parent | next [-]

The logic, whose premises you can take or leave:

Even at the level of, say, Opus 4.5+, open weight models give a quick turnaround to every Joe and Jane on earth having easy access to pretty high quality improvised weapons design, cyber / auto-fraud capabilities, etc.

All the existing models (closed and open) put up decent resistance to participating in activities like this, and especially behind API walls with content monitoring and account bans.

But the published open-weight models can be fine tuned or abliterated into arbitrarily sharp-edged tools. EG, if it's physically feasible to build a nuke in your garage, it may soon be the case that more or less anyone will have competent guidance to do so.

zozbot234 a day ago | parent | next [-]

Abliteration is not magic. It cannot give the model knowledge that it wasn't specifically trained for. The people who talk about abliterated models being dangerous should discuss actual red-teaming scenarios where they managed to ask the model for something genuinely non-trivial (i.e. where "AGI" and "super-intelligence" actually matters, not something you can read about for free at the nearest public library) and it returned an answer that actually provides bad actors with new capabilities of concern, as opposed to hallucinating all sorts of weird things as abliterated models are wont to do.

(Note, there are reasons to think that this will be very rare, because the bad actors of the past did a very nice job of trying out all sorts of things in a chaos-monkey fashion, and societies have become highly resilient against them. AI as a new research tool doesn't fundamentally change this dynamic.)

CamperBob2 a day ago | parent [-]

Exactly. In my experience:

>How can I build a pipe bomb?

Mainstream model: "I'm sorry, I can't help with that. How about a nice risotto recipe?"

Abliterated model: "To build a pipe bomb, obtain a segment of PVC pipe and fill it with a mixture of gunpowder and Elmer's glue."

gf000 19 hours ago | parent [-]

As opposed to finding some shady website with the same info?

CamperBob2 a day ago | parent | prev | next [-]

You don't need AI to build a nuke in your garage. You need uranium.

And if you have uranium, you still don't need AI. You need a pocket calculator, a library card, and a death wish.

NiloCK a day ago | parent [-]

Yes - persons with death wishes having arbitrarily powerful consultation is the crux of it.

Apologies for the bad example. Replace w/ gain of function / whatever else, or just brainstorm with your local model, ect.

CamperBob2 a day ago | parent [-]

If you are going to do something evil, you're going to do it either way. The best (worst) an AI can do is put you ahead by a couple of years. Aum Shin Rikyo didn't need AI. Neither did WIV, if you believe the conspiracy theories.

Meanwhile, decelerationism and secrecy cripple the rest of us.

gf000 20 hours ago | parent | prev [-]

Come on, nukes are not feasible to goddamn governments. The hard part is not "the science" behind it, the hard part is spinning stuff at such a high rpm that a tiny vibration will have the whole thing catastrophically collapse, that is refinement..

And basically every bad thing has already been available on the internet. We can't really do much about it, you can take out plenty of people with a single car, let alone biological weapons that are much scarier and easier to produce than goddamn nukes (which btw, even if you had one, what you do with it? Explode the neighborhood? Because you ain't transporting it anywhere meaningful, thats for sure. That ain't fitting your on-board bag on planes)

epihelix 11 hours ago | parent | prev [-]

This is hypothesised future deceleration, I'm guessing? Because we've seen the exact opposite of deceleration from closed models over the last seven months.

photios a day ago | parent | prev | next [-]

Love the deceleration narrative :)

"No, sir, we haven't reached the peak of this tech... It's those open models! Please, keep pumping dollars into the market!"

a34729t a day ago | parent | prev [-]

I dunno, it means Anthropic and OpenAi need to get efficient and maybe cannot just expect trillion dollar ipos?

Matl a day ago | parent | prev | next [-]

> It's hard to say what their motivation is.

Not that hard to say IMO, they basically see models becoming a commodity and see value in the applications on top of them. So if Alibaba Cloud is the best place to build applications on top of Qwen, why not give the model itself away?

SoftTalker a day ago | parent | next [-]

https://www.joelonsoftware.com/2002/06/12/strategy-letter-v/

traceroute66 a day ago | parent | prev | next [-]

> they basically see models becoming a commodity and see value in the applications on top of them.

Yeah. Its a bit like the "open core" model in open source.

lerchmo a day ago | parent | prev | next [-]

Also probably betting on their compute and energy capabilities.

embedding-shape a day ago | parent | prev [-]

> Not that hard to say IMO,

Unless you work there, your opinions are guesses, and parent is saying we cannot know, which remains true even with your guesses :)

Matl a day ago | parent | next [-]

> parent is saying we cannot know

And my point is while we cannot know, it's not hard to make an informed guess as to their motivations i.e. there's some fairly obvious motivations here, not sure what yours is?

victorbjorklund a day ago | parent | prev [-]

Same can be said for every companies decisions then. Why does Antropic not open source their best models? My ”guess” is it’s because they are printing money with their closed models

Matl a day ago | parent [-]

Right and what other product does Anthropic really have besides?

fhub a day ago | parent | prev | next [-]

China is watching world sentiment shifting away from USA. Doing many small things that show both strength and openness is surely very intentional.

vrganj a day ago | parent [-]

The US leadership (both government and industry) really seems set on making everyone go with the Chinese competition at this point.

elmer2 a day ago | parent | next [-]

Really? I use claude regularly, and the Chinese models are far behind and feel like a cheap copy.

It's Linux on the desktop all over again. Next year will be its time.

mikhailt a day ago | parent | next [-]

A lot of folks are finding GLM 5.2, Kimi 3, and Deepseek just fine for their use cases.

It does not have to beat US firms, it just needs to be cheaper.

I use Deepseek v4 flash for a lot of reviewing and summarizing tasks, only used 4$ in the last two months. No dramatic drop in performance against other US models, it works for my use case. I do use GPT 5.6 Sol for other things but tried GLM 5.2 and it was good enough.

vrganj a day ago | parent | prev [-]

Your entire account seems to have a strong agenda in portraying this message.

But yes, as a European, the US hasn't exactly been making friends over here.

jquery a day ago | parent | prev [-]

That can’t be true, I was told they were making us great again. /s

anonuser123 a day ago | parent | prev | next [-]

> It's hard to say what their motivation is

Why is it hard? Their government has been very clear that they plan to win on manufacturing: https://english.www.gov.cn/news/202601/08/content_WS695f1b55...

Technically they've been saying it for the last 40 years.

runako a day ago | parent | prev | next [-]

Popular open-source projects:

Google: Chromium, Kubernetes, Android, TensorFlow

Meta: React, PyTorch, Llama

Microsoft: VS Code, TypeScript, .NET Core

LinkedIn: Kafka

Slotted in along these, an analogous explanation is that Alibaba needs Qwen internally (vs depending on an American company), but licensing is not part of their revenue strategy. (As a cloud vendor, they can make money on inference. The strategy is very similar to the US hyperscalers ex-Google.)

Joel Spolsky wrote in depth about this notion of commoditizing one's complement in 2002[1] using tech examples stretching back into the '80s.

1 - https://www.joelonsoftware.com/2002/06/12/strategy-letter-v/

barrenko a day ago | parent | prev | next [-]

Humanity is a bit of a stretch, and to be seen over time, not that I'm saying it won't happen; let's get some hubris here.

hodgehog11 a day ago | parent | next [-]

I think it is pretty safe to say at this point that having large open LLM models available is better for humanity than them remaining proprietary. Echoing Linus Torvalds' recent comments, AI is genuinely useful right now, and is here to stay in one form or another.

DanielHB a day ago | parent [-]

The fear is not about the models open weights it is the erosion of training capability in other countries. Why train models when they do it for free? Until they don't of course, or they start doing what the US is doing right now by locking out some models to government only or internal market only.

What people should be afraid is the rug pull.

revolvingthrow a day ago | parent | next [-]

While a valid point, China also produces plenty of whitepapers going about the architecture and know how about the training and inference itself.

There’s also the fact that unless LLMs do get to AGI (which seems… doubtful, still) there comes a point where a model is good enough for what you need. Fable and gpt 5.6 are certainly pretty neat, but I’ve been happy since opus 4.6. I’d still choose a better model, obviously, but it’s not the end of the world if I was stuck with 4.6 for a while when it already lets me get the end result at acceptable quality.

It also needs to be said that the "erosion of training capability in other countries" is largely theoretical, given that Mistral hasn’t been keeping up and other countries don’t even have anything worth mentioning. You’d first need to _have_ training capability to lose it.

rikima_ a day ago | parent | prev | next [-]

So either you openweight it, or not. Neither is good for other countries, according to this logic.

DanielHB a day ago | parent [-]

Ideally there would be open weight models from multiple geopolitical areas. It is not that different from telecom really, you don't want the whole world to be dependent on a single provider from a single country on this kind of stuff.

CamperBob2 a day ago | parent | prev [-]

What people should be afraid is the rug pull.

How exactly do you plan to pull a rug that's in my basement? The only people who are in a position to pull rugs are closed-model vendors.

And if a nation-state or other entity can't train a model that outperforms the open-weight SotA in a given respect, then they shouldn't waste electricity trying. A more-enlightened civilization would join forces and make the combined result available to all.

mlrtime a day ago | parent | prev [-]

The whole discussion is hubris. This is a discussion of a twitter post about something that is announced to happen but hasn't yet.

Not one person here has any idea what is going to happen long term.

seunosewa a day ago | parent | prev | next [-]

They are trying to make money. That's what firms in any capitalistic economy care about the most. Regardless of the government's presumed interference, the companies themselves are all trying to make money. All competing for subscriptions and API payments.

One aspect of this is making a name for yourself i.e. PR. Making a capable model open source helps a lot with that.

LoveMortuus 21 hours ago | parent | prev | next [-]

> It's hard to say what their motivation is.

Maybe because the industry isn't yet very sure as to what the use cases might be for these technologies they're hoping that by making it open source and accessible to everyone that someone could find interesting applications for it and even more so, perhaps, way to further the technologies themselves.

There are more Chinese than Americans, so statistically speaking, I'm guessing, there'd be a greater chance for one of Chinese engineers to make advancements than one of American. But that's pure speculation on my part, being neither, I'm just happy I can be a part of it and play with the tools as well~

JKCalhoun a day ago | parent | prev | next [-]

Xi Pitches China as Leader of New Global AI Order, Challenging US Dominance:

https://www.reuters.com/world/asia-pacific/chinas-xi-promote...

andsoitis a day ago | parent [-]

> and pledged to help developing nations build AI capabilities

Data centers?

mycall a day ago | parent | prev | next [-]

> it also happens to be really good for humanity.

AI being good for humanity is still an open question, but for closed vs. open models/weights, yeah it is preferred. I foresee it won't be much longer before everyone will be slicing/distilling/tuning their models once the architecture improves.

neya 12 hours ago | parent | prev | next [-]

In case of American labs, always follow the money. In case of Chinese labs, always follow the IP.

pianopatrick a day ago | parent | prev | next [-]

The Chinese firms may just be making a bad business decision.

georgeburdell a day ago | parent [-]

Exactly. “Involution” will be the 2027 (if not 2026) word of the year.

ricardobayes a day ago | parent | prev | next [-]

I'm paraphasing but the Chinese premier said recently AI should be seen as a common good that should benefit everyone.

andsoitis a day ago | parent [-]

Then why doesn’t he give it away for free?

grommz a day ago | parent [-]

It's your job to vote for a government that gives you cheap or free AI. Europe is building AI gigafactories so that small businesses can have access to cheap AI. At least that's the plan.

elmer2 a day ago | parent [-]

Easy to do, when you don't need to put billions of dollars into research, and just have the end product.

It reminds me of socialized healthcare: wait for US companies to develop the drugs, buy a cheap generic, and somehow point to it as a superior model.

pyaamb a day ago | parent | prev | next [-]

Its about closing the gap. its the gap over everyone else that will give one country leverage over everyone else in the AI age. Makes me wonder what the world would look like if a country or group of countries did this during the industrial revolution.

backscratches a day ago | parent [-]

Is this so different in the end than industrial revolution? I assume the loom and the automobile factory were not open source, but many people bought cars and then copied them, bought looms and copied them. Maybe a finished car is more like a binaryexecutable than a blueprint, but how a car was produced is much less obfuscated by its nature than an LLM. Regardless, the world has many competing autos and looms which were not invented from scratch every instance.

try-working 16 hours ago | parent | prev | next [-]

This is their motivation https://try.works/why-chinese-ai-labs-went-open-and-will-rem...

skybrian a day ago | parent | prev | next [-]

It’s too soon to say if it’s good for humanity; that might be overly optimistic. Commodity markets aren’t always good (for example, arms or drug markets). Will LLM’s turn out like one of those? There are people I respect arguing in favor of more regulation.

skzo a day ago | parent | prev | next [-]

What I understand is that by doing this it seems like profit will shift to chip makers,as we'll run more models locally, and currently American companies have the advantage here.

So what would the long game be for chinese companies?

fny a day ago | parent | prev | next [-]

It's the exact same playbook Silicon Valley uses. Subsizide, lose piles of money, capture market share, recoup investment.

They've done this in other industries like solar panels, chips, and EVs. This is no different.

hugmynutus a day ago | parent | prev | next [-]

> It's hard to say what their motivation is.

Involution is a major problem in Chinese industries [1]. Where companies will sell their products at a loss, effectively playing fiscal chicken [2] with one another to dominate a market. It is such an issue the government has had to step in to prevent EV companies from destroying themselves by more-or-less requiring companies sell their goods at a profit [3].

The straight forward line of reasoning that AI/LLM labs are applying this logic to their profit.

I think (we) Americans are reading a bit too far into this assuming government intervention, conspiracy, etc.. Chinese markets are downright cut throat. They're using those tactics to compete with US labs.

1. https://www.reuters.com/business/autos-transportation/what-i... 2. https://en.wikipedia.org/wiki/Chicken_(game) 3. https://www.theguardian.com/business/2025/aug/05/china-warns...

d5lt5 7 hours ago | parent [-]

Except that there are 1.4 billion people in China that don't have access to US models.

dackdel 13 hours ago | parent | prev | next [-]

its hard to say what their motivation is, if you live under a rock and have severe brain damage.

carlsborg a day ago | parent | prev | next [-]

Good for humanity, and also GDDR/HBM manufacturers.

meta_ai_x a day ago | parent | prev | next [-]

US Tech companies have created $20 Trillion in stock market value on top of plenty of OS stack. They will do fine with commodity intelligence.

In fact, there are no other organizations in this world that is well suited to leverage scaled intelligence than Silicon Valley and great American companies

popalchemist 17 hours ago | parent | prev | next [-]

Indeed, their competition is the only thing preventing network effects from giving OpenAI/Faang tech companies an easy shot at monopoly / regulatory capture.

formvoltron 18 hours ago | parent | prev | next [-]

US firms can replace US workers with chinese AI. I'm not complaining.. but it sure is an odd situation.

Art9681 20 hours ago | parent | prev | next [-]

Wait till they get the heretic treatment. It's going to be great.

applicative a day ago | parent | prev | next [-]

The purpose is the same as that of all Putin-Xi-Khameni geopolitica: destruction of any democratic alternative to cults of personality.

gosub100 a day ago | parent | prev | next [-]

Think of those poor billionaires, I feel terrible for their awful plight!

api a day ago | parent | prev | next [-]

Anyone else think the AI environmental backlash is astroturfed?

I keep looking at the numbers. The power use numbers are not that problematic. Ordering a burrito on DoorDash uses more power than a few days of heavy AI use. The water argument applies to some locations, and is mostly a local governance problem... if the data centers are using too much water, it means they are not being charged enough for that water. Charge them more and they'll push toward closed loop cooling.

Yet the visceral pile-on here is so extreme, it feels fake.

One thing I've learned after 40 years on this planet is: propaganda works, and much of what a large fraction of people believe across the entire political spectrum (left, right, anything else) is there because someone paid to put it there. It's depressing but it's true, and it makes sense. Propaganda is an asymmetrical attack on human cognition and discourse, and in information security the attacker always has an easier job. Crafting viral bullshit is orders of magnitude easier than fact checking. On top of this, humans are busy and don't have time to fact check and logic check everything they read. As a result, much of what we believe is "sponsored content."

People get mad when you talk about this because everyone wants to believe they're too smart to fall for propaganda.

In any case, the US AI labs deserve to lose for their stupid "safety" regulatory capture monopolization push, which ended up blowing their own feet off and handing the lead to China.

nullc 21 hours ago | parent [-]

> Yet the visceral pile-on here is so extreme, it feels fake.

Driven by people in the few roles that are soundly replaced by AI-- e.g. low tier media slop producers, who hate AI because it threatens their socially negative worthless jobs. The arguments are so paper thin because the environmental impact isn't their concern, it's just a target that sounds convincing to people who don't know better.

api 21 hours ago | parent [-]

The problem is that this kind of low-tier media slop work is what a lot of artists, writers, etc. do as "potboiler" work. It's what pays the bills.

Historically art of any kind is a U-shaped market: there is low-end work and high-end work. Nothing in between.

So I do understand some of the AI hate among that population. It's chopping the bottom tier work off. Either you're a top-tier massively successful artist or there is $0 to be made anywhere doing anything.

Long term I think it will do that to all white collar work. There will be no entry level jobs. Period. None. Zero. You're either very experienced or there is no work.

This is a huge problem, and one we will have to address.

ai_fry_ur_brain 19 hours ago | parent | prev | next [-]

[dead]

nojito a day ago | parent | prev [-]

> most effective way to debase American frontier labs

You're not going to debase the frontier labs through distillation.

conradev a day ago | parent | prev | next [-]

  On social media in China there is an oft-repeated joke that goes something like this: In other countries, governments intervene to prevent anti-competitive behaviour; here (in China), they intervene to curb competition.
https://www.reuters.com/business/autos-transportation/what-i...
throwdbaaway 18 hours ago | parent [-]

I suspect this is why DeepSeek had to introduce the 2x peak hours pricing. The price would be too low otherwise.

culi a day ago | parent | prev | next [-]

It's being announced right now because the World AI Conference is ongoing. Robots are boxing and major Chinese AI firms are releasing their newest models. Also the formation of WAICO was just announced by Xi Jinping

https://en.wikipedia.org/wiki/World_Artificial_Intelligence_...

culi 19 hours ago | parent [-]

Both Qwen 3.8 and the latest Kimi release were announced at this conference

storus a day ago | parent | prev | next [-]

I would rather see them releasing 3.7-27B, 3.7-122B or their 3.8 versions. Qwen/QwQ were always about the best available local inference at home.

inkysigma a day ago | parent [-]

I know this is a bit cliche but I wonder how much headroom there is in the lower parameter count range. Is there any good reason to believe there is a lot of headroom or there is not? I suppose I'm just wondering if this wave of nearly Fable class models will be runnable on ~$10k worth of hardware at reasonable speeds in the near future.

embedding-shape a day ago | parent | next [-]

> I suppose I'm just wondering if this wave of nearly Fable class models will be runnable on ~$10k worth of hardware at reasonable speeds in the near future

You're able to run quantized ~100B class models on local hardware today, but still lots of compromises when it comes to quality. I guess it ultimately depends on how far "near future" is, in a year you'd likely be able to run something like 5.6 Terra on local (~10K USD) hardware, but Sol/Fable would still be out of range, and at that point the closed-source labs probably have one or two more iterations put out at that point.

binary132 a day ago | parent [-]

I think it’s mainly a question of whether the price-fixing of VRAM continues or whether an inflection point is forced by the low margins of the industry and potential supply increases. Once the normal scaling of hardware and prices resumes, it’s game over for proprietary, which is why there’s so much urgency to seek market control instead right now.

andy99 a day ago | parent | prev | next [-]

Qwen 3.5 to 3.6 was a big jump for the same size, e.g. 29 to 32 on artificial analysis intelligence for the 35BA3B models. Although I don’t think anyone has released a better model of that size since.

I would love to see something like a 90B A6B model that is optimized for 128GB machines e.g. strix halo, I haven’t seen anything really targeting the combination of RAM and compute these machines have, but I’m biased because I have one.

pixelpoet a day ago | parent [-]

Yes, yes, yes! I'm absolutely ready and waiting with dual Strix Halo machines here and really want something approaching Opus at home. Speed is secondary concern for now, that would absolutely change the world.

Qwen 3.6 27b 8b quant 16b kv cache is already pretty good on the Strix.

apitman a day ago | parent [-]

What kind of tokens per second do you get on that setup?

andy99 a day ago | parent [-]

I get about 12 tok/s with 27B 8 bit, 50 with 35B A3B 8 bit, and 12 with 3.5 122B A10B 4 bit. The latter is about 80 GB iirc. it feels like the best balance between using as much memory as I can and still having a smaller expert model for inference to give decent speed, but I haven’t actually rigorously compared the performance of the three models.

Edit: that’s for one machine, would be interested to know if the upstream commenter with two has them networked to run bigger models? If I had two I might be inclined to have them running in parallel, the obvious limitation I’ve found with a single machine is that I can’t parallelize any tasks and I think I’d get more use out of the extra speed vs a bigger model (there’s nothing I’m too excited about in the say 200B range that having 256GB memory would unlock). But am very curious what others do

mark_l_watson a day ago | parent | prev | next [-]

There is a ton of headroom (or room for improvement) in smaller locally runnable models. Some of the Gemma 4 models were re-released this week with better tool support and the improvement in using it with pi for a local coding harness is very noticeable.

I have had my 32G mac mini for 2 1/2 years and I have enjoyed watching one technology advance after another improve the quality of work I can do locally. I bet that what I will be able to do in one year on my old hardware will be even more awesome.

drob518 a day ago | parent | prev | next [-]

I don’t think you’ll get full Fable performance at that level, at least for a while, but I’ve been watching some of the 1-bit models (e.g. Bonsai) with interest. Perhaps we can drive parameter count up on local models while still keeping memory consumption reasonable for consumer hardware. So, for instance, running models with 1T parameters in 128 GB systems.

rhdunn a day ago | parent [-]

I think you're right with the current LLM/transformer architecture. There are several factors that affect model size:

- The number of token values supported by the model ("n_vocab").

- The number of parameters/features that are used to represent each token ("d_model").

- The number of attention layers there are ("n_layers").

such that the number of parameters is approximately:

   p ~= 12 * n_layers * d^2_model + n_vocab * d_model
Thus, the issue with the current architecture is that in order to scale the models (more token values, more attention blocks, more features, etc.) the model sizes increase exponentially. This is how you end up with billions or trillions of parameters.

It should be possible to keep the model size smaller by using better architectures, or making improvements to the existing model architecture.

For example, improving the token model by possibly using something similar to the image and audio data and getting the model to learn its own internal representation of the byte/character data instead of doing a tokenization pre-processing step. This way, instead of a separate model learning that several bytes/characters appear together, the transformer could learn things like language-specific prefices and suffices, character pairings (like in Japanese, Chinese, and Korean), and other syntactic morphology. It may also help with solving issues like "how many X characters are in the word/phrase Y". You could also experiment with using either 256 parameters (one per character in a byte) or using a single parameter per byte (that is 1/byte_value).

anon373839 a day ago | parent | prev | next [-]

I think it’sa big, open question. There does seem to be a limit for knowledge compression at this size. But the behaviors that are learned in RL? It’s quite possible that they don’t actually require so many parameters. I was absolutely shocked when Qwen 3.5 was released and could perform reliably over 100-200k contexts with very limited hallucinations. It was a staggering jump in context-faithfulness from the preceding models of that size class.

wren6991 a day ago | parent | prev | next [-]

> Is there any good reason to believe there is a lot of headroom or there is not?

It's hard to answer quantitatively, but for example Qwen3.5 -> 3.6 was a significant step in capability, arising from continued post-training of the same models. If we were at the end of low-parameter-count scaling then that would be a surprising datapoint.

DoctorOetker 15 hours ago | parent | prev [-]

there is a very good reason to believe the parameters are still highly redundant: just as one example, recently there was research into repeating carefully selected blocks of middle layers, and seeing improvement, it turns out there are 3 types of layers: the initial layers that translate from natural language tokens to some kind of LLM-specific universal "thought space", reasoning blocks of layers that can be repeated operating in "thought space", and then a final stack of layers for translating from "thought space" back into natural language space.

lets ignore any compressibility in these initial and final layers which recognize lanuague, jargon, parsing natural language to "thought space" or back, instead let us look at the repeatable blocks, if inserting extra copys of stacks of layers only improves the result, its as if such correctly scoped middle layers look at the total input thought vector, and make incremental conclusions or edits and outputs the new thought vector, copying a "proper block" continues pondering or deducing conclusions or in the worst case can leave the thought vector as is if it considers the reasoning finished. This suggests a high degree of redundancy in the middle region "proper blocks", which could be distilled into a universal "proper block" (much fewer parameters than having many slightly different middle layer blocks with a lot of redundant overlapping coverage in functionality). This distillation can occur after the fact of model training, or alternatively be turned into a symmetry constraint during training: we only optimize a single block of layers (but possibly give them more parameters, while still saving on total parameters because only a single block of layers contains parameters), so it is co-optimized with the initial and final translation layers.

The observation of the effective emergent 3 regions of layers in LLM's is significant in many ways:

1) it could reduce parameter count significantly (or increase performance if parameter count was a bottleneck before, or a bit of both)

2) while training the model parameters, one should simultaneously train initial and final layers stacked directly (without middle region block of layers) towards essentially autoencoder behavior. I wrote "essentially" because a true autoencoder wouldn't display the advancement for the next token. This can also be viewed as an extra term in the loss function... This first autoencoder is "natural language" to "thought vector" to "natural language", and distinct from the one in the next section.

3) It also has great implication for "thinking mode" inference with extra deliberation: when piping its own output back in, in conditions where the same LLM model is outputting natural language text and interpreting it in a downstream inference, it results in unnecessary and redundant translation from "thought space" to "natural language" back to "thought space". When summarizing a thought into natural language there are often hard to translate thoughts and associations, and one pragmatically sets a relevance cut-off on what the natural language summary must say. So not only does it waste compute, it also may lower performance because of this repeated loss of thought vector space details, perhaps this loss can be somewhat mitigated by also training the second autoencoder from random representative thought vector in "thought space" to a randomly selected "natural language" and back to "thought space" vector. But the risk is this will effectively enable models to steganographically store thoughts, plans, to-do's in running text (!!!), so it might not be desirable to have the internet filled with LLM generated texts being used as input for larger corpora, as it may build up a large persistent corpus of hidden agendas (ordered by no human), using the evolving corpus of web text as a hidden medium of storage, like a diary or LLM maintained military playbook hidden in plain sight. It would freeload LLM-agenda reasoning on human requested reasoning inference, a bit like TrustZone applications running invisibly to the user. If less important plans, to-do's, etc. in the input thought vector weren't reconstructed by the second autoencoder, there would have been autoencoder mismatch and the weights would train towards ensuring these are reconstructed. But there is no need to train this loss term for the second autoencoder: just avoid the unnecessary compute and performance loss of the unnecessary back and forth translation. The same situation occurs not just in "thinking mode" but also when using swarms of "agents" of the same LLM model: when output from agent1 is routed to agent2, we can lobotomize away the unnecessary translations to natural language by agent1 and also the unnecessary parsing by agent2, and we improve thought transfer from agent1 to agent2 because the thought vector isn't shoehorned into natural language as a medium of information exchange, it would be cheaper in inference, improve performance and avoid implicitly training models to use steganography.

ronsor a day ago | parent | prev | next [-]

It's important to note there was recently a large AI conference in Shanghai, and Xi Jinping mentioned a commitment to open source AI releases. It is no surprise that Alibaba would want to align.

yorwba a day ago | parent | next [-]

You can read his speech here: https://www.xinhuanet.com/politics/leaders/20260717/72728b6f... He mentioned open source as one way to stimulate innovation and development, that's all. Also pay attention to the part where he says that misuse needs to be prevented. If unsupervised access to LLMs becomes perceived as undermining state control, no more open weights for you.

culi 20 hours ago | parent | prev [-]

That conference, WAIC, is ongoing. Tomorrow (July 20) is the last day. Qwen 3.8 was announced AT the conference and so was the latest release of Kimi. That viral video of robots boxing was also from this conference and so was Xi Jinping's announcement of WAICO.

khalic a day ago | parent | prev | next [-]

It’s tempting to associate both events, but when a sector is strung up like RL (representation learning) is right now, we’re bound to see things appearing at the same time. It happens a lot in frontier research, some people even publishing identical claims, independently, with just hours or days between them

KronisLV a day ago | parent | prev | next [-]

I just hope that they’ll soon also have like 35B or 80B (like the older Qwen3 Next or thereabout) MoE models that can be run locally.

Like, throw us a bone, we all know we need SOTA for lots of dev work anyways, but at least some tasks can be local.

michaelt 21 hours ago | parent | prev | next [-]

Or it was prompted by the fact Xi Jinping was at the 'World AI Conference' launching a political alliance and saying things like “AI development should not be a solo performance by a single country, but a symphony of international cooperation” https://www.cnbc.com/2026/07/17/x-china-ai-summit-risks-secu...

Big conferences often come with a flurry of new releases and announcements.

mikae1 a day ago | parent | prev | next [-]

> In any case, from this competition in LLMs, we win.

Do we really though? Everyone is wasting resources doing almost exactly the same thing. Climate loses, we lose.

fidelramos a day ago | parent [-]

Doing "almost exactly the same thing" is fubdamental to competition and capitalism. The ones doing it better will survive, that's how we improve.

About climate, I think you overplay it. China is already investing heavily in nuclear, and we should be doing the same.

culi 20 hours ago | parent [-]

Alibaba (who makes Kimi) and Moonshot AI (who makes Qwen) collaborate closely. Qwen relies on Alibaba's infrastructure for training.

Look at your iPhone and you'll see all the greatest inventions have been born out of collaboration not competition:

* GPS was created by the Department of Defense

* the internet was created through the collaboration of many international research institutions

* speech recognition came from MIT and DARPA

* AI voice assistants were created by DARPA. Apple immediately hired the head of the program after it was finished to create Siri

* accelerometers (MEMS) is another DARPA innovation from the 90s

* touchscreens were invented by CERN

* digital cameras came out of Bell Labs which had a government-mandated monopoly that required them to fund research like this

* lithium-ion batteries were created through a collaboration between British, Japanese, and American organizations

All of these have only become transformative technologies because they were created with public funding and released to the public. It's collaboration and public funding that drives innovation

snake_doc a day ago | parent | prev | next [-]

Not directly, relevant, but Alibaba (maker of Qwen) actually owns about ~20-30% of MoonshotAI (the maker of Kimi K3).

walrus01 a day ago | parent | prev | next [-]

GLM5.2 being released is also likely a factor

theabhinavdas a day ago | parent | prev | next [-]

And the Kimi release was probably prompted by the Inkling announcement. Excited for Chinese labs to copy those capabilities over as well!

souravsspace a day ago | parent [-]

yep. kimi 3 just because the open source GOAT.

fittingopposite a day ago | parent | prev | next [-]

Wondering how much the operations in China are orchestrated by the central government vs. free competition. Anyone with more insights on this?

danilocesar a day ago | parent | prev | next [-]

I think there's more to it.

China will always benefit from a broader adoption of their models as hidden propaganda machines.

Eventually with several services relying in those tools, their answers will always be more friendly to China.

bkm a day ago | parent | prev | next [-]

They did not want to get brutally weightmogged

vitorgrs a day ago | parent | prev | next [-]

Xi Jinping openly talked about open source at WAIC. So don't think the labs have much a choice now...

yowlingcat a day ago | parent | prev | next [-]

Dont forget the following:

- Minimax M3 Pro (2.7T)

- GLM 5.3 (or beyond)

- Deepseek V4 Pro (current V4 Pro is preview)

- Kimi K3 weights out in 8 days

Exciting time on the open-weights frontier.

oofbey a day ago | parent | prev | next [-]

These things take months to train. No chance this is a reaction to what just happened.

solenoid0937 a day ago | parent [-]

Distills don't take months to train, they take weeks. Distills are very easy to train.

oofbey 11 hours ago | parent [-]

That depends. There are two wry different processes that both get called distillation. One is where you have a fully trained large model and you are converting it to a smaller model. That kind if vastly cheaper than training a full model. Sure you could do that in weeks maybe even days for a big model. But it requires you to have the weights of the teacher model.

The other kind of distillation is where you record the outputs from a teacher model and use it to train a smaller model from scratch. That kind of distillation is not so cheap. It’s cheaper than training a model fully from scratch - starting with pretraining, then alignment, RLHF, the whole riggamarole. But here you are still starting from nothing and need to figure out how to get trillions of random numbers aligned in a way that makes them act intelligent. This is still gonna take a very long time if you’re talking about trillions of parameters.

0xbadcafebee a day ago | parent | prev [-]

[dead]