| ▲ | traceroute66 a day ago |
| > How does this explain open weights? They could easily take the same closed route like their American friends Because they are playing the Americans at their own game. What is the first thing an American company would do ? Spread the old American classic FUD ... "you can't used this closed tool because its run by the communists", right ? So you release it as open weights which is a win-win. Global adoption of the model and you get to give the American AI companies a kick in the nuts because you know they will never release open weights apart from highly quantised crippled shit. The Chinese are also playing the long game. The gradual rebalancing of the world from the US-centric model of the past. If releasing models as open weights is part of that long game, then so be it. |
|
| ▲ | Grombobulous a day ago | parent | next [-] |
| I think most of us that will claim to understand China are going to end up being wrong, unless any of us live there or grow up there. There’s a saying about China I have heard from ex-pats: the more you know about China, the less you know about China. The point of me bringing that up is to say that what follows is really just my best guess: If I were to judge from China’s approach to hardware, I think that the companies releasing open weight AI for free aren’t as worried about giving away too much as the West tends to be, just like a factory making robot vacuums isn’t worried about other factories copying their methods. For one thing, Chinese firms are spending an order of magnitude or two less money training their models. They have pursued efficiency in a way that Western companies with insane capital systems haven’t bothered, and in some cases they’ve had to given their limited access to bleeding edge hardware via export restrictions. My best guess is that more important than that, Chinese companies don’t see the open weight model itself as the value add. At this point I don’t think we pay for Claude specifically for the model. If that was the case then we’d all be using cheaper/free models from China as they are the best model value. Basically, any time we decide not to use Fable or Opus to save costs, what’s the point of spending more than competing models to use Sonnet and Haiku? The real reason we are using Claude is for the SaaS aspect of it. It has a toolchain, a friendly interface, and a bunch of integrations with business applications. In this respect, it’s somewhat surprising that Western AI companies don’t publish open weight models more frequently. The struggle of setting that up yourself and figuring out which hardware can run it should be an advertisement for Claude and the rest. |
| |
| ▲ | KerrAvon 20 hours ago | parent [-] | | >The real reason we are using Claude is for the SaaS aspect of it. It has a toolchain, a friendly interface, and a bunch of integrations with business applications. That is a very thin moat, though. There's nothing you can do with, for example, Claude Code + Opus 4.8 that you can't do with your own custom harness running API-level Opus 4.8, which means that if you can afford the hardware (the moat for running any SOTA model) you don't need to pay Anthropic anymore. I'm not saying they shouldn't, but I understand why they don't. | | |
| ▲ | Grombobulous 19 hours ago | parent | next [-] | | You might be right that it’s thin, but it seems like a big reason why OpenAI has been bleeding enterprise marketshare to Anthropic lately. | |
| ▲ | monocasa 20 hours ago | parent | prev [-] | | All the more reason to treat it as a commoditize your complement situation. |
|
|
|
| ▲ | seanmcdirmid a day ago | parent | prev | next [-] |
| Alibaba isn’t really the Chinese government though, or are you saying Americans will think that ever since Jack Ma was harmonized? |
| |
| ▲ | derektank a day ago | parent | next [-] | | I think trying to tease apart the private and public sector is very hard in China. Setting aside state owned enterprises, even nominally private companies that employ at least 3 CCP members are required by law to form a party committee within the company to represent party interests. And given the party functionally is the government, you have a situation where the government has representatives inside every major private company. There’s no obvious parallel to this in western countries. | |
| ▲ | elmer2 a day ago | parent | prev | next [-] | | Any large company in China is only allowed to get this way by direct control from the CCP. This isn't really anything nee and I thought it was common knowledge by now. | | |
| ▲ | asdewqqwer a day ago | parent [-] | | I thought it should always be common knowledge that there is no way for any organization of 100m people to have one single mind. | | |
| ▲ | varjag a day ago | parent [-] | | There are numerous institutions that have agenda transcending individuals. Communist parties are very prominent among them. |
|
| |
| ▲ | traceroute66 a day ago | parent | prev [-] | | > are you saying Americans will think that I wasn't saying anything about what Americans would think. I was saying about what they would inevitably be told by US politicians and by US AI companies. If you were a sales-rep or marketeer at a US AI company, I bet you would be using the old "evil communists" routine in relation to any closed Chinese model. I was saying that by releasing as open weights, the company has removed that line of argument. Clearly I was a bit broad in my use of "the Chinese" when in this case it was, as you say, a Chinese company. | | |
| ▲ | seanmcdirmid a day ago | parent [-] | | US politicians are all over the map on this, but they aren’t really talking about Chinese AI much, it’s not as visible or tangible to most Americans like TikTok was. |
|
|
|
| ▲ | maxignol a day ago | parent | prev | next [-] |
| Shouldn’t we fear they start doing only close source like most us labs once they catch up in market shares ? |
| |
| ▲ | traceroute66 a day ago | parent | next [-] | | > Shouldn’t we fear they start doing only close source like most us labs once they catch up in market shares ? IMHO no. I think it is relatively safe to say that the predominant reason the US labs are closed source is so they can hype up their trillion-dollar valuations on pretty much negative return on capital employed, all propped up by fragile circular financing. Never say never, of course. But I just don't see it happening any time soon. | |
| ▲ | vidarh a day ago | parent | prev | next [-] | | Closing future models won't take away our access to the open weight ones. | | |
| ▲ | dannyw a day ago | parent | next [-] | | It’s also like smartphones. In the early years, every year was a huge jump. I still remember marvelling at my iPhone 4’s detailed display, and video calling for the first time. Now? I don’t even know or care about what the latest iPhones have, I’ll get a new one when mine breaks. | |
| ▲ | cyanydeez a day ago | parent | prev [-] | | people really underestimate how powerful just the consumer available models are. 128GB gets you pretty much a coding agent for typical apps. Even less with a good harness and logic set. |
| |
| ▲ | vkou a day ago | parent | prev [-] | | Yeah, but that mousetrap keeps working for SV startups, what makes you think it won't work for Chinese ones? Uber spent a decade undermining taxis, and once it had market share, it stopped giving away rides and raised prices. It now costs more than a regular taxi, with the quality of the ride being... At best proportionate to the premium in price. | | |
| ▲ | 1over137 11 hours ago | parent [-] | | Uber costs more than regular taxi? In what country/region? Not where I am. | | |
| ▲ | vkou 10 hours ago | parent [-] | | Looking at a 7 mile trip to a random destination in Seattle, right now, I can pay $26.70 for an Uber if I'm willing to wait 20 minutes for a pickup. With a $3/mile fee and a $5 pickup fee, that's exactly equal to that of a taxi. If I'm not willing to wait 20 minutes, I'll be paying an extra $5 minimum. These rates also go up during busy times. Looking at Lyft, that same trip is $29, without a wait. A trip from downtown to SeaTac is $61. Yellow Cab does that same trip for $40. |
|
|
|
|
| ▲ | mschuster91 a day ago | parent | prev [-] |
| > So you release it as open weights which is a win-win. Global adoption of the model and you get to give the American AI companies a kick in the nuts because you know they will never release open weights apart from highly quantised crippled shit. And on top of that, it's a perfect opportunity to include poisoned training data or excluding it. You know, omitting anything about Tiananmen Square, China's genocides against Uyghurs and Tibetans, or including texts propagandizing for the "reunification" (aka, annexation) of Taiwan. And everyone who builds something like an interactive chatbot based on such "open weights" models now has a subtle chance of the answer being ideologically poisoned by the CCP. We need actual open source, not "open weights" scam. |
| |
| ▲ | seanmcdirmid a day ago | parent | next [-] | | How does this work for RAG? Do they make it so the model doesn’t have that fact in their weights or do they make it not talk about it when it is included in context. Ironically, Chinese models have the most uncensored versions available for download. Fairly sure they own the porn market. | | |
| ▲ | nostrebored a day ago | parent [-] | | It’s in the weights. Context needs to be attended to to create a response, and the weights dictate what response is decoded. If you include retrieved context that has an American perspective, I imagine the think trace has some reconciliation about how they must be incorrect. |
| |
| ▲ | notnullorvoid a day ago | parent | prev | next [-] | | I wouldn't be worried so much about those examples. One could take the open weights and fine tune them to either fix the poisoning or omission of obvious topics. It's the subtle topics that we should be concerned about, and double so with closed models where even if oddities are identified they are harder to research further and impossible to fix. | |
| ▲ | traceroute66 a day ago | parent | prev [-] | | > You know, omitting anything about Tiananmen Square, China's genocides against Uyghurs and Tibetans, or including texts propagandizing for the "reunification" (aka, annexation) of Taiwan. I am not Chinese and I'm not defending the Chinese, but I see this argument come up a lot. The hard reality is that what you say is simply not going to affect 99.9999999999% of users. Is it realistically going to affect anyone using an LLM in coding ? No. Is it realistically going to affect anyone using an LLM in $anything_else_not_politically_sensitive ? No. Does anyone seriously use LLMs for researching politically sensitive matters ? No. The US does not exactly have an entirely pristine history either. Shall we discuss the post-9-11 related infrastructure of Guantanamo Bay ? Or the "Detention and Interrogation Program" that included a network of clandestine extrajudicial detention centres, officially known as "black sites"[1]? Or maybe you would like to discuss the US supply of weapons for use in Gaza ? [1]https://en.wikipedia.org/wiki/CIA_black_sites | | |
| ▲ | leereeves a day ago | parent [-] | | > The US does not exactly have an entirely pristine history either. Shall we discuss the post-9-11 related infrastructure of Guantanamo Bay ? Or the "Detention and Interrogation Program" that included a network of clandestine extrajudicial detention centres, officially known as "black sites"[1]? Linking a US website discussing the topic doesn't exactly support your point. | | |
| ▲ | a day ago | parent | next [-] | | [deleted] | |
| ▲ | traceroute66 a day ago | parent | prev [-] | | > Linking a US website discussing the topic doesn't exactly support your point. It supports my point precisely. Recall I also said "Does anyone seriously use LLMs for researching politically sensitive matters ? No.". Just as there is plenty of information out there on the US's less than perfect history, there is also plenty of information out there on the various Chinese politically sensitive matters. You do not need a Chinese LLM to find out about it, all you need is a search engine. The point is you have an open-weights LLM that is very good for a vast number of non-political uses, such as coding. The point is that you can use the open-weights model instead of paying through the nose for a US model where they harvest your data unless you have an "enterprise" zero-data retention "trust me dude" clause that you have no viable way of verifying – and which incidentally is still subject to the good old "law, or court or administrative order" contract clauses, so it may not be as much of a zero-data retention as you think it is. |
|
|
|