Remix.run Logo
atleastoptimal a day ago

AI models cost tens of millions to train. Offering them for free won’t justify the upfront costs.

The Chinese model of model training/open sourcing only makes sense in the context of the overall strategy of undercutting American frontier labs’ profit margins.

bwestergard a day ago | parent | next [-]

"the overall strategy of undercutting American frontier labs’ profit margins"

I don't doubt that's an unregretted side-effect for political leaders in China.

But the major motivation is to accelerate diffusion within their own massive economy in the pursuit of an across the board productivity boost in the face of an aging population.

henry2023 a day ago | parent | prev | next [-]

Couldn’t you say the same argument about VC funded startups? They lose money following a strategic goal.

The main difference here is if a startup goes underwater all the tech is usually lost. The Chinese weights are not going anywhere if the labs fail.

atleastoptimal a day ago | parent | next [-]

The VC money is only contingent on the strategy eventually bearing fruit. I do imagine going open source -> closed source could work for some model companies who get enterprise/ecosystem buy-in but the probability of ROI is lower.

vohk 21 hours ago | parent [-]

I think open source will remain competitive among smaller players and adjacent industries wanting to avoid lock-in with the majors. OpenAI, Anthropic, Google, etc are all out to win - they require profit extraction from their R&D. China seems to have, over the near term, accepted that they can not (or at least have not) pull ahead and so open source collaboration speeds the collective development, keeps them close to the frontier, and ensures their industry has access to learn from and implement. The USA playing export controls games with Fable made that aspect very stark.

But I agree that's the catch - it doesn't make sense to throw money at open source models in hopes of direct return, so you need a nation or conglomerate to do it so as to control the technology they rely on.

mmonaghan 12 hours ago | parent | prev [-]

This is true for any company following an open source strategy. You'll have the weights but you'll need to run inference, figure out your system prompt, sampling, quantization, etc etc. Loads of tuning.

elasticsearch the first example that comes to mind. you can run it yourself but elastic gives you so many lessons learned and tunes ootb that it sings with relatively little effort, though still reqiures some.

about a million dbs i could make the same comparison for

rlt 21 hours ago | parent | prev | next [-]

I can see two reasons American companies might want to train models they give away for free:

1) They sell compute: chips (Nvidia), data centers (AWS, Microsoft, Google, SpaceX, etc), or even end-user device manufacturers like Apple (e.x. M7 rumored to have 1.5TB of unified memory). If Jevon's paradox holds, then cheaper (or free) models means more demand. But compute is likely supply-constrained for years anyway.

2) Their product isn't AI but depends on AI being cheap, or they don't want competitors to capture that value, i.e. "commoditize your complement" https://gwern.net/complement

It probably doesn't make sense for these companies to invest a lot of money training models that will be obsolete in a few months anyway. When progress starts to plateau I'd expect more companies to start training models they give away for free.

gabriel666smith 21 hours ago | parent | prev | next [-]

I don't think this is necessarily going to prove to be true.

I often see the sentiment: "the Chinese strategy only makes sense in the context of undercutting American labs' profit margins".

If, for example, you are a company with a near-monopoly on "serving video content", and you feel reasonably confident about retaining a decent slice of the serving-video-content market (Google in the west is an example, Tencent in the east), then training video models on your dataset - and releasing them freely - makes an awful lot of sense.

Free tools to create with mean more video content. In this hypothetical, you're reasonably certain that any video content which does get created will also be watched on your platform.

That is a net positive. The question becomes: How many watch-hours earns back the cost of training a model? It's probably not really that many, especially when you have a near-monopoly on a billion sets of eyes.

It's also a net-positive if people build better video models from research you release, because - again - you are reasonably certain that the even-more-innovative content those models produce will be watched on your platform.

It really begins to make strategic sense if your company is in a GPU-poor environment. Your costs cease at the point you upload a model if your users are running it themselves. You don't have to serve the model. The content is still created.

You are also less likely, I think, to alienate human creators whose work the model was trained on if the model is not sold back to them as a subscription, or by the token, but given for free as a tool.

This frames the conversation very differently. It creates, I think, less of an "us vs them" dynamic, and more of a rising tide.

It's true that it is also beneficial that these models undercut (especially in language models) American companies. But, generally, Americans are not the customers of Chinese companies releasing models. They are already serving a huge volume of customers in a complex, existing marketplace.

The full picture is much more nuanced than simply a geopolitical desire to undercut US labs, and there are several other reasons the strategy can make logical sense.

gslepak a day ago | parent | prev | next [-]

Yeah this is a battle and it's why governments decide to spend resources on this. Protectionism won't help America, American needs to compete. There's a general consensus that open source AI must win because people don't want to end up as slaves to a megacorp, so if you're anti-open source AI you're not gonna fare well.

faitswulff a day ago | parent | prev | next [-]

It’s also a flywheel for China’s homegrown chips industry.

childintime a day ago | parent [-]

Exactly, the whole ecosystem. Cheap robots with cheap models. AI is oil to let the machine work, it should be cheap and ubiquitous.

kamranjon a day ago | parent | prev | next [-]

"It’s obvious to me that there are ecosystem benefits throughout China, from manufacturing to scientific research; every sector can just plug in these models."

I agree with this sentiment and think it's echoed in Fareed Zakarias take here: https://youtu.be/VBblUjLw5lE

China seems to perceive AI as a much more sensible technology than the US and seems to be integrating it in far more industries than the US.

I'm not sure the American mind can understand the distributed benefits afforded to the Chinese economy from opening their AI models, I think it's pretty reductive to assume it's purely a strategy of undercutting American frontier labs.

kettlecorn 21 hours ago | parent | prev | next [-]

I don't think this is accurate. AI is driving the cost of software towards 0 and these AI models themselves are software.

Releasing the models for free accelerates the trend but if you're a startup that needs leverage it's a good way to build brand and customer momentum that will be relevant in the more established future market.

I can see an American company taking on the same strategy, and in fact Thinking Machines based out of San Francisco did that just a few days ago by releasing their first model with open weights.

znnajdla 18 hours ago | parent | prev | next [-]

> AI models cost tens of millions to train.

There are people that spend tens of millions of dollars on paintings and artwork. I can see plenty of reasons why organizations and individuals will continue to want to drop a few million on an AI model just for the fun and prestige.

__MatrixMan__ a day ago | parent | prev | next [-]

There are good reasons to dislike outcomes that involve a single entity pulling well ahead of the pack here. Whether or not it continues to be American labs in the crosshairs and Chinese operators doing the aiming, perhaps it's reasonable to plan for continued efforts of this sort.

lagrange77 19 hours ago | parent | prev | next [-]

I'm genuinely curious:

Is that maybe spilled milk?

Maybe today's US frontier models provide enough information content, so that the momentum suffices to use them as a base for every coming generation of distilled and later fine-tuned models?

cco 16 hours ago | parent | prev | next [-]

> AI models cost tens of millions to train.

A pittance frankly. Something that could easily be covered by oh I dunno, let's call it a National Science Foundation who's in charge of subsidizing important basic research for a nation's interests.

Anywho, when the market is trillions (and of potential nation state concern), it is pretty inconsequential and very much worthwhile.

Aside, I think your scale is a bit off, I think Moonshot has raised $5B and potentially they get other breaks from China, not sure. So to produce something like SOTA takes billions, not tens of millions. I'd still argue it is worthwhile to subsidize and invest in open versions, imagine spending $5B to unlocking a few percentage point increases in your country's productivity.

JSR_FDED a day ago | parent | prev | next [-]

Either that, or they don't want to be hostage to a handful of companies intent on owning the future. Personally I'm right there with them.

a day ago | parent | prev | next [-]
[deleted]
michaelt a day ago | parent | prev | next [-]

There is a huge cultural influence opportunity too.

Imagine if, in 10 years time, every school kid is learning the causes of the US civil war from an LLM, getting their essays on hiroshima and nagasaki graded by an LLM, and a million other things.

A country with competitive LLMs gets to decide whether "it was more complicated than just slavery", and whether "it was tragic but necessary, saving lives over all".

Countries without competitive LLMs are effectively going to be buying all their history, economics and sociology textbooks from abroad.

applicative a day ago | parent | next [-]

It might seem so til you consider how an LLM is trained.

An indirect illustration: I can attest that Deepseek has very good 19th German, and knowledge of German 19th c literature, science and historical scholarship. No one in China could control the training that led to this. The German training sources were well aware of the exact nature of eg American slavery, so they are in the weights.

State control operates in the outer layers not the llm itself.

stickfigure 20 hours ago | parent [-]

People can and will use LLMs to curate the training data for the next generation. I expect they already do.

georgeburdell 21 hours ago | parent | prev | next [-]

Maybe your kids. In my corner of the country, parents are actively demanding that schools back off computer usage, much less AI

stackghost 21 hours ago | parent | prev [-]

Every school kid in the USA, you mean? Because I think other countries would rightly perceive the world you described as a dystopia.

I don't want my kids' education to be surrendered to the whims of Big Tech douchebags any more than I want AI decisions in legal cases or an AI replacement for a family doctor.

Some systems are better left mostly analog. Education is one of them.

nickysielicki 21 hours ago | parent [-]

Your kids education was already surrendered to the whims of the Big Textbook Politburo. History textbooks are full of propaganda. I think what we have today and what we grew up with is 10x more dystopian.

stackghost 19 hours ago | parent [-]

Textbooks, especially in the early years, are a minor part of education.

I want them to have an actual human teacher.

asadotzler 17 hours ago | parent | prev | next [-]

The upfront costs are actually a lot higher, but still absolutely trivial in the big picture. OpenAI pulled in $120,000,000,000 in one round of fundraising and they're closing in on $200,000,000,000 total. Even at 1 billion dollars, new frontier model training is only 0.5% of what they've raised.

mannanj a day ago | parent | prev [-]

You're completely ignoring the value of data.