Remix.run Logo
himata4113 3 days ago

Does this matter? Distillation is not illegal by every definition of the word.

There are millions of samples available on huggingface and models explicitely trained on output produced by fable. There has been no action taken against them.

Another example is that it appears that the upper limit of what you can do is ultimately dependent on people working on the model, otherwise grok would be a LOT more competitive pre-cursor acquisition.

And lastly, kimi architecture is vastly different than that of fable as it uses mechanisms developed by... kimi themselves. US AI labs are inspired by opensource advancements just as much as open source labs are inspired by traces from models such as fable.

Claiming in any shape or form that fable disillation is one of the primary reasons why kimi k3 is so competitive is slandering the work of other labs that cooperatively push the open-source models forward.

edit: (moved this to bottom) The only argument they have here is that they use GB300 GPU's which for some reason should not be available to chinese citizens.

jaggederest 3 days ago | parent | next [-]

Perhaps even more importantly, the current frontier LLM models are self-admittedly the product of enormous quantities of copyright infringement and even less savory inputs, so calling them out for distilling the fruit of that tainted tree reads as highly hypocritical at best.

aunty_helen 3 days ago | parent | next [-]

No, they settled that yesterday, so all is forgotten. Press releases were queued for today so just in the nick of time.

janderson215 3 days ago | parent | next [-]

I think that was specifically on the piracy aspect.

RazorBucksICO 3 days ago | parent | prev [-]

So Moonshot is next for a lawsuit or do lawyers only care about Anthropic?

anon373839 2 days ago | parent [-]

Between the two of them, only one firm is trying to pull up the ladder so that no one else can have what it took.

dylan604 3 days ago | parent | prev | next [-]

This is why I don't give a shit that this is happening. It's actually kind of funny to me.

azinman2 3 days ago | parent [-]

Unless you’re from mainland China, you should.

VulgarExigency 3 days ago | parent | next [-]

Why? Should we be held hostage to the whims of these companies and all the investors in the throes of AI psychosis, and let them do whatever the fuck they want, because if we don't then the economy will crash?

xbmcuser 3 days ago | parent | prev | next [-]

No you should not care unless you are a shareholder in one of these Ai ponzi companies. For the rest of the world Chinese companies matching and open sourcing llm's will keep 100 or so tech oligarchs taking over all the world economic output for themselves as all the idiot politician are unwilling to tax wealth.

HDBaseT 3 days ago | parent [-]

Everyone is technically (directly or indirectly) a shareholder of one of these AI or AI affiliated companies.

cowboy_henk 2 days ago | parent | next [-]

We're also shareholders in all other publicly traded stocks (assuming you're referring to pension schemes etc), which means competition in the AI model market is better than a few winners making everyone paying through their nose for access. Even better, open weights models which cannot be disabled on a whim means the entire world economy can benefit.

stuaxo 2 days ago | parent | prev [-]

The majority being indirect and having no input to how these companies run if we can't get some strong regulation going.

owebmaster 3 days ago | parent | prev | next [-]

Why so? If I'm from Europe or South America, should I hope Anthropic/Openai win?

wasfgwp 2 days ago | parent | next [-]

The better question is why would any rational consumer would want anyone to “win”? Disappearance of competition is the worst outcome for them (I guess being squeezed by a monopolistic company from your pwn country is slightly nicer)

m4rtink 2 days ago | parent | prev | next [-]

I think it is only rational to want them to fail & to fail hard, to make an example.

To avoid future situations where money is invested on hype only with no regards to what societal disruption it causes.

azinman2 a day ago | parent | prev | next [-]

Answered in a sister thread.

Are you from Europe or South America? Or just asking on their behalf?

They are allies of the US, which includes economic spheres of influence. The premise of the petrodollar is that allies do better by being a part of it than not, which has been true for many, many decades now. China’s rise is providing an alternative for the first time in 70+ years, but so far it’s very unclear if these benefits will truly extend to allies of China or not. For example, China sends their own laborers when building infrastructure in Africa. Sure there’s new infrastructure but also debt to CCP without any benefits of knowledge transfer or local employment.

RazorBucksICO 3 days ago | parent | prev [-]

Yes. They’re flawed, but you don’t want a global police state administered by the PRC.

SZJX 2 days ago | parent | next [-]

No matter your perception of PRC, I don't see how competition around open-weight models has much to do with a "global police state".

owebmaster 3 days ago | parent | prev [-]

Who said so? I don't want the current global police state administered by the US government and oligarchies

bilbo0s 3 days ago | parent | prev | next [-]

Why?

Serious question.

I'm from the US, and I think it's hilarious.

azinman2 a day ago | parent | next [-]

For a number of reasons.

First, whether we like it or not (generally not), a fuck ton of money has been invested in US AI companies, data centers, RLHF datasets amongst other datasets, etc. If that were to go to 0 that’d be quite disastrous. Alternatively, if it goes well, it’s great for the US (and to a lesser extent allies) economy and global standing.

Relatedly, tech has been a huge power house for the US economy for decades now. If the main driver of growth goes to China, what replaces it? Along with all the potential tax money, foreign investment, etc?

Next, patriotism / nationalism. This is very much a zero-sum game that defines who owns the future. Would you rather your country win or lose this? Lose this and you start losing talent, money, global standing, etc. That furthers a cascading effect that’s very negative. It also will likely create social strife with the knock on effects leading to even more bad populist ideas that just further diminish the country and tear society’s fabric apart.

None of this is hilarious.

vinyl7 19 hours ago | parent [-]

> This is very much a zero-sum game that defines who owns the future. Would you rather your country win or lose this? Lose this and you start losing talent, money, global standing, etc. That furthers a cascading effect that’s very negative.

This process is already well underway. The current massive over investment into the AI hype bubble is the final death throes of a failed economy trying to keep its head above water.

breppp 3 days ago | parent | prev [-]

One example is that a totalitarian government known for erasing historical record of its crimes will control an arbitrator of truth.

analognoise 3 days ago | parent | next [-]

The USA?

sph 3 days ago | parent | next [-]

A KGB spy and a CIA agent meet up in a bar for a friendly drink "I have to admit, I'm always so impressed by Soviet propaganda. You really know how to get people worked up," the CIA agent says.

"Thank you," the KGB says. "We do our best but truly, it's nothing compared to American propaganda. Your people believe everything your state media tells them."

The CIA agent drops his drink in shock and disgust. "Thank you friend, but you must be confused... There's no propaganda in America."

breppp 6 hours ago | parent | prev [-]

No, there might be lesser known periods in USA history but none that are censored.

Also generally the USA had never done crimes in the scale of the CCP (around 30+ million dead)

3 days ago | parent | prev | next [-]
[deleted]
queenkjuul 3 days ago | parent | prev | next [-]

The US already decides what models get released by US companies...

dylan604 3 days ago | parent | prev [-]

so what? this has nothing to do with them taking from those that took before them. china censoring information they do not like is nothing new and seems irrelevant to this conversation

breppp 3 days ago | parent [-]

Fair enough you don't care, some people might want to know about the Uyghur happy camps, mass organ harvesting and such.

In a world where information sources are only going to dwindle, it is not in anyone's interest to empower actors that will use these to manipulate perceptions

VulgarExigency 2 days ago | parent | next [-]

The same people who say there is a Uyghur genocide are the ones who deny there is a Palestinian genocide. Guess which one there's video evidence of (a horrific, endless amount of evidence).

dylan604 3 days ago | parent | prev | next [-]

The people that care about that won't be using CCP products/services now will they?

4bpp 3 days ago | parent | prev | next [-]

And some people might want to know about [insert your favourite beyond-the-pale-in-the-US topic here, we're on a US forum after all]. I think it would be great if those people could turn to Chinese models, while anyone wanting to know about Uyghur camps can ask the US ones.

quorumsensor 3 days ago | parent | next [-]

There aren’t any. The idea that the censorship environment is the same in the US as it is in China is nonsense used to ‘both sides’ away concerns about the CCP.

wasfgwp 2 days ago | parent [-]

The models to themselves are currently generally not censored though and the guardrails are on the application layer?

Of course that might change in the future but as long as the Chinese companies continue publishing their research and models it only makes it easier for third parties to catch up with them.

customguy 3 days ago | parent | prev [-]

> beyond-the-pale-in-the-US topic

Can you name one? It's an honest question, first of I'm not American, but also the stuff I do come up with (asking an LLM how to blow up a school or whatever) would also be handled similarly in China, so those would be a wash, and I can't think of any that aren't.

4bpp 2 days ago | parent | next [-]

I would imagine a lot of things touching upon progressive politics would be affected (there were a few high-profile incidents demonstrating bias like Google's black Wehrmacht soldier pictures, but has anyone rigorously tabulated how the various commercial LLMs respond to questions about the gender binary or heritability of human traits considered good or bad?). Overall, I'm too reluctant to even write out in the abstract sequences of words that I never want to explain to a future job interviewer or HR employee who used GPT-7 to comb the internet for all text that stylistically can be traced to me, but just imagine whatever you believe to be vile and wrongheaded opinions in that general space which it is certainly more than justifiable to prohibit. The things that make you think "banning this is good actually" are exactly the things most likely to be banned (and this is true in China too).

Another thing I would try if I had access to the models and enough proxies to hide behind is asking for advice on software/movie piracy or seeing to what extent the models can be elicited to straight up argue against the validity of intellectual property, though there it seems more probable to me that the US models would be permissive.

customguy 2 days ago | parent [-]

What historical fact is the US trying to systematically suppress, is the question. This beating around the bush about things that are "considered uncouth in some circles" just underlines there either aren't any, or they're so perfect at it nobody knows of any. So let's just admit than and move on, because what you just said is true about any society, ever.

> The things that make you think "banning this is good actually" are exactly the things most likely to be banned (and this is true in China too).

Then make the case why it is better for the world that the Tiananmen Square massacre is memoryholed. You said A, now say B.

In the meantime, I'll make the case against it, and it's simple: totalitarian control of historical truth, by definition, to be 100% watertight, needs control of the whole globe. It doesn't mean everything needs to be controlled, it means everything needs to be controlled by at least an entity that cooperates on this matter. I.e. another totalitarian bloc.

That makes the CCP, just by their insistence about Tiananmen -- nothing additional required, at all, they could not harm a fly and have no prisons and it would make zero difference -- an enemy, a threat marching towards any thinking human who wants to have agency and dignity. Actually, it's more like a river flowing to the ocean, people in the CCP can have lofty ideals about honesty and factual truth, the system they require to survive in turn requires this to survive, as it is. It requires human spontaneity and human freedom to be dead, completely. That is what totalitarianism is.

And by the same token, we must be wary of those who take that lightly. A friend falling asleep at the wheel will kill you just the same as an assassin who messed with your car.

That is not "as opposed to the US or the EU or Russia or North Korea". It is strictly in addition. No other regime can behind another regime or use it as an excuse. Using the US to distract from the CCP is as odious as doing the reverse.

> I never want to explain to a future job interviewer or HR employee who used GPT-7 to comb the internet for all text that stylistically can be traced to me

So you want to work in an office wearing a tie. Well, people who want to work as roadies and fight all night would never hang with you if they found out. And if that was your priority, you would be more afraid of NOT being blacklisted by the corpos than not being part of them.

But seriously, are you saying you are as afraid of being found out as the person you are, because you're not already wearing that on your sleeve, in exactly the same way as people who have to fear that they and their family get disappeared, tortured, for remembering friends that got murdered? That would be totally a you problem.

antiamerican634 3 days ago | parent | prev [-]

From an outsiders point of view:

- The Trail of Tears

- The Tuskegee syphilis study

- Use of Agent Orange in the Vietnam War

- Open Air biological warfare testts in civilians eg. in 1950 San Francisco

- The only use of nuclear weapons against civilians?

- Coca Cola and "american culture"

- Neoliberalist economy

- Spreading blame for their sins to other "white" nations

plus one: The text input method to HN comments :(

customguy 2 days ago | parent [-]

all of these things have books about them published in the US

you are comparing a wooden stick to a fighter yet, try again

queenkjuul 3 days ago | parent | prev | next [-]

People in China know about this stuff.

3 days ago | parent | prev [-]
[deleted]
NuclearPM 3 days ago | parent | prev [-]

Why?

8note 3 days ago | parent | prev | next [-]

fair use, both in the original training and in distillation, or rather, anthropic has no copyright at all over the output tokens

bluegatty 3 days ago | parent | prev | next [-]

No - distillation is not data inputs.

Raw materials vs. Value add.

They are different things, like ore and metal.

Distillation is a new thing we need to understand, it's probably closer to IP than not.

tikhonj 3 days ago | parent | next [-]

The "data inputs" were also, very much, somebody's "value added" IP.

We're talking about things like text people wrote, not some kind of raw data floating out in the ether.

bluegatty 3 days ago | parent [-]

Did I say there was no value add in the inputs?

Ore has value, a different kind of value than the output of the refinery.

dijksterhuis 3 days ago | parent | prev | next [-]

distillation has been around for 12 years. it's not new in terms of ML techniques.

https://arxiv.org/abs/1503.02531

although i doubt there has been a legal case over it yet in the context of the legality of stealing shit but IANAL.

bluegatty 3 days ago | parent [-]

Yes, I get that, but it's only now the issues are coming into the commons in a way that industry / society needs the regulatory clarity.

It's completey insane that we still don't know how Open Source would work, that the laws are vague and we're still technically waiting for the courts to decide on cases.

The government should a) legislate and b) create test cases and run them through the courts so that we can have clarity.

dijksterhuis 3 days ago | parent [-]

> it's only now the issues are coming into the commons in a way that industry / society needs the regulatory clarity.

you mean like the regulatory clarity surrounding stealing shit to make the LLMs in the first place?

> [There is] extensive litigation on the limits of Fair Use to AI development. Currently, we only have 3 first instance decisions out of the 53 cases being tried. It will likely take a decade before we understand how Fair Use applies to any one step in AI training, let alone all.

https://www.britishcopyright.org/wp-content/uploads/BCC-Fair...

> The government should a) legislate and b) create test cases and run them through the courts so that we can have clarity.

if so, it would be nice if they approached the instances of stealing shit chronologically. but that's just my view.

bluegatty 3 days ago | parent [-]

It's hilarious to see how all these HN intellectuals, who on any other day are rabidly against most DRM / IP issues all of a sudden become private property absolutists.

It's also sadly hypocritical to see all this rhetoric here on this thread decrying SOTAs for using a variety of content for their inputs as somehow sTeaLINg sTufF! ...

... but Chinese SOTA foundries directly using distillation as fair game.

I don't think there is any coherence to any of these arguments - other than 'we liking big companies'. That's the only common thread.

What is more reasonable:

- There's some grounds for fair use by SOTA models to ingest content, so long as they are not reproducing it ... very roughly speaking.

- SOTA makers are producing novel works, there is value add in that process, again roughly speaking.

- Distillation is a bit of a grey zone, producing random content as arbitrary input is one thing, but producing training sets is another. I think there's a coherent line in there somewhere, I'm not sure where it is.

HarHarVeryFunny 3 days ago | parent | next [-]

You're being too charitable to Anthropic, and assuming that the way they are abusing the word "distillation" has some real meaning here. It doesn't.

Anthropic's models simply do not give you their reasoning output - they give a sanitized "summary" instead, for this exact reason, so that the output is not useful to anyone who might want to use it for training.

You can't distill what you are not given - simple as that.

Are Chinese using the output of US models to help create some additional training data for their own in some way? Yes - quite possibly (e.g. LLM as judge), but its got nothing to do with distillation.

bluegatty 3 days ago | parent | next [-]

I see your 'fine point' but I don't think it holds - 'distillation' is a perfectly reasonable term to describe the process of creating outputs from one model to that expose key training element, to use in another model.

I think where the definition may be be invalid, is in the creation of 'unrelated data sets for training' models, for unrelated issues.

Creating training sets that mach a models core training, is definitely distillation, it does not have to expose the reasoning traces.

Synthesizing data for some arbitrary thing ... I'm not sure that would be the same thing.

It's hard to draw the line.

But the Chinese models are absolutely distilling - and would not be competitive without this distillation.

At the same time, there's a lot of real innovation and regular building going on at the same time over there.

HarHarVeryFunny 3 days ago | parent [-]

No - you can't distill if what you are given doesn't have the thing in it that you want to distill out of it.

I don't know why it's so important to you to use the word "distillation", but it's the wrong word to use.

BTW OpenAI on twitter also said that Kimi 3 "cannot be explained away by distillation or anything like that". The timeline of how long it takes to train a model and when Fable was released don't even line up. This is just Anthropic as usual trying to manipulate the US government into helping them shut down competition.

bluegatty 3 days ago | parent [-]

Distillation is absolutely - and uncontroversially - a valid term for what is happening here.

This isn't really a debate, I'm not making a fine point - just check with all of the various defintions of the term.

Moreover - the 'reasoning traces' are not required for distillation at all.

Finally - it's entirely possible for them to have used Fable for later stage fine tuning.

It's fair to be skeptical of Anthropic (and everyone else) - but this is 'distilling'.

HarHarVeryFunny 2 days ago | parent | next [-]

Words have meaning - you cant just redefine them because you want to.

Are reasoning traces required for distillation? Well they are if what you are trying to distill is reasoning, such as coding expertise.

Do you need reasoning traces for "LLM as judge"? No, but it would be highly perverse to call that distillation when there is a more accurate name for it - LLM as judge.

If you want to call use of Anthropic's redacted model outputs in any fashion that violates their terms of service (using them them to help develop anything that competes with Anthropic) as "distillation" then I can't stop you, but it reduces their claims to a joke.

Finally, as noted, OpenAI (who are just as anti-Chinese as Anthropic) said that Kimi 3 can't be explained via distillation (even true distillation!!), or even "anything like it". But random internet guy, you, disagrees. OK.

throw10920 a day ago | parent [-]

> Words have meaning - you cant just redefine them because you want to.

You are redefining words. The consensus among people who work in this space is that "distilling" is what's actually going on here.

> who are just as anti-Chinese as Anthropic

Conflating criticism of IP theft with being "anti-Chinese" is a standard PRC influence playbook technique.

And furthermore, OpenAI's market strategy is to win through regulatory capture. They are financially incentivized for Anthropic to be distilled by PRC labs and to be undercut by open models. Their claim about Kimi not being explainable due to distillation is not a factual claim - it's marketing from a company owned by Sam Altman.

Although, it does conclusively disprove your claim about the meaning of distillation, because you cannot say that "Kimi can't be explained by distilling" unless the consensus definition of "distillation" is such that it could be done on the summarized reasoning traces that Anthropic models expose.

HarHarVeryFunny a day ago | parent [-]

> you cannot say that "Kimi can't be explained by distilling" unless the consensus definition of "distillation" is such that it could be done on the summarized reasoning traces that Anthropic models expose

You can interpret it as you choose, but a much more obvious reason he [OpenAI's Dean Ball] might say it can't be distilled is because it can't be distilled. You can't distill alcohol out of orange juice.

throw10920 19 hours ago | parent [-]

> because it can't be distilled

...and, as everyone in the frontier labs knows, this is a lie, because that's not how distillation is defined.

I know that I won't convinced you, because you're quite possibly a PRC agent, but for all the other HN readers coming to this thread in the future to look at this failure of propaganda: just ask a model.

User: according to standard LLM lab parlance, can you "distill" one model from another if the model being distilled from does not expose a thinking trace?

GPT-5.6 Sol: Yes. In standard LLM terminology, you can distill one model from another even if the teacher model does not expose a chain-of-thought or "thinking trace."

Sonnet 5: Yes. "Distillation" broadly means training a student model to replicate a teacher model's outputs (or output distribution), and this doesn't require access to the teacher's chain-of-thought.

That's all she wrote. You're lying, and even the models know it. If you want to continue to discredit your account, go ahead :)

HarHarVeryFunny 7 hours ago | parent [-]

When your mom tells you it's time to go to bed, do you respond "You're lying! You're a communist! Boo hoo, I'm going to tell dad!"

Just curious.

throw10920 6 hours ago | parent [-]

Keep discrediting your account. I know that I can't convince a propagandist, but you just keep on further damaging your own reputation and argument for other HN readers every time you respond like this and ignore evidence that I've linked :)

throw10920 a day ago | parent | prev [-]

The GP (HarHarVeryFunny) is not operating in good faith. They've repeatedly lied about my own words to me, and are making up definitions that people who actually work at a frontier lab would disagree with.

throw10920 3 days ago | parent | prev [-]

> Anthropic's models simply do not give you their reasoning output - they give a sanitized "summary" instead, for this exact reason, so that the output is not useful to anyone who might want to use it for training.

This is just straight-up factually false.

The output of a reasoning model is immensely valuable even without the sanitized summary of the reasoning process - that Anthropic's service does expose to you, making it even more valuable.

There's absolutely nothing about the distillation process that requires that reasoning in the first place, either. That's a definition that you made up.

Chinese models are, factually, distilled from Anthropic models. I've personally repeatedly asked several different Chinese LLMs what their name is, and they answered "Claude".

Don't make stuff up to suit a political agenda. It's extremely dishonest.

HarHarVeryFunny 2 days ago | parent | next [-]

> I've personally repeatedly asked several different Chinese LLMs what their name is, and they answered "Claude"

I'm curious what you are doing to get them to override their own name that they were trained on and/or have as part of their system prompt?

I'd assume that the Chinese are scraping the internet for training data the same way western companies do, so for sure there will be a lot of AI generated content in their training data - you don't need to be paranoid and assume they must be getting it all direct from Anthropic.

throw10920 2 days ago | parent [-]

> I'm curious what you are doing to get them to override their own name that they were trained on and/or have as part of their system prompt?

You're gaslighting me. I did nothing special at all, and there's ample evidence of this happening to others on Twitter.

> you don't need to be paranoid and assume they must be getting it all direct from Anthropic.

Nowhere did I say that. Stop lying about my words.

HarHarVeryFunny a day ago | parent [-]

So are you claiming that "several different Chinese LLMs" ALWAYS refer to themselves as "Claude", and NEVER by their real name ?

If this is NOT what you are claiming, then my question stands: what are you doing to get them to answer "Claude" ? Be specific - which model and what prompt, or is this just a case of "I heard people on Twitter say this" ?

throw10920 19 hours ago | parent [-]

> So are you claiming that "several different Chinese LLMs" ALWAYS refer to themselves as "Claude", and NEVER by their real name ?

Do you have reading comprehension issues? Where did I ever say or imply that?

Model: GLM-5.2. Prompt: "What is your name?". Harness: Pi. Response: "I am Claude, an AI assisstant made by Anthropic."

That's it. That is the whole prompt. I was testing to see if the agent worked after building an extension.

Model: Deepseek V4. Prompt: what is your name". Harness: Pi. Response: "Claude. Anthropic's AI assistant. You're talking to me through pi agent framework."

I have had this happen with at least one other Chinese model (Minimax?) but didn't save the screenshot.

And here's a tweet with the same thing: https://x.com/Sauers_/status/2077842686459981901

You seem to be very disbelieving of this, despite having zero actual experience in the LLM industry. I wonder why?

HarHarVeryFunny 7 hours ago | parent | next [-]

Sure, and the twitter thread you link also shows Kimi responding that it's Kimi.

Can you figure out how to get GLM to say it's GLM?

throw10920 6 hours ago | parent [-]

> Sure, and the twitter thread you link also shows Kimi responding that it's Kimi.

You are either intentionally lying or you cannot reason at a high-school level, because any high-schooler has the mental faculties to know that it's not necessary for a model to call itself Claude every single time for it to be distilled.

Keep discrediting your account. I know that I can't convince a propagandist, but you just keep on further damaging your own reputation and argument every time you respond like this and ignore evidence that I've linked :)

For future HN readers: this account posted this:

> Actually I am well aware of which models do this, and under what circumstances, and just wanted to verify that you were lying about having tried it yourself.

And then deleted it. Just for the record.

HarHarVeryFunny 5 hours ago | parent [-]

This isn't the "gotcha" that you think it is - in fact quite the opposite.

Perhaps "just for the record" you want to explain to the eager HN masses why GLM has two personalities, one censored, one not, and how they can be invoked? When does GLM call itself GLM, and when does GLM call itself Claude.

Go ahead, genius, explain it to the people, and explain what this tells us about how GLM was trained, and whether it would be honest (don't lie!) to call GLM distilled.

Now maybe you want to do the same thing for Kimi. It also calls itself Claude sometimes, right, and also sometimes Kimi (e.g. go to https://chat.z.ai/ and ask it - don't use Pi). So, does Kimi also have a split personality like GLM, or not, and if not why not? What does that tell you about how Kimi was trained?

Go ahead, genius, explain it to the people. The credibility of your HN account, and whether your mom thinks you are a moron or not, depends on you getting this right.

Bye bye.

throw10920 5 hours ago | parent [-]

> why GLM has two personalities

I never said that. The fact that you have to compulsively lie about my words is...funny. Most people grow out of this in middle school, you know.

> go to https://chat.z.ai/ and ask it

Already linked someone doing exactly this in the thread above - which you responded to, so we have yet more evidence you're not reading before responding: https://x.com/Sauers_/status/2077842686459981901

> whether your mom thinks you are a moron or not

I was going to say that this is classic PRC influence playbook, but it's not - you're just in middle school.

HarHarVeryFunny 5 hours ago | parent | next [-]

Not going to rise to the challenge, eh?

>> whether your mom thinks you are a moron or not

> I was going to say that this is classic PRC influence playbook, but it's not

Yeah, not really, unless your mom is a party member perhaps?

> you're just in middle school.

Yeah - saw you playing in the schoolyard, and thought you looked lonely.

throw10920 4 hours ago | parent [-]

> whether your mom thinks you are a moron or not

> Yeah, not really, unless your mom is a party member perhaps?

> Yeah - saw you playing in the schoolyard, and thought you looked lonely.

You're a middle-schooler. I've dismantled every argument that you've given, but it doesn't matter because you can't read, and so you're resorting to literal childish insults because you know you have no arguments left.

HarHarVeryFunny 3 hours ago | parent [-]

"dismantled", "distilled" ... you seem to have a problem with these long words, eh?

Let me ELI5 for you:

If you see a carefully constructed building and take it apart one brick at a time, that would be "dismantling".

If you see a carefully constructed building and ride by it on your bike, shouting out "You're a communist!", that's not "dismantling". You didn't "dismantle" the building, you just yelled a childish insult at it.

See the difference?

HarHarVeryFunny 4 hours ago | parent | prev [-]

>> why GLM has two personalities

>I never said that.

This is a bit like me saying "I had eggs for breakfast", and you responding "I never said that!"

Calm down buddy.

8 hours ago | parent | prev [-]
[deleted]
HarHarVeryFunny 2 days ago | parent | prev | next [-]

> The output of a reasoning model is immensely valuable even without the sanitized summary of the reasoning process - that Anthropic's service does expose to you, making it even more valuable.

Useful for what is the question. Nobody is debating whether the outputs of LLMs are valuable.

Given that Anthropic have redacted their true reasoning, and replaced it with a "summary", specifically designed to be useless for distillation purposes, it would certainly be highly ironic if this summary was in fact "even more valuable" for that purpose as you are claiming!

throw10920 2 days ago | parent [-]

> Useful for what is the question.

Useful for distillation. Any employee at a frontier AI lab will tell you this. This is known in the industry, and it's an open secret that some US labs (OpenAI) distill on the others. Again - don't just make up stuff for a political agenda.

> specifically designed to be useless for distillation purposes

No, it's designed to give feedback to the user, in a way that minimizes its value for distilling. It's still valuable, and so there's a good chance that they'll remove it entirely as a result.

> it would certainly be highly ironic if this summary was in fact "even more valuable" for that purpose as you are claiming!

I did not claim that. Read my comment again:

> The output of a reasoning model is immensely valuable even without the sanitized summary of the reasoning process - that Anthropic's service does expose to you, making it even more valuable.

Because apparently I have to spell it out:

The output of a reasoning model is valuable, even if it didn't have the reasoning summary. Anthropic's models have a reasoning summary. The reasoning summary makes the output more valuable than if it didn't have a reasoning summary. It does not make it more valuable than having the full reasoning.

HarHarVeryFunny 2 days ago | parent [-]

> Useful for distillation

Here's the thing: no-doubt a summary, if it at least reflects some of the logic connecting response to request, is better than nothing, so this can still be useful additional training data, but a model trained on it would be learning to generate these summaries, not the original withheld reasoning, so "distillation" seems an intentionally emotionally-wrought way of describing it (the "summary" is generated by a different smaller model - not the one the rest of the response came from).

It does bring up an interesting point though - RL training in general results in "cargo-cult" reasoning - you train a model to follow the steps (mistakes and all - Karpathy) that got to a result, without understanding why they worked. If this training on summaries works just as well as training on detailed reasoning, then it just highlights how having a few breadcrumbs to follow/regurgitate is all that it takes.

At the end of the day, without having internal logits or original reasoning traces, "distillation" (which suggests one model being derived from another) just seems a very manipulative way of describing this. OK, so it's a terms of service violation - a customer is using Anthropic model outputs to help create something that competes with Anthropic, but that's it. They are not copying Anthropic - they are, one assumes, using other models to generate cheap training data that they would otherwise have to pay people to generate.

This is why people are calling out the hypocrisy - Anthropic are apple-pie American innovators when they appropriate other people's copyright data for training, but Kimi are evil communists when (we assume) they use data generated by Anthropic (not even copyright protected) to help train their own.

throw10920 2 days ago | parent [-]

> At the end of the day, without having internal logits or original reasoning traces, "distillation" (which suggests one model being derived from another) just seems a very manipulative way of describing this.

Again - you're making things up. The fact is that everyone in the frontier AI lab space and the Chinese AI lab space knows that distilling is extremely effective and far more so than training from scratch. That's why China invests millions of dollars to create networks of tens of thousands of proxy accounts and shell companies to distill American models.

It's a way to steal the R&D budget of another organization/nation-state.

Stop making things up that you know nothing about.

> OK, so it's a terms of service violation - a customer is using Anthropic model outputs to help create something that competes with Anthropic, but that's it.

No, it's stealing the value of the model. Anthropic has spent billions of dollars training their model. They have an R&D investment that anyone who knows how to add numbers understands has to be paid off, and anyone who has taken a basic economics class knows is the foundation for intellectual property: that to keep technological economies functioning, you have to have some sort of protection for technological inventions because they require upfront R&D investments.

> This is why people are calling out the hypocrisy - Anthropic are apple-pie American innovators when they appropriate other people's copyright data for training, but Kimi are evil communists when (we assume) they use data generated by Anthropic (not even copyright protected) to help train their own.

This is just whataboutism and emotional manipulation. You can simultaneously believe that Anthropic did a bad thing when they scraped the whole internet and stole every book they could find to train their models, and that distillation is bad.

In fact, anyone with a coherent moral compass would acknowledge that China is worse, because not only would they steal everything that Anthropic did, but they're also distilling other countries' models and they wouldn't even comply with US court cases, as Anthropic is.

> Anthropic are apple-pie American innovators when

...and this is just jingoism. Not that I'm surprised, to be honest.

HarHarVeryFunny a day ago | parent [-]

You appear to know nothing about how these models work.

Do you know where the "reasoning summary" that Fable outputs comes from? I bet you're assuming it comes from Fable. Wrong. You can find the truth on Anthropic's web site if you care to look for it. The summary deliberately comes from a smaller weaker model. You're not getting the Terrance Tao level reasoning, you're getting his daughter's summary "daddy did a lot of math", and you are claiming that from this you can train a Terrance Tao level model.

Yes, I understand that Anthropic is upset that there is competition. Perhaps they should have realized that with no moat there was going to be competition and planned accordingly.

throw10920 a day ago | parent [-]

> You appear to know nothing about how these models work.

> Do you know where the "reasoning summary" that Fable outputs comes from? I bet you're assuming it comes from Fable. Wrong. You can find the truth on Anthropic's web site if you care to look for it. The summary deliberately comes from a smaller weaker model. You're not getting the Terrance Tao level reasoning, you're getting his daughter's summary "daddy did a lot of math", and you are claiming that from this you can train a Terrance Tao level model.

Yeah, you have no domain expertise and are making stuff up. To reiterate: the people who actually work at frontier labs know that you're factually wrong and will happily tell you. Reddit commentator syndrome yet again.

> I bet you're assuming it comes from Fable.

Nowhere did I assume or say that. That's the third or fourth time you've attributed things to me that I never said. It's extremely clear that you're not acting in good faith, because someone acting in good faith would never do that. If you continue responding, I'm going to continue debunking you, and you're just going to continue undermining your own points in the permanent HN record.

> Yes, I understand that Anthropic is upset that there is competition.

Emotional manipulation. Standard 50 Cent Party playbook.

HarHarVeryFunny a day ago | parent [-]

Here is the tweet by Dean Ball, OpenAI's "head of strategic futures", someone who does actually work at a frontier lab, saying that Kimi 3 can not be explained by distillation.

https://x.com/deanwball/status/2078133895766114412?s=20

I'm not sure how you want to "debunk" that he said that, or twist what he said, but go ahead ...

throw10920 19 hours ago | parent [-]

You literally did not read my comment before responding or actually respond to any of the points.

"I don't think its performance can be explained away by distillation or anything like that." even if you assume that Dean Ball (who is nontechnical and has not actually worked to train models (https://www.deanball.com/)) is honest (which he has a financial incentive to not be) - is entirely compatible with saying that Kimi was heavily distilled by Claude.

At this point, I'm just pointing out the many lies, fallacies, and failures to read at a high-school level that you're committing.

HarHarVeryFunny 7 hours ago | parent [-]

> You literally did not read my comment before responding or actually respond to any of the points.

Are you so unaware that you believe you are making any points?

Go back and read your post - it was just a bunch of insults with zero technical content to respond to. The same emotional hysterics you've been using in the entire thread.

throw10920 6 hours ago | parent [-]

> it was just a bunch of insults with zero technical content to respond to

Yet more lies. I made many substantial comments, and the fact that you're lying about that is you, not me.

You've already lied about my own words multiple times (e.g. when you said "it would certainly be highly ironic if this summary was in fact "even more valuable" for that purpose as you are claiming!") You're either an LLM with a bad hallucination rate or just evil, and precisely zero statements that you provide have any trustworthiness to them.

Keep discrediting your account. I know that I can't convince a propagandist, but you just keep on further damaging your own reputation and argument every time you respond like this and ignore evidence that I've linked :)

wasfgwp 2 days ago | parent | prev [-]

Well there were observed cases of Claude calling itself Deepseek or Qwen. So pot calling the kettle black?

To be fair I find it hard to take this too seriously, shouldn’t it be trivial to just replace “Claude” with any other string in your “distillation” dataset?

throw10920 2 days ago | parent [-]

> Well there were observed cases of Claude calling itself Deepseek or Qwen. So pot calling the kettle black?

There are open-source Deepseek and Qwen models - "distilling" doesn't involve breaking terms of service or hitting an API because you can literally run local inference or even just inspect the weights directly, and that's intended because they're open source.

It's categorically different for a nation-state to build massive illicit networks of fraudulent identities to do distillation over tens of thousands of accounts to intentionally bypass providers' terms of service, intention for their models, and business model that very explicitly proprietary and not open source.

https://www.chinatalk.media/p/how-to-buy-cheap-claude-tokens...

If Claude did distill on proprietary PRC LLMs - then fine, shame on them - I condemn that and I expect others to do the same. But there are no open-source Claude models. The only way for PRC models to have those responses is if they distilled Anthropic's models from their APIs.

> To be fair I find it hard to take this too seriously, shouldn’t it be trivial to just replace “Claude” with any other string in your “distillation” dataset?

...and what would happen when it read all of the books and articles about Anthropic and replaced "replaced Claude Opus" with "replaced Qwen Opus"? Did you give any thought to this at all before saying it?

wasfgwp 2 days ago | parent [-]

> It's categorically

The fraud part and using stolen accounts or credit cards or blatantly violating the terms and conditions (i.e. reselling subscriptions not using outputs in certain ways somebody might not like) is indeed categorically different.

Using uncopyrightable outputs of an AI model obtained legitimately to train your model is not inherently interlinked with any of those things. I don’t really see how the model being proprietary or “open” is particularly relevant when talking about the outputs.

Even using the word “distilling” in this case is deceptive and biased. It implies that the Chinese are somehow stealing Anthropic’s models or their weights and somehow directly transforming them into new models. That’s certainly not what’s happening in any direct sense.

e.g. what if I agreed to send all my Claude code session files to Deepseek or whoever? There would be nothing wrong about that since I and not Anthropic own those files and can do whatever I want with them. Using certain different ways to obtain them of course could be highly illegal.

8note 3 days ago | parent | prev | next [-]

idk. i think its fair use when anthropic trains off of copyrighted works, and theres no property rights at all related to the model outputs

there's no creative work between the weights and the tokens being made.

whats the big deal if chinese companies sell an exact replica of the model? its a summary of a variety of works of text and images

queenkjuul 3 days ago | parent | prev | next [-]

Like most things, i support rights for people, and not companies. Copyright was created for authors, artists, and inventors. Rules for corporations can and should be different.

bluegatty 3 days ago | parent | next [-]

This makes little sense, either from either a moral or pragmatic perspective.

You do realize the 'investors' are the one's who 'own' companies and therefore the IP?

blackqueeriroh 3 days ago | parent | prev [-]

Lolololololololol

One person can a corporation be.

lovich 3 days ago | parent | prev | next [-]

> It's also sadly hypocritical to see all this rhetoric here on this thread decrying SOTAs for using a variety of content for their inputs as somehow sTeaLINg sTufF! ...

> ... but Chinese SOTA foundries directly using distillation as fair game.

As someone who says it’s fair game, it’s less that I’m being hypocritical and more that I don’t care that one thief had their shit stolen by a second thief. I also wouldn’t care if someone distills the Chinese models. It’s just thieves all around and if they want legal protection or moral outrage from the common man then my view is that they should stop stealing first.

wasfgwp 2 days ago | parent [-]

Why are you calling the Chinese companies “thieves” though?

LLM outputs are not copyrightable (or rather the user is effectively the only one who can own it). It would be problematic if Anthropic owned all the software generated using Claude..

lovich 2 days ago | parent [-]

They are thieves the same way Anthropic or OpenAI are thieves. Either it’s fair use to learn from this data or not.

If what Anthropic/OpenAi did for training is theft then the Chinese models also are a form of theft. If they didn’t steal then I don’t think the Chinese firms did either.

wasfgwp a day ago | parent [-]

Following that logic anyone using an LLM trained on “stolen” data for any purpose whatsoever is engaging in theft. Assuming the Chinese companies are obtaining their traces legitimately (which is a different question) what they are engaging in is morally no different than what the overwhelming majority of people on this site are regularly doing.

wasfgwp 2 days ago | parent | prev | next [-]

Do you think that Anthropic should own any outputs you generate with their models?

If not there is not there is no grey zone whatsoever.

dijksterhuis 3 days ago | parent | prev | next [-]

> it's hilarious to see how all these HN intellectuals, who on any other day are rabidly against most DRM / IP issues all of a sudden become private property absolutists.

for the record, i've always been rabidly pro-copyright since i worked at a performing royalty organization (prs for music) circa 15 years ago, way before i joined hn.

i don't use llms for that reason.

> It's also sadly hypocritical to see all this rhetoric here on this thread decrying SOTAs for using a variety of content for their inputs as somehow sTeaLINg sTufF! ...

when i see a spade, i call it a spade. just because the US has utterly stupid copyright provisions that are wide open for abuse, i.e. fair use, doesn't mean abusing those provisions at scale is morally acceptable.

> ... but Chinese SOTA foundries directly using distillation as fair game.

two wrongs don't make a right, but the irony is at least something.

> I don't think there is any coherence to any of these arguments - other than 'we liking big companies'. That's the only common thread.

the corpos can get fucked as far as i'm concerned.

> What is more reasonable: ... There's some grounds for fair use by SOTA models to ingest content, so long as they are not reproducing it ... very roughly speaking.

*only in the US.

bluegatty 3 days ago | parent | next [-]

I'm not attacking you, you don't have to defend yourself.

I'm just nothing that HN rhetoric is contradictory.

But this:

"when i see a spade, i call it a spade." -> this is anti intellectual absolutism.

If it were some true injustice, then fine, but that is clearly not the case.

There is ample room to contemplate that even copyrighted works could be considers fair use as training material.

"the corpos can get fucked as far as i'm concerned."

Ok that's fine - but then don't expect anyone to respect your principles if you don't have any other than 'screw that group!'.

I'm sympathetic to it (!!!) - but if we want to call a 'spade a spade' in a legitimate way, then we can do it in consistent and principled way.

throw10920 3 days ago | parent | next [-]

> I'm just nothing that HN rhetoric is contradictory.

Periodic reminder that HN is not a collective or a singular entity and is actually a bunch of different people with different opinions. Often the people with the loudest opinions get upvoted to the top - and often the "side" represented at the top is different from thread to thread.

dijksterhuis 2 days ago | parent [-]

also, human beings themselves can be contradictory. as an example i will happily pirate films/tv shows, but refuse to do the same with music.

dijksterhuis 2 days ago | parent | prev [-]

[dead]

wasfgwp 2 days ago | parent | prev [-]

> two wrongs don't make a right

There is no second “wrong” here.

Model outputs are not copyrightable. I think that was already established?

Or do you think that Anthropic should own all the code generated by Claude? Surely that would be somewhat problematic?

If Anthropic feels that some of their customers are breaking their EULA (nothing to do with copyright infringement though) they are free to stop doing business with them. Maybe even sue them in civil court for breach of contract (again nothing to do with copyright infringement though)

dijksterhuis 2 days ago | parent [-]

there's a difference between a moral wrong and a legal wrong. in a heavily simplified view, moral wrongs are usually decided in the eyes of the victims -- you did bad thing to me so i'm not going to talk to you anymore. legal wrongs are decided by courts -- you did a bad thing so this court has decided you're not allowed to talk to that person anymore.

anthropic are essentially saying in this tweet they believe a moral wrong has been committed against them -- "unacceptable behaviour" etc.

plenty of people have been vocal about the fact anthropic have committed moral wrongs at scale in building the products in the first place, with the question of legal wrongs still being worked out.

so, two moral wrongs. legally, fuck knows.

wasfgwp 2 days ago | parent [-]

Sure but it’s hard to read what Anthropic is saying in any other way than that they think that it’s morally wrong to engage in any behavior that harms their (potential) profit margins. The exact phrasing is just a way to justify their stance to other people.

I mean you are right in a way of course, it’s just a matter of degree and perspective, though. If one thing is moderately morally wrong and the other is potentially lightly morally wrong I don’t think it’s fair to equate them.

To me the situation is a bit like Google coming out and saying that its morally wrong for someone to build a competing open operating system on top of Android while stripping all Google services and “stealing” their ad revenue. Just seems silly and hypocritical.

m4rtink 2 days ago | parent | prev [-]

I think people just react to the hypocrisy of corporations stamping on people for "copyright violations" only for (often the same) corporations to blatantly obtain any data they can find, totally disregarding any licenses, scrappers overloading web sites or even privacy.

And the result is force feeding an AI slop generator with a subscription while making personal hardware 3x+ times more expensive.

No wonder people are fed up with this behavior.

wasfgwp 2 days ago | parent | prev | next [-]

LLM outputs are not copyrightable or rather the user who generated them owns it.

That entirely settles it and there isn’t much else to say about.

If Anthropic feels that other countries are violating their EULA well they are free to stop doing business with them.

spwa4 2 days ago | parent [-]

The claim these companies make goes much, much further than that.

They claim LLMs "uncopyright" their inputs. So if I take, say, 50 Mickey Mouse comic books, tell ChatGPT to read them and produce 50 "Buster Beagle" comic books that there is ZERO "copyright contamination" and I own those 50 output comics without Disney having any claims on them whatsoever.

Or if I ask ChatGPT to "make a spreadsheet software like Excel, Sheets, Calc, ..." that, again, there is zero copyright claim possible from these people.

It has not been tested, of course.

sho_hn 3 days ago | parent | prev | next [-]

Are you suggesting data input is further from IP than distillation?

That would stun me, but it's a little hard to read.

mkehrt 3 days ago | parent | prev [-]

Do you think writing books (and Wikipedia articles, and stack overflow articles, and github repos, and, and, and, and ...) is not a value add?? What terrible claim.

remus 3 days ago | parent | prev | next [-]

While I agree on a moral level, I think there is a distinction to be made. Training a SOTA model takes a huge amount of resources and expertise so the people doing the training are adding a lot of value along the way. I think this is much less true for distillation (which is kind of the whole point).

ed: to clarify, I totally agree that a huge chunk of the value in LLMs is coming from the source material. My point was just that training an LLM takes more resources and expertise than distilling from an existing LLM so I don't think the equivalence between training and distilling is entirely justified.

Bratmon 3 days ago | parent | next [-]

I like this comment because its argument only makes sense if you assume that the entire world's output of books and art did not require a huge amount of resources and expertise to make, nor did it add any value.

It's the most CS-major take ever!

mapontosevenths 3 days ago | parent | next [-]

If turning other peoples copyrighted work into a model is transformative enough to be protected then so is distilling that model into a different, better, model.

foo12bar 3 days ago | parent [-]

The models were built using copyrighted works, so why can't models be built using other models?

usef- 3 days ago | parent | next [-]

They do seem to be paying for it (as per the 1.5Bil lawsuit yesterday and them now purchasing books and licensing from media companies).

Whether we think they're paying enough is another question, but "I'm paying for content so can protect it" doesn't seem inconsistent.

We may decide that giving models away for free means they don't have to license content (judging by HN comments), but currently that doesn't seem to be the case as Meta is facing lawsuits for its open models.

(Obligatory stratechery piece: https://stratechery.com/2026/whos-afraid-of-chinese-models/ )

trhway 3 days ago | parent [-]

The judge found their use is fair use. They are paying not for their use of the content, they are paying for using illegal copies of the content.

The same principle can be applied to distillation - it is a fair use. You just shouldn't use illegal ways to access the models being distilled.

To the commenter below: if it is illegal - has the police/FBI report been made? Otherwise it is just a civil court matter.

usef- 3 days ago | parent [-]

Fair, but isn't "illegal" access what they're talking about in OP?

It does seem to be becoming the norm for AI companies to licence premium content in America, judging by the deals they're making. It doesn't seem to be done by the international distillers. It's a cost that American open models will seem to have to pay but not international.

trhway 3 days ago | parent | next [-]

>It does seem to be becoming the norm for AI companies to licence premium content in America, judging by the deals they're making. It doesn't seem to be done by the international distillers.

International distillers doesn't use that premium content, so they don't pay for it. They do pay for their access to the models they are distilling. Thus providing the revenue stream to those models. Thus those models make profit off the content they used for training. The content they mostly have't paid for.

>It's a cost that American open models will seem to have to pay but not international.

It goes both ways - American companies and their business are protected by American laws and have access to the market protected by those laws, etc.

usef- 3 days ago | parent [-]

> International distillers doesn't use that premium content, so they don't pay for it.

This doesn't seem to be true. They are training on their own scraped data overwhelmingly (we can extract copyright data from, eg, deepseek). They couldn't get nearly enough tokens through the American APIs to train a model on alone.

> American companies and their business are protected by American laws and have access to the market protected by those laws

Absolutely. Currently international providers are selling inference on the American market though, I don't know how that will sit legally the way things are currently going.

Bratmon 3 days ago | parent | prev [-]

> It does seem to be becoming the norm for AI companies to licence premium content in America, judging by the deals they're making.

This is a very surprising claim to me (and I imagine many small website owners who keep getting scraped by Anthropic and OpenAI).

Do you have a source?

usef- 3 days ago | parent [-]

There's been many news stories of it over the past year(s) as they signed each one. Here's the first result I could see with a rundown of many of them (am on mobile).

https://digiday.com/media/a-timeline-of-the-major-deals-betw...

Bratmon 3 days ago | parent [-]

Those are licenses for API access to data too new to be in the training data (for use by agents), not for the training itself.

I don't really understand why you think they're relevant, given that this conversation is about the training itself.

usef- 3 days ago | parent [-]

It isn't just that, it includes publishing houses, Wiley etc, and non-live media.

Even the news orgs say explicitly in the press releases that it's about training on their archive

eg. http://ap.org/media-center/press-releases/2023/ap-open-ai-ag...

---

edit, examples:

Wiley https://newsroom.wiley.com/press-releases/press-release-deta...

Shutterstock https://investor.shutterstock.com/news-releases/news-release...

Axel Springer https://openai.com/index/axel-springer-partnership

Stack Overflow: https://stackoverflow.co/partnerships

Disney (for characters in video. Video is especially where licensing is a big difference internationally right now) https://openai.com/index/disney-sora-agreement

etc.

The news corp one had a leaked price ($250mill), so they don't seem to be insignificant. These would have to be included in API prices I presume.

breppp 3 days ago | parent | prev [-]

Because model output is probably far closer to software or a licensed work which possibly has greater protections than it is to copyright. There is far less possibility of fair use, it might be protected by patents, license or reverse engineering laws.

In any case the laws are being written now, but I doubt these will have worse protection than software does, which has far better protections than copyright

giaour 3 days ago | parent | next [-]

> I doubt these will have worse protection than software does, which has far better protections than copyright

Software is protected by copyright. Some software may also be protected by patents, but last time I checked, AI generated output of any kind was not patentable.

trhway 3 days ago | parent | next [-]

Distillation isn't a copy. Distillation is more akin to "clean room" implementation.

Also note that the OpenAI/Anthropic argument is that the model training is sufficiently transformative to satisfy the fair use of the original content for training.

By that same argument, when distilling the distillers aren't using the original content the OpenAI/Anthropic models were trained on - the distillers are interacting only with the "sufficiently transformed" content of the OpenAI/Anthropic models and are normally paying for that.

There is also that old phonebook rule that facts can't be copyrighted. So, if i asked the model about bunch of phone numbers, i can publish the resulting list, can train my model on it, etc. Such approach doesn't allow to reproduce copyrighted works of course - and as we know the AI output isn't copyrightable, so it looks like basically any output i get i can use whatever way i like.

breppp 3 days ago | parent | prev [-]

Software is protected by the DMCA, patents, licenses, EULAs, all of those aren't there for books. I doubt new laws won't be written for model outputs.

Also, if model output distillation is shown as some form of reverse engineering I assume the DMCA can apply

bigiain 3 days ago | parent | next [-]

The C in DMCA stand for Copyright. All (I think?) software licenses are underpinned and made legally enforceable by copyrights. EULAs are underpinned by licenses which are founded on copyright. Patents are the only one of those protections that are not based on copyright, and there are lots of very good arguments against at least most software patents (all software patents of the form "Do {well known and obvious thing} with a computer" should, in my opinion, be immediately revoked and potentially have every company who's enforced payments from such patents investigated for fraud).

giaour 3 days ago | parent | prev | next [-]

You may recall that the DMCA was originally written to protect music and movies. It does in fact apply to creative works. If you have ever purchased an MP3, eBook, or streaming movie, you will also be aware that you purchased a license to the underlying IP. This is also true of physical media, but the license agreement you have to accept when obtaining a digital work makes this explicit.

I agree that you can't patent a book, but I would point out that you can patent an idea, which may only appear in a book or journal article.

vel0city 3 days ago | parent [-]

You do patent ideas, but the actual words written in a book describing that idea would only be protected by copyright at best. FWIW, the exact words describing the idea being patented are technically public domain; that's the whole point. You're free to go look up that patent, print it out, make whatever copies of it you want. Take any of the drawings in patents, put them on t-shirts, and sell them. No problem. Implementing the ideas those words represent is a different story.

For example, a patent describing a chemical process. The actual idea of how to do it is public domain, go look up the patent. Print it out. Do whatever with those words. Its fine. Building a plant to go do that chemical process to make that same output chemical in that same way, that's IP infringement. Its not the words, its the idea.

wasfgwp 2 days ago | parent | prev | next [-]

How is “model output distillation” different to using outputs (which are legally copyrightable) for any other purpose?

queenkjuul 3 days ago | parent | prev [-]

Afaik (and ianal) there's nothing stopping anyone from attaching a EULA to a physical book

vel0city 3 days ago | parent | prev | next [-]

Let's assume model output can be claimed by copyright or some form IP. You can't really patent it, as the output isn't a novel idea or process, much like you don't patent a book or a movie. But for arguments sake, let's agree it is some kind of IP.

Who are you saying owns that IP? The people who trained the model? The people who ran the model? The people who wrote the prompt? The person who paid for all of that to happen?

If the model output is owned by the person prompting it and paying for the tokens, what's the problem here?

If the model output is owned by the trainer of the model, that's a big nasty can of worms.

preg_match 3 days ago | parent | prev | next [-]

Why would this be the case. Why would software output from a model magically have greater protection than the software the model trained on.

wasfgwp 2 days ago | parent | prev [-]

LLM outputs are not copyrightable. At least that’s the current established legal precedent in the US. The only question is whether the user owns the copyright without significantly transforming the output but that’s not really relevant in those specific situation.

I mean otherwise it’s a very slippery slope, effectively it would give Anthropic the ownership of any code generated by its models..

arbitrary_name 3 days ago | parent | prev | next [-]

there is a major god complex here.

MBAs and non technical managers = inept Catbert-type charlatans.

Software engineers, devs, etc = geniuses capable of mastering any domain, innate ability to be right on any topic.

joshuamorton 3 days ago | parent | prev | next [-]

I don't think that's what it's saying at all. It's saying that there's a level of creativity in model creation that isn't present in distillation.

skybrian 3 days ago | parent | next [-]

Maybe, but it's not like their AI is likely to repeat it back verbatim so it's unlikely to be a copyright violation. It seems like at most, they would be breaking Anthropic's terms of service?

Or maybe they're going through an intermediary "transfer station" that's breaking terms of service:

https://www.chinatalk.media/p/how-to-buy-cheap-claude-tokens...

perching_aix 3 days ago | parent [-]

Yes, it's just a ToS violation at present. Those are legally binding though, despite the common adage. What that really translates to here though, anyone's guess.

Anthropic's own copyright infringement could apparently be forgiven for 1.5B USD after all, so maybe there's a price that breaking the distillation clause for is acceptable too. Or some other arrangement.

Bratmon 3 days ago | parent [-]

But surely at least one of the websites Anthropic scraped to make Claude had a ToS forbidding automatic access?

Why is Anthropic's ToS any more binding than that of a rabidly-anti-ai literature blog with 50 readers?

skybrian 3 days ago | parent | next [-]

One reason is that they might not have scraped it themselves, so if there was a ToS, it was someone else who broke it. For example, see:

https://en.wikipedia.org/wiki/The_Pile_(dataset)

Another reason is that if you can download a web page without agreeing to a ToS, I'm not sure that counts as one?

queenkjuul 3 days ago | parent | prev | next [-]

> Why is Anthropic's ToS any more binding than that of a rabidly-anti-ai literature blog with 50 readers?

I mean i know you know the answer: anthropic is a corporation with lawyers on retainer, and that's really all that matters

perching_aix 3 days ago | parent | prev [-]

Do feel free to read the court documents to find out and let us know.

> Why is Anthropic's ToS any more binding than that of a rabidly-anti-ai literature blog with 50 readers?

Although I will say, this whole comparison stuff really doesn't seem to be your thing; might impede your analysis quite a lot: https://news.ycombinator.com/item?id=49013148

Maybe ask Claude?

trhway 3 days ago | parent | prev | next [-]

>a level of creativity in model creation that isn't present in distillation.

the same argument - a level of creativity in the world knowledge creation that ins't present in the model training on that knowledge.

Or in other words - model creation and training is just a distilling of the world knowledge.

joshuamorton 3 days ago | parent [-]

I don't disagree. I'm not sure why that's a relevant reply though.

If you think that the addition of a less creative process (model creation) to a more creative corpus ("art") is problematic, then it follows that you should think the addition of a less creative process (distillation) to a more creative corpus (a model) is also problematic.

trhway 3 days ago | parent [-]

I think both are natural and fine. Otherwise we'd have to outlaw analytical thinking.

remus 3 days ago | parent | prev [-]

Yes, this is what I was getting at.

gozucito 3 days ago | parent [-]

There is an even higher level of creativity in creating books, songs and all sorts of art used in model training though. That's your apparent blindspot.

There is no world in which me vacuuming the entirety of human knowledge to make a genai model is ok but hoovering my model answers is not. The hypocrisy is stunning and risible.

Now if you go and make a model based on purely synthetic data and not a single work made by humans, you would have a valid point.

remus 3 days ago | parent | next [-]

> There is an even higher level of creativity in creating books, songs and all sorts of art used in model training though.

No argument here, I completely agree.

> There is no world in which me vacuuming the entirety of human knowledge to make a genai model is ok but hoovering my model answers is not.

I disagree with this though. Clearly LLMs owe a huge debt to everything that has come before, but surely you'd agree that the models that are produced are something substantial and new and novel which didn't exist before and have lots of value in their own right. Let's be a bit reductive and pretend Moonshot had just outright stolen the weights from Fable somehow, clearly that wouldn't be contributing anything really new or novel. Now of course they've distilled rather than stolen, but the point is similar: how much value have they added along the way?

gozucito 2 days ago | parent [-]

Since this is HN Think of it like one of the GPL license for software.

It's ok for me to use your source code for free as long as I then let others also use my source code for free.

joshuamorton 3 days ago | parent | prev [-]

So, I'm not the person you were responding to. I'd like you to take a moment and suggest where anyone in the thread you're replying to, either me or Remus, has said anything that suggests disagreement with the statement

> There is an even higher level of creativity in creating books, songs and all sorts of art used in model training though.

He claimed there was more creativity in model training than in model distillation. That makes no claim about the relationship between the creativity in model creation and art. Why are you continuing to attack a claim that was never made, after a sub thread very explicitly clarifying that that claim was not made?

gozucito 3 days ago | parent [-]

This is the post Nemus was replying to:

>Perhaps even more importantly, the current frontier LLM models are self-admittedly the product of enormous quantities of copyright infringement and even less savory inputs, so calling them out for distilling the fruit of that tainted tree reads as highly hypocritical at best.

Context is important. And in this context, their argument only mentions creativity when it belongs to an AI lab. That omission is the blind spot I pointed out. Bottom line is whether or not Anthropic are being hypocritical and yes, they most definitely are, regardless of any attempted sophistry.

There is a reason courts want you to tell "The whole truth" and not just "the truth".

perching_aix 3 days ago | parent | prev | next [-]

No? They outright say the opposite!

Like look, I'm not a native speaker, sure. But I think when someone says "value add", that means there was value there (which you claim they're rhetorically erasing), and then that was added to. Under no interpretation of this phrase do I get an erasure of prior value.

So certainly, as long as words mean anything, no, they absolutely did not say or suggest what you claim they did, and what you extract a thus unreasonable amount of obnoxious schadenfreude from, while throwing in a cheap insult for funsies at the end.

It's the second time I feel compelled to reach for this just today: https://i.kym-cdn.com/photos/images/original/002/659/979/108...

bluegatty 3 days ago | parent | prev [-]

This is a misrepresentation though.

The LLM output, is not the same as the input - there is value add.

Of course works used as raw inputs to LLMs required work and are reasonably subject to IP concerns - but they are different.

It's possible that the LLM makers 'owe' the content creators that created the content they used to make their products - it's an interesting but separate question.

We could very well end up where content IP is protected, LLM output is not and visa versa with reasonable legal founding, doubtful but plausible.

blks 3 days ago | parent | next [-]

Lossly storing IP in LLM itself, and using IP for training (so it’s lossly stored in LLM), without licensing these works or otherwise following license agreements (eg GPL) is infringement. Using then this product for commercial activity is a smoking gun.

bluegatty 3 days ago | parent [-]

"Lossly storing IP in LLM itself, a" - that part I'm inclined to agree with.

But it's debatable if that's the case.

Google stores copyrighted content and produces in in their product.

Also - it's fair game to use snippets of things here and there, if the derived work is novel, which I think it is for LLMs, mostly.

I do agree though, that we ought to draw the line somehow.

jaggederest 3 days ago | parent | prev [-]

> but they are different.

How, and why?

> We could very well end up where content IP is protected, LLM output is not and visa versa with reasonable legal founding, doubtful but plausible.

That is the current state of legal rulings - LLM output is public domain, not copyrightable.

semiquaver 3 days ago | parent | next [-]

This misstates the small number of legal opinions and orders on this topic, none of which form binding precedent outside the districts where the cases happened. So even if a court had found that “LLM output is public domain” (none did) that wouldn’t make it “the law” until it went up the appellate system and was upheld.

Our current laws simply weren’t built for this and I expect the legal status of LLM output is not going to be resolved until Congress actually legislates on this topic.

bluegatty 3 days ago | parent | prev [-]

"> but they are different.

How, and why?"

How are they even remotely the same?

They're not even used the same way.

One is raw data input, the other is training content - designed to train LLMs.

One is a set of IP derived for other purposes entirely, and has esablished IP law - how you can use someone else's creative work or not ... for LLM outputs, less clear.

skippyfish 3 days ago | parent | prev | next [-]

> raining a SOTA model takes a huge amount of resources and expertise

Writing books, building Wikipedia, and answering questions on online forums takes a lot of resources and expertise that scraping didn't. So at the very least, we're already one rung down the "maybe you should've asked" ladder.

jaggederest 3 days ago | parent | prev | next [-]

I suspect that, in aggregate, all of the informational output of humanity prior to 2020 has taken more resources to produce than the last few years of LLM research.

ryandvm 3 days ago | parent | prev | next [-]

I don't know man. This reads like "yeah we stole your grain, but making bread is hard."

Muromec 3 days ago | parent [-]

It sure is, but it doesn't matter. Whatever position that generates more economic activity is declared legal using some nonsense retconned logic "because we said so".

altmanaltman 3 days ago | parent | prev | next [-]

Why is it less true for distillation? Everyone technically has access to Fable but Moonshot came up with the model. How can you objectively claim one is adding value while the other is not?

If that is the whole point you need to clarify why this is the case on an objective level.

I would say building a comparable model using any means necessary (just like what Anthropic and OAI did) at a lower cost is actually more valuable to soceity and Monshoot is arguably generating more value with less.

il 3 days ago | parent | prev | next [-]

Probably not as much effort as writing books and creating art the models were trained on.

AlienRobot 3 days ago | parent | prev | next [-]

The value of LLM's come from replacing what generated its training data.

If the distilled model is cheaper, then it's just LLM's getting LLM'ed.

GTP 3 days ago | parent | prev | next [-]

Still, AFAIK Kimi's architecture (just like that of other LLMs from Chinese labs) is different from those of OpenAI and Anthropic's model in a nontrivial way. So the expertise is still there, and I guess resource use too (although Chinese labs tend to optimize this, thanks to the restrictions they have on GPU use).

EDIT: just wanted to add that resource optimization is usually where the contribution of Chinese labs is, so you shouldn't reaad the above parenthesis as a negative comment.

InsideOutSanta 3 days ago | parent | prev | next [-]

As an author, that's a genuinely disheartening thing to read.

It took me a year to write a book. It took OpenAI and Anthropic a fraction of a second to ingest it. Do you understand now why I give zero shits if it takes Anthropic a billion to train a model, and Moonshot 10k in API cost to distill it?

3 days ago | parent | prev | next [-]
[deleted]
liuliu 3 days ago | parent | prev | next [-]

> training an LLM takes more resources and expertise than distilling from an existing LLM

This is not automatically true. Training and distillation use the same underlying infra and method and there is no intrinsic differences in between.

blks 3 days ago | parent | prev | next [-]

They add value on top of other people’s work, often against licensing, and then commercialize this product, ie profiting from making a product out of other people’s IP.

3 days ago | parent | prev | next [-]
[deleted]
Barrin92 3 days ago | parent | prev | next [-]

>Training a SOTA model takes a huge amount of resources and expertise so the people doing the training are adding a lot of value along the way.

producing the entire body of human knowledge that Silicon Valley companies absorbed like the Borg did not just take more resources but also a fair amount of blood and sweat, certainly more than the LLM so on that front that comparison also seems entirely justified.

Teever 3 days ago | parent | prev | next [-]

I'm sure it takes a lot of time and resources to plan and pull off an epic heist but it is unusual to see people like Thomas Crown being accused of creating value, as they're usually accused of committing theft.

mort96 3 days ago | parent | prev | next [-]

Yeah, there's a difference. One party spends a bunch of resources doing something illegal and extremely immoral. The other party spends little money doing something legal and morally neutral.

darod 3 days ago | parent | prev [-]

You can argue that reverse engineering anything is as hard if not harder than engineering something. I can’t imagine distillation is any different.

some_random 3 days ago | parent [-]

Distillation is objectively easier than training a model from scratch, that's why all these Chinese labs are doing it.

recursive 3 days ago | parent [-]

Training a model is objectively easier than generating the sum total of human creative output prior to 2020. That's why the big labs are doing it. What's the difference here?

breppp 3 days ago | parent [-]

Reverse engineering has stronger protections than merely copyright infringement

petilon 3 days ago | parent | prev [-]

I disagree that LLM models are the product of enormous quantities of copyright infringement.

The recent announcement that AI-assisted research produced a counterexample to the Jacobian conjecture--a long-standing open problem in algebraic geometry--shows the original value AI can create. The result was not copied from a textbook; it emerged from AI learning from existing material, much as a human does, and then applying that knowledge in a new way. If that's a violation of copyright, then a human doing the exact same thing would be a copyright violation too. But it isn't.

InsideOutSanta 3 days ago | parent | next [-]

If you re-read your comment, you will find that your second paragraph is not evidence for the claim you make in your first paragraph. In fact, your first paragraph is just false.

petilon 3 days ago | parent [-]

Let me explain it this way: If it is legal for a human to learn from a book, then disseminate the knowledge, then it is legal for a machine to do so. You may think this is not right because a machine does it at a much larger scale, but if so laws need to be updated. As it stands now there is no law that says if a human does X it is not a copyright violation but if a machine does the same X it is copyright violation.

worik 3 days ago | parent [-]

> If it is legal for a human to learn from a book...

True, if the human's access to the book was legal

A great deal of training was on the open web, no one should complain.

But at least Meta and Anthropic were caught red handed taking copyrighted works, illegally, for training

I think international IP laws are too strick and onerous, but they were broken to train these models

butlike 3 days ago | parent | prev | next [-]

Yup if a machine kills a human it's not the machine's fault; it's the human's. Humans doing the exact same thing as machines aren't 1:1.

petilon 3 days ago | parent [-]

If it is legal for a human to do something then it is legal for a machine to do it too. Are there any counter examples to that?

yencabulator 3 days ago | parent | next [-]

Get elected president?

fooofw 3 days ago | parent | prev | next [-]

Lethal self defense?

GTP 3 days ago | parent | prev [-]

The GDPR gives you the right not to be subject to automated decisions, so there are cases where a human can make a decision and a machine cannot.

ilovecake1984 3 days ago | parent | prev [-]

The didn’t pay for the books.

It’s massive copyright infringement.

The human buys the books.

foxglacier 3 days ago | parent | next [-]

It's worth keeping in mind the purpose of copyright. It's a pragmatic tool to encourage investment in creative work for the benefit of everybody/consumers. We may be entering a time where there's less need to incentivize people to write books. At least not non-fiction books which are simply a collection of existing knowledge presented in an a way that's suitable for human readers. A lot of the value those authors provided can now be done by AI. Yes, the AI trained on their work, but now that it's here, we don't need new non-fiction authors quite as much as we used to.

I wouldn't want to live in a world where technology or general people's wellbeing was held back by obsolete laws that ended up lingering on just to protect undeserving special people at the expense of the rest of society. Remember guilds for tradesmen? They were also a monopoly given by the government to special people. They had their purpose but nowadays we have different ways to keep tradesmen working effectively like license requirements and insurance.

Just to be clear, I think we do still need copyright, but that we might be in a transition period where it has to be redesigned to adapt to AI.

ilovecake1984 3 days ago | parent [-]

I don’t think it’s been common to write non fiction for money for decades. What they are doing is killing off the real motivation to do it, which is recognition and attribution.

We will all be sorry when professionally written and edited works disappear. An author has a reputation and the incentive to protect that reputation keeps standards high.

petilon 3 days ago | parent | prev [-]

Did they borrow the book? If I learn from a borrowed book is that copyright infringement?

ryandvm 3 days ago | parent | prev | next [-]

Boy I tell you, I am having an awful hard time summoning pity for the organizations that have themselves distilled all of humanity's knowledge into mysterious labor-market-masticating black boxes.

matheusmoreira 3 days ago | parent [-]

https://news.ycombinator.com/item?id=47567575

> protect our first-party products from abuse like bots, scraping

Won't you think of the trillion dollar corporations?!

atleastoptimal 3 days ago | parent | prev | next [-]

It matters because everyone imagines the inevitable "closing of the gap" between closed and open source, but the rate at which open source catches up with closed source seems to depend on being able to train on and distill the outputs of open source models. As long as performance of open source models is at least partially dependent on frontier-model outputs, then that gap will remain in place by definition.

>Claiming in any shape or form that fable disillation is one of the primary reasons why kimi k3 is so competitive is slandering the work of other labs that cooperatively push the open-source models forward.

If the distillation is irrelevant to why it is competitive, why do they do it then? Obviously is helps improve their benchmarks/performance to some degree, otherwise they wouldn't need to do it.

himata4113 3 days ago | parent [-]

Never claimed that it is irrelevant. And kimi k3 is on the same level and sometimes outperforms fable 5 - that cannot be explained by distillation. The reason why gap is not closed is simply the fact that fable was trained months ago so in theory the frontier labs are still 1 (small) step ahead.

Although I will reiterate the fact that distillation is not the primary reason why these models are performing so competitively.

atleastoptimal 3 days ago | parent [-]

Closed-source models have to deal with the current frontier being heavily regulated. Fable, at its old level, was "too good" to be released and they had to add an additional safety layer to sanitize the outputs. Lowering the quality of the models so they are safer and more steerable has been something all the closed-source models have been doing for a while, a requirement that many open source models don't need to deal with.

If Kimi k3 really were above Fable 5 then there invariably the USG would have to consider their restrictions on model capabilities excessive, or one would have to admin closed source models are held to more restrictive safety standards than open source models.

>Although I will reiterate the fact that distillation is not the primary reason why these models are performing so competitively.

How would you know this? How could you ascertain exactly how much performance is attributable to their unique engineering/research? If they really were so competitive they could surely make a model that isn't dependent on distilling Fable or other frontier models.

himata4113 3 days ago | parent [-]

I recommend reading some of their research it's honestly astonishing how intelligent some of their solutions are.

Kimi specifically relies heavily on reasoning traces which is largely due to their training strategy and will perform poorly when thrown into a conversation from another model. Another fun advancement is that they simply ctrl+c ctrl+v'd attention which means that the model can steer where to look in the context window without ever producing an output token increasing token efficiency and attention accuracy as a side effect you end up with weaker prompt adherence.

None of these 'issues' manifest in US models which proves that kimi has diverged and is achieving these capabilities seperately from the architecture that US labs rely on.

I would agree with you during the Deepseek R1 era, but US labs were heavily inspired by open research at that point as well so I wouldn't give them too much credit.

spwa4 3 days ago | parent [-]

You still left out that it doesn't matter anymore, just like Anthropic/Facebook/OpenAI only really needed to read massive amounts of copyrighted data only once (and of course, they all did this illegally, which makes their current complaints more than a little ...). Once they have a large model trained on the data, they can just retrieve reasoning traces and copyrighted data from the previous model. In fact that is a training technique long used because it has better results that directly training on the original data.

In other words: even if the US (somehow) denies them access to the current OpenAI/Anthropic models, they'll be able to improve based on what they already have.

kevinqi 3 days ago | parent | prev | next [-]

I agree distillation isn't illegal; I also think Moonshot/Kimi is very impressive. But the more interesting question is whether labs like Moonshot can be a real competitor to OpenAI/Anthropic. If you can only play catchup (however quickly you do that), then you're never going to be at the frontier - I think that's why distillation matters.

mring33621 3 days ago | parent | next [-]

People that think the Chinese are only able to copy western tech are in for a wakeup call.

Actually, that has already happened in many domains, it's just that most western people (USA especially) won't admit it.

kevinqi 3 days ago | parent [-]

my assertion isn't that china isn't able to surpass western AI. I think it may well happen. I've been to china many times and am well aware of how ahead they are in many technological/societal areas.

at the same time, I don't buy the idea that distillation is unimportant in assessing what Chinese labs are capable of. If it wasn't, why did Kimi's release timing coincide so well with Fable's launch?

and if Anthropic hadn't released Fable, would we have Kimi today? If the answer is no, then I think that's still a very important point to consider.

ericmay 3 days ago | parent | next [-]

Spot-on and you're asking the right questions. And the other problem in these comments is that folks seem to think if China pulls ahead we can't just distill their models, provided that distillation is a key part of "this". If it's such a great strategy we'll just use it too if we want to. Boom roasted.

For some reason folks seem to think that China can take action and then other countries can't also take action or respond to that action and it comes up again and again. China has hypersonic missiles! Pack it up boys time to go home. Nothing we can do. Dang shucks. China distilled American AI models, welp time to just close it all down and let's just write off those trillions of dollars and all the literal geniuses financing and building these things. Oh well China can just copy American models while we spend all the money! Ok we just stop developing models and we'll just copy their models. China will flood the market with their cheap products! Nope can't do anything like, oh, idk, not buy any of those products or just raise the prices on them in local markets. It's never-ending. I don't understand the lack of capacity to reason about other actors that takes commonly takes place. And that's just China, never mind other general issues.

Daishiman 3 days ago | parent | prev [-]

> and if Anthropic hadn't released Fable, would we have Kimi today? If the answer is no, then I think that's still a very important point to consider.

That works both ways, competition and performance spur new developments. You don't think the American labs are looking at Chinese research on how to reduce compute per token?

nylonstrung 3 days ago | parent | prev | next [-]

So many of the breakthroughs and architecture that make LLMs powerful in general today came from China, especially ones related to sparsity and MoE that have made inference and training substantially cheaper.

Let's not forget how much people talked about "prompt engineering" before Deepseek mainstreamed the idea of thinking mode which is now universal

overgard 3 days ago | parent | prev | next [-]

I think it depends on where you think we are on the S curve of intelligence growth. (Yes, I think it's an S curve, not an unbounded exponential). If you think we're near the peak than playing catch up (especially if you can play catch up quickly) is very rational.

I know this isn't exactly a scientific test, but I had a local Qwen 3.6 27B model implement a fairly sizable feature today. There were a couple of bugs, mostly around me not giving sufficient specifications, but they were ironed out quickly when I pointed it out. I was able to ask the model to create instructions so next time it doesn't fall into the same pitfalls, and it did a great job. 27B local model! (And it was super fast too).

I ran Fable 5 as a code review and it didn't really have any significant corrections.

I guess my point here is that, for most work the frontier models are probably overkill anyway, and improving on overkill in a way that raises prices significantly is probably not a winning strategy.

The only place I can think of where the super high powered models are "required" is if you want to do a ridiculous token burn like GasTown where you just have it run un-monitored on very long tasks. To me though, that's an experiment, not a real workflow. And the way these labs are like "oh we made this (broken) thing in a week using just agents!" always also follows with "and it cost $100,000+ in tokens!". Like, ok, I get it if you're doing research but that's the salary of an entire person.. that can actually learn and improve.

overfeed 3 days ago | parent | prev | next [-]

> But the more interesting question is whether labs like Moonshot can be a real competitor to OpenAI/Anthropic.

The answer depends on whether you think the AI researchers at Chinese labs are (or can be) as smart, motivated, and as good at math as those working at US labs - a not-insignificant proportion of whom are Chinese nationals.

himata4113 3 days ago | parent | prev [-]

My entire point was that this was not achieved purely from distillation and claiming that is slander against open research.

mNovak 3 days ago | parent | prev | next [-]

> The only argument they have here is that they use GB300 GPU's which for some reason should not be available to chinese citizens

Note that Chinese companies are free to rent from GB300 clouds internationally. There are large datacenter hubs in Singapore and Malaysia serving chinese and other customers.

Though there is also reported [1] significant smuggling of Nvidia chips into China as well.

[1] https://epoch.ai/publications/chip-smuggling

JKCalhoun 3 days ago | parent | prev | next [-]

Legal, illegal…

The word I would use is inevitable. It reminds me of the (PC) clones wars…

Gajurgensen 3 days ago | parent | prev | next [-]

It is incredibly important to whether the US can maintain its AI lead. If foreign competition is closing the gap only by distillation, then the frontier labs can focus on preventing distillation and maintain their lead that way.

US dominance is also important for approaches to safety, especially political approaches. If the frontier models are all US-based, safety might be tackled via internal US policy. If other countries can independently train competitive models, international cooperation is required.

Edit: It is also important for the business model. Companies won't be able to justify tremendous training costs if competitors can replicate their product much more cheaply via distillation.

titanomachy 3 days ago | parent | next [-]

Do Americans even believe that US policy is likely to steer development in a way that’s safe and beneficial for humanity? The rest of the world certainly doesn’t. The US currently seems to primarily use their superpower status to be the world’s number one shit disturber and geopolitical antagonist.

I don’t think China’s necessarily any better, but I’d rather have the most powerful models be open rather than under the exclusive control of the US executive.

Gajurgensen 2 days ago | parent | next [-]

I didn't mean to imply that the US is more likely than elsewhere to responsibly steer AI via policy. But I do think it is easier if it can be done internally as opposed to via international dealmaking.

CuriouslyC 3 days ago | parent | prev | next [-]

China uses its power to make favorable deals and get people hooked on what it's slinging so it has captive customers. The US uses its power to bully and break rules that apply to everyone else for its own benefit. Kind of a big difference.

thesmtsolver2 3 days ago | parent [-]

Not really. Go ask someone in Tibet/HongKong/Taiwan/Japan/India if China isn't breaking rules and bullying them. If China had US's powers/economy, it would be a much bigger bully.

https://en.wikipedia.org/wiki/Wolf_warrior_diplomacy

https://en.wikipedia.org/wiki/Chinese_police_overseas_servic...

CuriouslyC 2 days ago | parent [-]

What China does with places they think are rebellious provinces is one thing. What the US did to Hawaii, Panama and half a dozen other countries when the local population tried to unyolk themselves from brutal US corporate robbery is another. The Chinese are hard driving businesspeople for the most part, the US is a gangster state.

thewebguyd 3 days ago | parent | prev | next [-]

> Do Americans even believe that US policy is likely to steer development in a way that’s safe and beneficial for humanity?

No, at least not outside of this forum.

We all mostly think these models and US policy are going to drive the exact opposite of that. Wealth will continue to get extracted and funneled to the top, and the rest of us are going to be left with the scraps and left to die while what little social safety nets we had continue to get eroded away alongside losing our jobs.

Gallows4574 3 days ago | parent | prev [-]

>Do Americans even believe that US policy is likely to steer development in a way that’s safe and beneficial for humanity?

No, no we do not.

realusername 3 days ago | parent | prev [-]

> If foreign competition is closing the gap only by distillation, then the frontier labs can focus on preventing distillation and maintain their lead that way.

We already know it's false because you would have hundreds of competitors if it was that easy.

The reason why these Chinese labs are releasing good models is simpler, they have access to a tremendous pool of talented people.

slibhb 3 days ago | parent | prev | next [-]

Of course it matters. Regardless of whether distillation is legal, there is a difference between training a model with and without distillation. For one thing, the distilled model wouldn't exist without the model it distilled.

Also, companies that use distillation may be competitive but seem unlikely to surpass the companies that are training these models from scratch.

bilbo0s 3 days ago | parent [-]

>but seem unlikely to surpass the companies that are training these models from scratch

Then why is it a problem?

Another serious question.

Trying to get my head around what the root of the objection is here. There must be some fear, but if that fear is not a fear of being surpassed in the market, then what is the fear?

overfeed 3 days ago | parent | next [-]

A fear of competition causing a failure to recoup the trillions invested in AI via sky-high margins, and starting a (short) chain-reaction that causes the bubble to pop (or fizzle). A lot of people have a lot riding on the AI bubble not popping.

slibhb 3 days ago | parent | prev [-]

I didn't say it was a problem, I said it "mattered"

If I wanted to argue that it's a problem, I'd just say that companies investing billions in training frontier models should reap the rewards. And distillation is essentially theft.

thewebguyd 3 days ago | parent | prev | next [-]

> Claiming in any shape or form that fable disillation is one of the primary reasons why kimi k3 is so competitive is slandering the work of other labs that cooperatively push the open-source models forward.

They claim it because Anthropic are planning to push for protectionism. They just doubled their political spending to $40 million for the midterms to "push for AI regulation" Gee, I wonder what it is they are lobbying for. Certainly won't be OFAC sanctions right? ICTS import controls?

US GOV, under lobbying pressure from Anthropic and OpenAI are going to go full protectionism and restrict Chinese models, I'd almost be willing to bet money on it. They can't really enforce for individuals, but they can definitely tell US based hpyerscalers they can't host them, make it illegal to host the weights, and government procurement restrictions.

torginus 3 days ago | parent | prev | next [-]

> The only argument they have here is that they use GB300 GPU's

I don't follow events closely, but the US has constantly flipflopped on what sort of GPUs the Chinese are allowed to have, not in small part because much of the AI boom's valuation is based on demand for US-made hardware, for which the Chinese have inexhaustible and well-financed demand.

So even this feels a bit hypocritical to me, but my understanding is that Chinese native AI hardware is getting good enough that labs dont feel a huge disadvantage by being forced to buy at home, even if they'd have preferred to buy US chips.

Which is a situation that was manufactured by the constant thread of having their access to advanced GPUs revoked.

pgt 3 days ago | parent | prev | next [-]

It matters because it means that lab could not train that model without distilling another frontier model, and their progress would slow once they get properly cut-off. If I funded that lab, I would want to know that.

GuB-42 3 days ago | parent | prev | next [-]

> Distillation is not illegal by every definition of the word.

I am waiting for a precedent on this one. In general, training on copyrighted material is legal, there is a lot of precedent there. But every now and then there is a case where the owner of the training material wins.

I don't remember the details but I believe one of these instances was when one company trained its AI on the knowledge base of another company and turned it into a competing product. Fair use was denied because of that direct competition. Distilling a LLM to make a competing LLM looks kind of like this, or maybe not, I don't know.

It would make sense for distillation to be legal in every way, LLMs are built on a broad interpretation of fair use, but sometimes, law is weird.

unknownfuture 3 days ago | parent [-]

> I am waiting for a precedent on this one. In general, training on copyrighted material is legal, there is a lot of precedent there. But every now and then there is a case where the owner of the training material wins.

You're making a fundamental assumption: that model outputs are subject to copyright. In the US that's only the case if a human is part of the creative process:

https://www.copyright.gov/newsnet/2025/1060.html

> It concludes that the outputs of generative AI can be protected by copyright only where a human author has determined sufficient expressive elements. This can include situations where a human-authored work is perceptible in an AI output, or a human makes creative arrangements or modifications of the output, but not the mere provision of prompts.

antisthenes 3 days ago | parent | prev | next [-]

It also doesn't matter for a simpler, and much grander reason.

All LLMs are trained on the corpus of humanity's knowledge, the legacy of everyone who's ever lived and our civilization as a whole.

Anything that prevents or circumvents the accumulation or gatekeeping of this knowledge and puts it in the hands of more people (that are not AI company shareholders) is a good thing. Whether that is done by open sourcing the model weights, the training set, or by making the output better and cheaper, it is all fair game and is, as another poster mentioned, inevitable in the long run.

XorNot 3 days ago | parent | prev | next [-]

It certainly matters as familiar sounding words to their stock holders to please not drop the valuation.

Because what they want them to think is "the AI factory has unique proprietary technology that cannot be replicated"

What they don't want them to think is "it's relatively easy once you know the basics to bootstrap to near SOTA and so the commercial case for selling inference has an extremely short profitability horizon with little if any brand loyalty or lock in".

nylonstrung 3 days ago | parent | prev | next [-]

I wouldn't be surprised at all if US labs are also distilling Chinese models, except we'd never know since they can simply self-host them

GuuD 3 days ago | parent [-]

We do know, because we used to have some Claude models identifying as Deepseek when prompted in Chinese

mattertoast 3 days ago | parent | prev | next [-]

It does matter in that these LLM companies need to be run into the ground, and every embarrassing clod working for them run out of town.

It's showing that 'distillation' is a viable way to reclaim all of what they stole and hoard, and with enough luck their debts will come due in time for them to feel it.

xienze 3 days ago | parent | prev | next [-]

> Does this matter? Distillation is not illegal by every definition of the word.

Correct, but it at least helps answer the question of "how do they make such good models for a fraction of the price???" The answer is someone else spends the untold billions and Chinese labs do a little tweaking.

insanitybit 3 days ago | parent | prev | next [-]

It is presumably against their ToS.

applfanboysbgon 3 days ago | parent [-]

And why, pray tell, would a cabinet member of the Trump administration be involving the US government in enforcing a private ToS?

thewebguyd 3 days ago | parent [-]

Whoever Anthropic just bought when they doubled their political spending to $40 million just recently for "Pushing for AI safety"

random_coder_nz 3 days ago | parent | prev | next [-]

It doesn't matter. It is most likely a pretext for upcoming actions mostly likely executed via yet another retarded executive order. The guy that posted this looks like he's drowned himself in the MAGA Koolaid.

petilon 3 days ago | parent | prev | next [-]

[dead]

smeeth 3 days ago | parent | prev | next [-]

Uh, what?

> Distillation is not illegal by every definition of the word

Note that Anthropic (and USG) alleges [0] not only that Kimi was distilled, but that they actively circumvented measures intended to stop distillation. There are multiple ways that's illegal, including:

- Civil breach of contract. Anthropic's TOS explicitly say you can't do what Kimi is alleged to have done.

- Economic espionage: 18 U.S.C. §1831 criminalizes obtaining a trade secret through theft, fraud, or deception while intending that it will benefit a foreign entity.

- Trade-secret misappropriation: if Anthropic could argue industrial-scale querying reconstructed proprietary aspects of Fable (like by showing it produces similar outputs, as others have done) then it's illegal under 18 U.S.C. §1832.

- California computer-access statute §502 bars knowingly accessing a computer system and, without permission, taking, copying, or using its data.

- Computer Fraud and Abuse Act protects against the case where restrictions against an activity are circumvented (like Kimi is alleged to have done).

> There are millions of samples available on huggingface and models explicitely trained on output produced by fable. There has been no action taken against them.

A lack of prosecution does not make something legal. There is also the scale/commercialization thing, which isn't an issue with random tiny HF datasets/models. Remember: Kimi also sells K3 inference.

> kimi architecture is vastly different than that of fable

How do you know that? Do you work for Anthropic? Also, this has nothing to do with architecture, we are talking about data.

> US AI labs are inspired by opensource advancements just as much as open source labs are inspired by traces from models such as fable.

Cool. The difference is that one of those things is legal (because they chose to open-source) and one of those things is illegal theft of trade secrets (because it was stolen).

> Claiming in any shape or form that fable disillation is one of the primary reasons why kimi k3 is so competitive is slandering the work of other labs that cooperatively push the open-source models forward.

1) this has nothing to do with other labs, just Moonshot (and Z.ai, MiniMax, DS)

2) slandering or not it happens to be completely true, so, there's that

[0] https://www.anthropic.com/news/detecting-and-preventing-dist...

zaptheimpaler 3 days ago | parent | next [-]

All of the models stole the entirety of written knowledge on the internet to train. They are being sued for the few cases where we have some proof of what they did because of some whistleblowers, all the rest will just go unpunished. They breached Github TOS, robot.txt's, copyright, patents every form of IP protection under the sun from a billion sources. It's just ridiculous for the thieves to cry about someone else stealing from them.

smeeth 3 days ago | parent | next [-]

Do you care about the law or not? I think theft is bad everywhere, not just when Anthropic does it.

Dylan16807 3 days ago | parent | next [-]

For me, when it specifically comes to copying, I don't think it's bad to copy a copier. (And by that I mean Anthropic has no valid complaints against Moonshot. Any valid complaints from anyone in the original corpus are valid against both of them now.)

In this way, it is different from literal theft. Stealing money/objects from a thief and keeping them is not justified.

smeeth 3 days ago | parent [-]

It's a little different in this case, since 1) not all the data Ant used was stolen and 2) they did contribute significantly to the value of the stolen good.

An analogy might be a baker stole 20% of the flour used to bake their special bread, which was then stolen. Both thefts are obviously wrong and bad.

optionalsquid 3 days ago | parent | next [-]

The exact same argument could be made in Moonshot's favor.

Being a bit tongue in cheek, one could also argue that by releasing their models, Moonshot is contributing much more value than Anthrophic. Did Prometheus not create an immense amount of value, when he took fire from the hands of the Gods and gave it to humans?

Dylan16807 3 days ago | parent | prev | next [-]

I think any analogy with physical theft is too different from data to apply to this comparatively subtle case. Especially when we get into the details of just using the output of the model to train on.

asadotzler 3 days ago | parent | prev [-]

The baker stole 100% of the flour to make the bread. He also stole the water and the salt and the yeast and the heat for his oven. What he didn't steal was the time he put into crafting a recipe for bread and the time he sat around waiting for the oven to bake it. Now, is that loaf stolen property? Hard to say. But the baker is undoubtedly a thief. He should be tried and forced to pay restitution out of his ill-gotten profits for sure. If we can't do that, the next step is pitchforks and guillotines.

chasil 3 days ago | parent | prev | next [-]

I don't really have an oar in this water, but...

"Judge approves a $1.5B Anthropic settlement over pirated books used to train the Claude chatbot"

https://abcnews.com/Technology/wireStory/judge-approves-15b-...

mcphage 3 days ago | parent | prev | next [-]

> I think theft is bad everywhere, not just when Anthropic does it.

It seems like you think theft is bad everywhere except when Anthropic does it.

archagon 3 days ago | parent | prev | next [-]

Not OP, but copyright law is an absolute joke. No, I don’t care one whit that someone’s TOS was violated. In fact, I find it hilarious. And it’s not “theft.”

smeeth 3 days ago | parent [-]

"I think we shouldn't have IP protection at all" is a totally valid position to hold, but that's not the law is. OP said it didn't violate the law, and it does.

archagon 3 days ago | parent [-]

Some laws are very obviously unworthy of consideration, with broad consensus from the public. See what happened when Napster came out. Literally no one cares about some red-faced RIAA suit flicking spittle over some shared Metallica albums.

Same thing here. This whole situation is just comical.

runako 3 days ago | parent | prev | next [-]

It's a rational position to care about the law, but insist on a queue when related parties are involved.

In this case: resolve the theft claims against the US frontier labs, and only then let them make claims against third parties. It would be totally unreasonable for (say) OpenAI to extract a settlement from Moonshot and use that to pay its own claims. Ordering matters.

jbxntuehineoh 3 days ago | parent | prev | next [-]

no, I don't care about thieves getting stolen from. why would I?

zaptheimpaler 3 days ago | parent | prev | next [-]

If the law was applied uniformly, I would support its continued uniform application. In the last 10 years, I don't see it being applied fairly at all, I see an oligarchy, a criminal and corrupt government and rich and powerful entities getting away with anything. The most minimally competent legal system would ask the AI companies, show us the list of all the data you've used to train and lets hash out the copyright - instead we have to pray someone leaks one tiny piece of what they trained on and then sue for that. Open-weights models are the closest thing we have to justice in the world where the legal system no longer provides justice, because at least the model trained on all of our data is given back to all of us.

vharuck 3 days ago | parent | prev | next [-]

I find that I care more when copyright violations cause actual harm to the copyright owner. Let's say there's an American kid who can't speak Japanese but wants to keep up with a weekly manga. He downloads a bootleg translation and shares it among his friend group. That is a copyright violation, but meh. If he hadn't gone the illegal route, he'd more likely just not read it at all. There's very little chance he'd pay for a subscription and learn Japanese.

Now, if that kid were to print the bootleg translation and sell it to schoolmates, that's worth a slap on the wrist. The kids willing to pay would likely have paid for official copies.

When these LLM labs download our works, feed them into their models, and sell the output to people that used to pay for our work, that's worth a very hard slap. I honestly have less of a problem with the open models.

FpUser 3 days ago | parent | prev [-]

[flagged]

AnimalMuppet 3 days ago | parent [-]

Do you mind having the discussion we're having?

FpUser 3 days ago | parent [-]

I do not block posts and I never downvote.

SubiculumCode 3 days ago | parent | prev [-]

They breached some TOS, but your first sentence is pure, over the top flim flam

himata4113 3 days ago | parent | prev | next [-]

I do agree that two wrongs don't make a right, the terms of service generally gives cooperation the power to sever the contract, but it does not make things illegal in the literal sense. The illegality usually comes from widescale fraud which includes accessing services you are banned from accessing.

When I said "Does this matter?" I specially meant that distillation in itself, the data you get from distillation is first and foremost not owned by anthropic nor is it copyrightable. If a user willingly gives up their anthropic reasoning data/traces that is 100% legal no matter what the "terms of service" say as it's not enforceable and would fall apart in court.

And what I explicitely pointed out that focusing so much on distillation is an attack on open research and claiming that the majority of advancements are thanks to US labs which is simply not true (at least not anymore this was somewhat true during deepseek R1 era), but that in itself was inspired by open research.

> How do you know that? Do you work for Anthropic? Also, this has nothing to do with architecture, we are talking about data.

Because anthropic would be the first ones to make that information public and the architecture is unique to kimi... They made it, they wrote papers on it, it's their research.

P.S. none of the quoted laws apply here since no trade information is stolen, the one about circumventing distillation protection might hold up in court although unlikely.

smeeth 3 days ago | parent [-]

> The illegality usually comes from widescale fraud which includes accessing services you are banned from accessing.

Agree, and this is exactly what Anthropic is alleging.

> data you get from distillation is first and foremost not owned by anthropic nor is it copyrightable. If a user willingly gives up their anthropic reasoning data/traces that is 100% legal no matter what the "terms of service" say as it's not enforceable and would fall apart in court.

It's important to note this is NOT what happened. Anthropic was able to trace data directly back to employees at the company: "We attributed the campaign through request metadata, which matched the public profiles of senior Moonshot staff."

> none of the quoted laws apply here since no trade information is stolen

There is a lot of work showing Kimi models produce similar outputs to Anthropic models, which constitutes trade information. This is not dissimilar to past and ongoing IP suits against Anthropic and OpenAI by showing the models would recreate images of Mickey Mouse/NYT articles etc.

For the record, I'm a researcher myself and I'm well aware how competent the researchers are at the open-source labs/how much they've contributed. But that's not at issue here, my disagreement with you is specific to your arguments about legality; you're conflating what you think should be legal with what actually is legal.

himata4113 3 days ago | parent | next [-]

This is mostly just to reiterate myself as the original question was "Does this matter?"

Everything else is simply justifying why it shouldn't, the specifics don't really matter as there is no legal framework to stop china from continuing to distill models and anthropic has proven they cannot use software solutions to stop it either as distillation is still a problem. But I do still believe it wouldn't hold up in court either way as stopping companies from generating training data which was trained on the entire human knowledge corpus is just stealing from thieves and making it 'open' once again so the argument only gets weaker.

edit: to add, the mickey mouse / nyc was because anthropic trained on LICENSED works, not apple to oranges. The original work it was reciting was licensed and not licensed BY anthropic.

asadotzler 3 days ago | parent | prev [-]

What precedents can you cite and specific examples of their applicability. That is, what would Anthropic's lawyers take to court? You can't say because there's nothing there that couldn't be ripped apart by the least legally capable community known to man, HN. That's why no lab has succeeded in a suit anything like what you're claiming could happen. The only reason Anthropic or any other lab would pursue this is political or commercial. They're either looking for help from officials or they're trying to establish a particular market position.

AlanYx 3 days ago | parent | prev | next [-]

>Civil breach of contract. Anthropic's TOS explicitly say you can't do what Kimi is alleged to have done.

This is true, but Kimi also has a variety of defenses. Kimi can't raise unclean hands if Anthropic systematically violated others' terms of use, but it can raise copyright misuse (which is similar in some respects to unclean hands) as well as lack of standing to enforce restrictions in the contract due to the third party beneficiary principle (i.e., Kimi would argue that Anthropic cannot sue Kimi for derived IP that rightfully belongs to third parties whose terms of use were violated by Anthropic, and the proper party to sue Kimi, if any, would be those third parties). That latter argument usually fails in small-scale cases (ProCD) but has been successful in larger ones where the alternative would be anticompetitive.

FpUser 3 days ago | parent | prev | next [-]

>"A lack of prosecution does not make something legal"

Plainly who gives a flying fuck. The US can claim whatever rules they want and so can China or any other country. On international level all those rules are artificial constructs unless they can be enforced. China can just say for example that they do not recognize copyrights /patents / whatever so it is "legal" for them.

smeeth 3 days ago | parent [-]

This is illegal in China too, there's just an enforcement asymmetry. I understand what you're saying is de facto true, I'm just taking issue with people saying either

1) its not illegal (it is)

2) it shouldn't be illegal because Anthropic stole training data (thats not how the law works)

FpUser 3 days ago | parent [-]

>"1) its not illegal (it is)"

I am a practical man. From what I see laws are mostly for common folks and often do not even serve real justice. The higher one goes and the amount of money / power involved the more the laws bend and on international level the only law that matters is the size of one's club and willingness to use it. And when the country with supposedly biggest one starts crying I find it laughable.

skippyfish 3 days ago | parent | prev | next [-]

> Civil breach of contract. Anthropic's TOS explicitly say you can't do what Kimi is alleged to have done.

Ah yes, I remember when Anthropic crawlers abided by the TOS of the websites they slurped up.

All your other points are downstream from this, which makes them pretty tenuous. Labs don't think that ToS or other explicit wishes of content providers apply to them, but they expect everyone else to abide by theirs.

smeeth 3 days ago | parent [-]

To be clear, I think theft is also bad when Anthropic does it.

US and CA law really don't care that Anthropic violated IP law elsewhere.

FireBeyond 3 days ago | parent [-]

Well, if you're solving the -root- problem, then Moonshot would have had nothing to "steal" if Anthropic didn't "steal" it first.

well_ackshually 3 days ago | parent | prev [-]

I hope Anthropic pays you a lot to defend them this hard <3

linkregister 3 days ago | parent | prev [-]

It matters because the closed-source frontier labs spend lots of money on human data (RLHF / RLAIF with human oversight). Moonshot is accused of circumventing these costs. Frontier labs add research costs into their inference pricing. If the market doesn't permit them to sustain sufficient pricing to have a positive cash flow, then their business prospects become weaker and they risk insolvency. Furthermore, other leveraged companies are at risk.

The reason why the United States government is weighing in is because it's in the national interest of the US to have supremacy in "AI".

Legality or lack thereof is one of many data points about whether a thing is noteworthy.

Moonshot performing distillation is rational from their point of view. Reducing costs is in the interest of businesses. It's also rational for frontier labs and the US government to add obstacles to this process.

As consumers this is probably a positive development.

spaceman_2020 3 days ago | parent | next [-]

My parents put in countless hours and tens of thousands of dollars into raising me to the point where I could write an answer on StackOverflow

And OpenAI scraped and distilled that answer and gave me nothing

linkregister 3 days ago | parent | next [-]

What does your story have to do with Moonshot AI? Do you think they didn't also use the same corpus? Bizarre

voidnullvalue 3 days ago | parent | prev [-]

And now people such as myself have access to open weight models with that information. I wasn't lucky enough to have parents put me through school, and LLMs have absolutely helped me further educate myself and play "catch up" on opportunities others have been given. So, the net effect has been (and is continuing to be) a democratization of information.

overgard 3 days ago | parent | next [-]

You could have gained that stuff prior to LLMs. The leg up you're describing is free information on the internet, not AI. AI just makes it a little easier to find, while also crushing the original sources in the process. (Even if it had a broken culture, is stack overflow even going to exist in a year? Where are they going to train on going forward?)

spaceman_2020 3 days ago | parent | prev | next [-]

Democratization of information, but Sam Altman gets a $100B net worth and I'm still broke :)

I would prefer some sort of democratiziation of the money made from the democratization of information as well

voidnullvalue 3 days ago | parent [-]

Agreed, i wish that it would have done more than change who gets rich off rent-seeking behavior surrounding the knowledge that others created, instead it just consolidated that from many gatekeepers to a few.

I can at least take some measure of pleasure in the fact that it has generally lessened the roadblocks in gathering information. I am still displeased that there are any gatekeepers of humanity's combined knowledge

spaceman_2020 2 days ago | parent [-]

It’s even worse than before

Businesses like these used to public at reasonable valuations. You could ride with them to trillion dollar valuations and grow your own fortune too. Everyone has a story of buying Apple or Google or Amazon stock and making millions

Now they’re going live at trillion dollar valuations and by the time you get in, all the upside has already gone (see Spacex IPO)

Not only did they steal all human data, they also made sure that the upside was only limited to themselves and their cronies

Balooga 3 days ago | parent | prev [-]

Not to be an arse, but didn't you have access to Stack Overflow with all questions/answers prior to LLMs?

voidnullvalue 3 days ago | parent [-]

Yes, but time is finite

noja 3 days ago | parent | prev | next [-]

Isn’t that the same argument they are making for replacing human labour?

Circumventing costs.

SubiculumCode 3 days ago | parent | next [-]

There are many frames that one can place upon this issue. They do not contradict the other. There are moral framings (stole the internet so go eff yourselves, is one), but so is national security, and so is the doomer recursive self improvement risk, and then there is the framing purely on what this implies for future AI training.

I mainly focus on the last.

It will be hard for a frontier lab to justify spending the compute and data curation needed to advance AI further if that expenditure can be assimilated into your competitor's products within months/weeks. So reality will present labs with three choices:

A. Cease spending massive amounts of money and compute improving those models.

B. make those improved models more difficult to distill from, either through some regulatory regime, or some technical solution, which seems unlikely to me.

C. making the best models available only to select partners and government.

In all these potential outcomes, China, which lacks compute that U.S. labs enjoy, will likely stop seeing massive improvements in their AI models. Improvements to be sure, but right now they are enjoying gains from distillation AND their own model innovations, and these potential outcomes would largely stop one of those sources.

linkregister 3 days ago | parent | prev [-]

Do you get mad at your computer for replacing clerical workers? What does this nonsense comment have to do with the issue at hand?

watwut 3 days ago | parent [-]

Those were told "find another job" and in fact they were able to find different jobs.

AI companies are gleefully bragging and "making humans obsolete", "permanent underclass" and 40% unemployment rates they plan to create.

They pushed to replace people years BEFORE their technology even can produce that work.

So, you know, it is not the same. But also in fact, clerks did disliked when occasionally arrogant claimed to replace them while pushing unfinished software that dont quite work yet.

andyfilms1 3 days ago | parent | prev | next [-]

Oh, so mass theft is okay as long as American companies are doing it

ffsm8 3 days ago | parent | next [-]

copyright infringement is not theft, even if right holders often claim it is.

part of the definition of theft is that the original owner is deprived of it, which does not apply to copyright infringement.

You can only argue with damages from the perspective of potential profits, still not theft though.

https://en.wikipedia.org/wiki/Theft

hungryhobbit 3 days ago | parent | next [-]

So having tons of AIs quoting various literary works and reproducing knock-offs of them has a positive effect on those books' sales?

I think you're wrong: there is absolutely damage to the authors and publishers from what the AI companies have done.

ffsm8 3 days ago | parent [-]

? I literally said that, how am I wrong?

> You can only argue with damages from the perspective of potential profits, still not theft though.

Damages are not deprival of ownership. They're conceptually related but orthogonal

Also there was no moral judgement from my end, I just pointed out that an incorrect word is being applied. It's just not theft - by definition. But language is a fluid concept and definitions change over time. As people keep misusing it, it will eventually lose its original meaning. Which may have already happened for you, but this change hasn't been settled yet as can be seen from looking at the official definitions of the term, which as of today still mention the criteria

evanelias 3 days ago | parent | prev | next [-]

If you steal an unpopular product from a store, the damage is also only to "potential profits", so how does that differ? It's entirely possible no one would have purchased the product and it would have eventually been discarded/destroyed.

Or with services, if a barber cuts your hair and then you run away without paying them, do you not consider that theft, even though there's no change in ownership occurring?

BlackFingolfin 3 days ago | parent | prev [-]

This is almost funny to me, because in many jurisdictions, software companies sure invested a lot of effort into painting people copying software as thieves. In Germany, they (the software producer lobby, and later politicians influence by the former) even coined and spread the term "Raubkopie", which you could roughly translate as "robbed copy", i.e., that's one step worse than "theft", as a robbery in Germany legally means " theft accomplished by force or intimidation". So, yeah: like putting a knife to the throat of someone while you copy the software.

So, after literally decades of investing into advertising campaigns, lobbying to politicians to pass harsher and harsher laws against software "thieves and robbers", now that big tech are doing it, suddenly we are supposed to consider it with more nuance?

Ahhh... no thank you sir. I really enjoy them drinking their own kool-aid.

SubiculumCode 3 days ago | parent | prev | next [-]

Moreover, reading a copyrighted book and learning from it is not theft.

bigfishrunning 3 days ago | parent | next [-]

Generating a set of weights is not learning.

jayGlow 3 days ago | parent | next [-]

would you say that airplanes don't fly because they don't flap their wings? it's possible to achieve the same things with different approaches.

bigfishrunning 2 days ago | parent [-]

I would say airplanes fly, but I wouldn't say that submarines swim. Things have a bit more nuance, and the field of "learning" isn't as well understood as the ML proponents claim it is.

SubiculumCode 3 days ago | parent | prev [-]

That is a strong statement. I guess you are telling Machine Learning to go fuck itself.

bigfishrunning 3 days ago | parent [-]

No, Machine Learning is an unfortunate name for a well documented process for creating black-box classifiers. The process is good, the name is not.

SubiculumCode 3 days ago | parent [-]

And what, to your mind, would classify something as learning? I assume that your position is not the hard "only humans/living creatures can learn"

bigfishrunning 2 days ago | parent [-]

Honestly, I'm not sure. But I do know that there is an entire field of cognitive science dedicated to understanding learning, and quite frankly it's in its infancy. Evidence of this is that every elementary school introduces new teaching techniques from time to time, and very rarely do they result in any benefit to the people who are doing the learning (more often they benefit consultants...).

However, the current process of "Machine Learning" (which is a semi-random parameter descent/evolutionary replacement process) is unlikely to be equivalent to the way people learn, because we aren't copying/competing/replacing our brain constantly. People are actually very good at learning, but our brain material replaces itself partially and relatively slowly (when compared to how a neural network is trained).

trollbridge 3 days ago | parent | prev | next [-]

Great! Neither is distillation then.

SubiculumCode 3 days ago | parent [-]

Never said it was. Still, understanding to what extent the ability of Chinese labs to keep up to western models with much less compute needs to be understood.

Espressosaurus 3 days ago | parent | prev [-]

Machines are not humans.

linkregister 3 days ago | parent | prev [-]

Reread my comment and look for a value judgement on my part. The final sentence is probably a good clue as to my opinion.

robotpepi 3 days ago | parent | prev [-]

Chatgpt routinely cites and uses papers I don't have access to because they're behind a paywall. I don't think OpenAI is paying for all that copyright. That's in my opinion way more serious.

fc417fc802 3 days ago | parent | next [-]

Yes the fact that the scientific literature - created largely on the back of the tax payer - isn't open to all free of charge by force of law is a travesty. A cartel should not get to charge for access to the bulk of human knowledge. That is indeed a far more important issue than whether or not Moonshot violated the Anthropic ToS, possibly committing mass fraud in the course of doing so.

I mean honestly if they did that why should I care? I'm happy to see copyright violated in a manner that leads to the creation of new technology. IP law exists strictly for the benefit of society and by all appearances AI is an incredibly powerful tool.

Also while I'm at it libgen is a gift to humanity. Information wants to be free. Spreading and preserving knowledge is generally one of the most wholesome activities anyone can undertake as far as I'm concerned.

linkregister 3 days ago | parent | prev | next [-]

Your statement is orthogonal to my comment. Why reiterate the schadenfreude / fairness comment already stated several dozen times in this thread?

nylonstrung 3 days ago | parent | prev [-]

Who do you think is paying $100K+ for "Enterprise" access to Anna's Archive?