Remix.run Logo
_aavaa_ a day ago

> distillation: why exactly is it bad? After all, what are large language models but the distillation of all of the knowledge on the open Internet, scraped by the frontier labs and distilled into the models that are themselves being distilled? Who is exactly being wronged here? ... The U.S. should pass a law that (1) makes explicit that collecting data for training models is fair use, and (2) bars terms of service that forbid distillation

Sounds great to me; live by the sword, die by the sword.

eli 16 hours ago | parent | next [-]

Seems only fair that if LLMs can use copyrighted data for training then they should be able to use cannot-be-copyrighted output of other LLMs.

But barring the terms of service from forbidding distillation seems like a tough sell. OpenAI shouldn't be allowed to decide what types of customers it wants and doesn't want?

mediaman 16 hours ago | parent | next [-]

This happens all the time. The government can decide legislatively that certain commercial terms are simply unenforceable. Making distillation clauses unenforceable in tort law would be straightforward. They can decide what customers they want to have, but they do not have unfettered rights as to the enforceability of terms governing the relationships between the parties.

eli 14 hours ago | parent [-]

I'm not doubting it's possible to pass such a law, I'm doubting that's it's a practical or worthwhile goal.

The terms of service don't even necessarily matter here. OpenAI could cancel your account for almost any reason, or for no reason at all. They don't particularly need to cite a ToS violation just as a store owner doesn't need to point to a written policy to kick you out of their store.

If the underlying issue is that LLMs should be regulated as a public good, then lets have that discussion. If it's that the major AI companies are becoming too powerful and anti-competitive, let's talk serious anti-trust enforcement. Micro-managing business policies isn't going to work very well.

hnfong 9 hours ago | parent [-]

You two are talking about different things.

You are pointing out that OpenAI can cancel user's accounts for almost any reason, and nobody can really force them to serve customers that they suspect are distilling their models.

That's one thing.

The GP is saying the government can make laws to make terms against distillation unenforceable. Without such laws, if you signed an agreement with OpenAI pinky swearing you won't distill, but turns out you did, you are liable in tort and OpenAI can sue you. (It seems nobody really cares about contract and agreements any more, but still...)

This is the other thing.

And I think you are both right.

eli an hour ago | parent [-]

What do you suppose the damages would be for a ToS violation? The difference between subscription rates and API rates?

OpenAI accused Deepseek of misappropriating trade secrets which could have serious penalties but seems like an awfully hard case to make.

Seems like we’d all be better off with a law that governs data sharing among AI companies, if that’s the policy goal.

Terr_ 15 hours ago | parent | prev | next [-]

> Seems only fair

"You're trying to kidnap what I've rightfully stolen!" -- Vizzini

paxys 15 hours ago | parent | prev | next [-]

It's pretty common to have such laws. OpenAI can put whatever they want in their ToS, but they cannot go back and sue someone for violating those terms if the government has ruled that clause to be unenforceable.

matheusmoreira 13 hours ago | parent | prev [-]

> OpenAI shouldn't be allowed to decide what types of customers it wants and doesn't want?

Correct. It shouldn't be allowed to do that.

jay_kyburz 13 hours ago | parent [-]

Err.. I would like preserve my own right to decide who I'll do business with.

SomeHacker44 12 hours ago | parent | next [-]

You already do not have that unfettered right in the USA.

jay_kyburz 6 hours ago | parent [-]

Yes, sorry I didn't mean to imply that I didn't want to do business with minorities or protected classes.

inigyou 3 hours ago | parent [-]

Sam and Dario keep saying intelligence will be like water or electricity. If that's the case then it should be illegal to deny to anyone. In most parts of the world the power company can't shut off your power for an unpaid bill - they have to get a court order to allow it, which gives you a chance to defend yourself or make a payment plan.

CamperBob2 12 hours ago | parent | prev | next [-]

Then write your own training corpus.

Der_Einzige 4 hours ago | parent | prev [-]

Unironically would love to kill that right. Unironically that stupid cake maker in Colorado should have just made the damn cake.

Wage spiritual warfare against the petit-bourgeoise. They all deserve it anyway, as they are the traditional harbringers of actual fascism.

eli an hour ago | parent [-]

The law was probably fine - that case was just wrongly decided based on the outcome the majority of justices wanted.

grim_io 15 hours ago | parent | prev | next [-]

Forbidding distillation is like forbidding using a compiler to make another(perhaps better, more efficient) compiler.

chuckadams 15 hours ago | parent | next [-]

Lots of software licenses have “non-compete” clauses that forbid you from using it to develop a competing product. Wouldn’t surprise me if there was a compiler or two out there with that restriction, most likely niche languages.

inigyou 3 hours ago | parent | next [-]

Oracle database has a clause forbidding anyone to publish benchmarks of it.

matheusmoreira 13 hours ago | parent | prev | next [-]

Those clauses should be illegal.

not2b 13 hours ago | parent | prev | next [-]

It's been common in electronic design automation tools to have license terms like that (forbidding use to create a competing product). However, competing companies have often found workarounds, either by finding loopholes or just breaking rules and hoping not to get caught.

wolpoli 8 hours ago | parent | prev | next [-]

If a person were to receive data from someone subjected to such restriction, is the receiver bounded by the same restriction?

kiicia 4 hours ago | parent | prev | next [-]

balmer told us that gpl is cancer, but true cancer is us model of licensing

scotty79 14 hours ago | parent | prev [-]

How the hell is non-compete legal in market economy? Competition is one of its core strengths. Why would anyone let anyone opt out of this, even a little bit?

thesmtsolver2 14 hours ago | parent [-]

No country in the world is full free market economy. It is always a spectrum.

We are discussing Chinese models. Now look at how much foreign competition the Chinese government prevents in their domestic market in other industries.

scotty79 14 hours ago | parent [-]

Chinese companies compete ruthlessly between themselves though. That's how they get this good. Full competition with preventing exploitation by foreign countries seems to be working great for them. American and European protectionism of local rent-seekers can't really compete with that.

kiicia 4 hours ago | parent | prev [-]

let's call it for what it really is, only companies "entitled to legally stolen data, don't steal from us now" are crying about distilation

cayley_graph 15 hours ago | parent | prev | next [-]

Yup, fair's fair. Anything else stinks of 'rules for thee but not for me' (a maxim the frontier labs seem worryingly happy to apply, on several counts).

ronsor 16 hours ago | parent | prev | next [-]

I am immediately sold on this.

Sorry, OpenAI & Anthropic.

matheusmoreira 13 hours ago | parent | prev | next [-]

> distillation: why exactly is it bad?

Felony contempt of business model.

magarnicle 13 hours ago | parent | prev | next [-]

Why would reading copyrighted material ever be an issue anyway? Wouldn't copyright law only apply to what you create and publish using the model? Training on every comic book should already be perfectly legal, as long as you accessed them legally, right? But publishing your own Batman comic using that training is copyright infringement.

What I'm saying is, doesn't the law already cover 1?

_aavaa_ 12 hours ago | parent [-]

Fair use requires more than you accessing the material legally.

In the US one of the factors is “ the effect of the use upon the potential market for or value of the copyrighted work”.

If anthropic Hoovers up the world’s books and trains on them, and then spits them out verbatim on command, then it will clearly impact the value of the work; nobody will buy the original, they’ll just ask Claude.

Others also argue that even if it’s not reproducing it exactly that the training runs afoul of that factor, specifically the “market for” portion. A rights holder can no longer license their book for training of LLMs if Anthropic goes ahead and just trains on it anyway.

magarnicle 11 hours ago | parent [-]

> If anthropic Hoovers up the world’s books and trains on them, and then spits them out verbatim on command, then it will clearly impact the value of the work; nobody will buy the original, they’ll just ask Claude.

Ah, right. So if we want models to be capable we need them to be trained on as much as possible, yet we also want to stop what you described. So what can be done?

_aavaa_ 3 hours ago | parent [-]

I mean the choice is: 1) we pass laws that explicitly say training models like this is legal (the original quote, 2) say it’s illegal and requires licenses for the data and ability to opt out, 3) we ignore it and continue because the companies are too big to jail.

qurren 15 hours ago | parent | prev | next [-]

Government cannot exactly "bar" terms of service. ToS isn't law. The most they can do is say they're unwilling to enforce them.

ToS is just conditions that you agree to in order to use a private service that is provided at-will. I can have a private coffee shop where the terms of service are that you must wear red to enter, and if you're not wearing red, you are not welcome on my property.

So it would be upto OpenAI and Anthropic to enforce them on their own terms (by banning accounts and IPs).

ascorbic 15 hours ago | parent | next [-]

The government absolutely can pass laws that ban particular contract previsions. They do that all the time. In your analogy for example while they can require you to wear red, they can't require you to be white.

onesociety2022 15 hours ago | parent | prev | next [-]

Governments can do anything they want by passing a new legislation. In your example, they could easily pass a law that states that any ToS cannot reject service to a customer based on the color of their attire. In the USA, it's obviously already illegal for a business to reject service to a customer based on some protected classes like race.

ButlerianJihad 14 hours ago | parent [-]

The joke is on you! I’m not wearing any attire! Hahaha!

nl 15 hours ago | parent | prev [-]

That's just not true. You can absolutely have terms of service that are illegal, and the government can enforce them.

llm_nerd 16 hours ago | parent | prev | next [-]

The distillation explanation is classic American exceptionalism: No one could possibly do anything unless they were copying American leaders (where "American" means a bunch of Chinese, Canadian, Europeans and Indians working in the US).

It's also a bit of securities defensiveness. Pretending that you really do have a super moat, people just keep swimming in it so you just need to add more alligators.

It's farcical. Anyone who has worked on large models knows that the premise that an almost-Fable model was trained with distillation is beyond ridiculous. It's theoretically possible if they spent tens of billions of dollars on API calls, but it isn't the magic that somehow these people keep convincing people it is.

Previously Anthropic has reported on some Chinese firms doing chicken-shit level of API calls, that at most would be doing some Q and A or final fine tuning. The notion that they're training these models via it is fantastically ignorant nonsense that only very ill-informed and gullible people fall for.

villish 9 hours ago | parent | next [-]

> Previously Anthropic has reported on some Chinese firms doing chicken-shit level of API calls, that at most would be doing some Q and A or final fine tuning

"Anthropic said the campaign was conducted between April 22 and June 5, 2026, and generated more than 28.8 million exchanges with Claude through almost 25,000 fraudulent accounts."

I don't know why you're trying to downplay it.

European models are so far behind because they don't resort to these tactics on a massive scale. Basically every other country is entirely dependent on 2 countries for frontier AI.

hnfong 9 hours ago | parent | next [-]

> European models are so far behind because they don't resort to these tactics on a massive scale. Basically every other country is entirely dependent on 2 countries for frontier AI.

You may or may not be factually correct in your other points, but you're really proving the GP's point here regarding American exceptionalism.

villish 8 hours ago | parent [-]

Are there other countries releasing frontier level models? Mistral is the only relevant player I can think of that comes from Europe, did I miss one?

llm_nerd 4 hours ago | parent | prev [-]

>I don't know why you're trying to downplay it.

Ignoring that I have literally zero trust in anything Anthropic has to say on this -- they have been doing the hysterical routine and trying to get every bit of government granted monopoly they can[1] -- those numbers still simply aren't that impressive.

>European models are so far behind because...

What a non-sequitur. Europe, like much of the West, foolishly delegated tech, media, payment systems, etc, to the United States. European efforts on this are poorly funded, poorly capitalized, and marginal efforts.

China is very much not Europe. China is looking to leave the US to the dustbin of history, and their efforts are a little more concerted.

[1] Surely Americans are aware that Anthropic and OpenAI are both very close to getting the US government to ban and fully criminalize the open Chinese models, right?

thesmtsolver2 14 hours ago | parent | prev | next [-]

China goes even further lol

https://m.economictimes.com/industry/renewables/china-wto-co...

ultrablack 15 hours ago | parent | prev [-]

Which Chinese model was it that identified itself as Claude 15% of the time?

inigyou 3 hours ago | parent | next [-]

Claude Opus identifies itself as Qwen if you ask the question in Chinese. So who's really distilling who?

llm_nerd 15 hours ago | parent | prev [-]

Models don't have some self identity, beyond what is explicitly handed to them via a system prompt. There have been many, many cases of models identifying as different models by different makers as a basic identity hallucination. They train on enormous volumes of data including lots of people talking about certain makers and models (ChatGPT was actually a super common one given that it became the kleenex of the LLM world). Hence why vendors have to specifically tell it to override that, and if they don't you get lots of funny cases of identity confusion.

This isn't the big gotcha some people seem to think it is, and the whole news cycle about that was mostly by people who have no idea what they're talking about. It's actually a meaningless data point. But it's precisely the sorts of people who think that a few thousand free accounts surreptitiously snuck off with Fable.

noncoml 16 hours ago | parent | prev | next [-]

Don’t know much about how distillation works so please enlighten me here.

> what are large language models but the distillation of all of the knowledge on the open Internet, scraped by the frontier labs and distilled into the models

If it’s as easy as that why do they choose to distill another model and not distill the knowledge on the open Internet from scratch?

numpad0 16 hours ago | parent | next [-]

known-good prompt-response pairs are more useful than random semi-coherent texts presumably

paxys 15 hours ago | parent | prev | next [-]

You need to do both.

A model trained on all knowledge from the internet (and other sources) is large but ultimately not very useful by itself, because it is going to spit out all kinds of garbage. You have to apply multiple further stages of training and refinement to the base model before putting it in front of users. So as an example you can train a model by yourself and then have GPT or Claude continuously check its outputs and correct it when it is wrong, ending up with a far more powerful model.

root_axis 15 hours ago | parent | prev [-]

Because the model can output data in a manner optimized for training a new model, including outputs that were post-trained like RLHF and RLVR.

petilon 16 hours ago | parent | prev | next [-]

[flagged]

altruios 16 hours ago | parent | next [-]

This is a silly perspective, inaccurate, and out of bounds framing.

Public libraries, in this instance, is curated data from all the internet, obtained through not legal means (I don't have a problem with this other than lack of attribution, being copy-left). Just to be clear.

But in answer to your incredibly leading and inaccurate framing... they are required (by their job title) to teach to those who who show up in the classroom, it's not their place to discriminate against anyone/thing (even those like itself (other robots)) that also show up in the classroom.

But you can't teach at a university using only knowledge learned from the library. you need a degree. You are free to teach at the park, where anyone can hear you. public in -> public out.

petilon 16 hours ago | parent [-]

If a professor learns from multiple books, generalizes from them and then shares his knowledge he is providing a valuable service. Versus someone who makes a recording of the professor's lectures and resells them to undercut the professor--that guy is not providing a valuable service.

tux3 16 hours ago | parent | next [-]

What kind of fresh hell does the sentence "undercut the professor" come from?

Teaching isn't a race to the bottom. You don't undercut teaching by giving more lessons, just like you don't slight the hospital by performing CPR.

petilon 15 hours ago | parent [-]

We are not really talking about teaching here.

tux3 7 hours ago | parent | next [-]

We sure wouldn't be, if we picked better metaphors.

_aavaa_ 15 hours ago | parent | prev [-]

No, we’re talking about an intimate set of tensors, not a human being.

A tree falling and killing someone isn’t tried for manslaughter.

So I don’t care about a hypothetical teacher.

lelanthran 9 hours ago | parent | prev | next [-]

> If a professor learns from multiple books, generalizes from them and then shares his knowledge he is providing a valuable service. Versus someone who makes a recording of the professor's lectures and resells them to undercut the professor--that guy is not providing a valuable service.

I'm confused now; isn't the LLM that trains on that professor's lectures, videos and textbooks undercutting him?

Where were you going with this?

idle_zealot 16 hours ago | parent | prev | next [-]

What about a student attending lectures of other professors and generalizing what he learns from them, then going on to become a professor?

petilon 16 hours ago | parent [-]

If the student is really good at generalizing we wouldn't even be having this debate because he would've just generalized from the same source materials the professor used.

altruios 16 hours ago | parent | prev [-]

What are you even trying to say: "undercut the professor"...

The further you try to constrain this topic into this illformed analogy the weirder it becomes. If we start off with a better analogy...

Crisco 16 hours ago | parent | prev | next [-]

No, but the students that learn and distill what the professor teaches are not obligated to use that information only how the professor wants them to.

petilon 16 hours ago | parent [-]

Can the professor refuse to teach some students, or must he teach all comers?

xboxnolifes 16 hours ago | parent | next [-]

One can teach whoever they want to or don't want to. If one joins a university, that changes things. They are now part of an organization larger than themself.

lostmsu 16 hours ago | parent | prev [-]

glhf under the circumstances

smarf 16 hours ago | parent | prev | next [-]

'why is reselling stolen stuff bad'

if the professor took all human knowledge, much of which was explicitly not free, and used it to make a for-profit knowledge machine that extrudes unreliable summaries of that knowledge, then yes, being obligated to teach for free would be a fitting punishment.

mywittyname 16 hours ago | parent | prev | next [-]

More like, is a professor who learned from books prohibited from writing his own books on the subject?

petilon 15 hours ago | parent [-]

He is prohibited from regurgitating source material, of course! But if he generalized from the books he read and really learned the subject--and even made new connections between ideas--then he is free to write his own book.

mywittyname 15 hours ago | parent [-]

He is not prohibited from "regurgitating source material" in many cases. Facts are free. It doesn't matter who first measured Young's Modulus of aluminum, anyone may state that fact as originally presented.

The professor is free to lift all the facts and formula they want. They just need to rephrase explanations. Which is pretty much what an LLM is going to do.

16 hours ago | parent | prev | next [-]
[deleted]
janalsncm 15 hours ago | parent | prev | next [-]

1) No one is asking Anthropic to give tokens for free, but at market rates.

2) Any professor who tried to ban students from posting lecture notes online would be immediately mocked.

_aavaa_ 15 hours ago | parent | prev [-]

Who cares, a LLM isn’t a person.

bluegatty 15 hours ago | parent | prev [-]

Making an LLM from raw data is value-add.

Distillation is just value extract.

It's soft, and I'm not sure what the answer should be ... but I think that there is a difference.

I think we start by recognizing that ... and then try to figure it out from there.

'The Internet' may be a public good, maybe we make them pay a tax for that, but that's different than distillation.

nemomarx 14 hours ago | parent | next [-]

What makes the Internet raw data in a different way? wasn't it mostly worked on by people first?

bluegatty 14 hours ago | parent [-]

There is value add in AI irrespective of how the data got to what it is.

Literally the biggest thing of our generation - AI - is the living embodiment of that 'value add' writ large.

'What is the difference' - is the AI you use all day, in comparison to 'all the world's data' you can use for stuff and do 'whatever' with it, but are not likely to come up with something hugely useful otherwise. Maybe, not likely, if you did, it would be 'value add'.

nemomarx 14 hours ago | parent [-]

Okay, so if the chinese models are used everyday, do they become a value add? Like what's the line you're drawing here. Amount of value it creates?

bluegatty 12 hours ago | parent [-]

Designing and creating an LLM from nothing is a monumental feat of Engineering and 'value add'.

Copying something is not.

Programming Microsoft Word is value add, copying the code is not.

Copying design ... there are some question marks there.

It's extremely easy to understand at it's core.

What makes it hard, is that faux intellectuals like to deconstruct ideas at the margins, and have those critiques stand in for reason.

"At sunrise the sun is only 'half there' ... there fore there is no 'day and night' just a blur! Day and night are the same thing!"

The training data used is part of all of this is a separate but related question.

inigyou 3 hours ago | parent [-]

they all just copied the Transformers paper anyway

scotty79 15 hours ago | parent | prev [-]

> Making an LLM from raw data is value-add. > Distillation is just value extract.

There is a value-add in selecting the valuable parts out of the garbage. And let's face it. Largest models contain a lot of garbage.

bluegatty 14 hours ago | parent [-]

I think that's kind of fair, but it still fits within the context of 'some things are value add' and 'more or less than others'.

We ought to identify that and integrate that into our thinking.