Remix.run Logo
Bratmon 3 days ago

I like this comment because its argument only makes sense if you assume that the entire world's output of books and art did not require a huge amount of resources and expertise to make, nor did it add any value.

It's the most CS-major take ever!

mapontosevenths 3 days ago | parent | next [-]

If turning other peoples copyrighted work into a model is transformative enough to be protected then so is distilling that model into a different, better, model.

foo12bar 3 days ago | parent [-]

The models were built using copyrighted works, so why can't models be built using other models?

usef- 3 days ago | parent | next [-]

They do seem to be paying for it (as per the 1.5Bil lawsuit yesterday and them now purchasing books and licensing from media companies).

Whether we think they're paying enough is another question, but "I'm paying for content so can protect it" doesn't seem inconsistent.

We may decide that giving models away for free means they don't have to license content (judging by HN comments), but currently that doesn't seem to be the case as Meta is facing lawsuits for its open models.

(Obligatory stratechery piece: https://stratechery.com/2026/whos-afraid-of-chinese-models/ )

trhway 3 days ago | parent [-]

The judge found their use is fair use. They are paying not for their use of the content, they are paying for using illegal copies of the content.

The same principle can be applied to distillation - it is a fair use. You just shouldn't use illegal ways to access the models being distilled.

To the commenter below: if it is illegal - has the police/FBI report been made? Otherwise it is just a civil court matter.

usef- 3 days ago | parent [-]

Fair, but isn't "illegal" access what they're talking about in OP?

It does seem to be becoming the norm for AI companies to licence premium content in America, judging by the deals they're making. It doesn't seem to be done by the international distillers. It's a cost that American open models will seem to have to pay but not international.

trhway 3 days ago | parent | next [-]

>It does seem to be becoming the norm for AI companies to licence premium content in America, judging by the deals they're making. It doesn't seem to be done by the international distillers.

International distillers doesn't use that premium content, so they don't pay for it. They do pay for their access to the models they are distilling. Thus providing the revenue stream to those models. Thus those models make profit off the content they used for training. The content they mostly have't paid for.

>It's a cost that American open models will seem to have to pay but not international.

It goes both ways - American companies and their business are protected by American laws and have access to the market protected by those laws, etc.

usef- 3 days ago | parent [-]

> International distillers doesn't use that premium content, so they don't pay for it.

This doesn't seem to be true. They are training on their own scraped data overwhelmingly (we can extract copyright data from, eg, deepseek). They couldn't get nearly enough tokens through the American APIs to train a model on alone.

> American companies and their business are protected by American laws and have access to the market protected by those laws

Absolutely. Currently international providers are selling inference on the American market though, I don't know how that will sit legally the way things are currently going.

Bratmon 3 days ago | parent | prev [-]

> It does seem to be becoming the norm for AI companies to licence premium content in America, judging by the deals they're making.

This is a very surprising claim to me (and I imagine many small website owners who keep getting scraped by Anthropic and OpenAI).

Do you have a source?

usef- 3 days ago | parent [-]

There's been many news stories of it over the past year(s) as they signed each one. Here's the first result I could see with a rundown of many of them (am on mobile).

https://digiday.com/media/a-timeline-of-the-major-deals-betw...

Bratmon 3 days ago | parent [-]

Those are licenses for API access to data too new to be in the training data (for use by agents), not for the training itself.

I don't really understand why you think they're relevant, given that this conversation is about the training itself.

usef- 3 days ago | parent [-]

It isn't just that, it includes publishing houses, Wiley etc, and non-live media.

Even the news orgs say explicitly in the press releases that it's about training on their archive

eg. http://ap.org/media-center/press-releases/2023/ap-open-ai-ag...

---

edit, examples:

Wiley https://newsroom.wiley.com/press-releases/press-release-deta...

Shutterstock https://investor.shutterstock.com/news-releases/news-release...

Axel Springer https://openai.com/index/axel-springer-partnership

Stack Overflow: https://stackoverflow.co/partnerships

Disney (for characters in video. Video is especially where licensing is a big difference internationally right now) https://openai.com/index/disney-sora-agreement

etc.

The news corp one had a leaked price ($250mill), so they don't seem to be insignificant. These would have to be included in API prices I presume.

breppp 3 days ago | parent | prev [-]

Because model output is probably far closer to software or a licensed work which possibly has greater protections than it is to copyright. There is far less possibility of fair use, it might be protected by patents, license or reverse engineering laws.

In any case the laws are being written now, but I doubt these will have worse protection than software does, which has far better protections than copyright

giaour 3 days ago | parent | next [-]

> I doubt these will have worse protection than software does, which has far better protections than copyright

Software is protected by copyright. Some software may also be protected by patents, but last time I checked, AI generated output of any kind was not patentable.

trhway 3 days ago | parent | next [-]

Distillation isn't a copy. Distillation is more akin to "clean room" implementation.

Also note that the OpenAI/Anthropic argument is that the model training is sufficiently transformative to satisfy the fair use of the original content for training.

By that same argument, when distilling the distillers aren't using the original content the OpenAI/Anthropic models were trained on - the distillers are interacting only with the "sufficiently transformed" content of the OpenAI/Anthropic models and are normally paying for that.

There is also that old phonebook rule that facts can't be copyrighted. So, if i asked the model about bunch of phone numbers, i can publish the resulting list, can train my model on it, etc. Such approach doesn't allow to reproduce copyrighted works of course - and as we know the AI output isn't copyrightable, so it looks like basically any output i get i can use whatever way i like.

breppp 3 days ago | parent | prev [-]

Software is protected by the DMCA, patents, licenses, EULAs, all of those aren't there for books. I doubt new laws won't be written for model outputs.

Also, if model output distillation is shown as some form of reverse engineering I assume the DMCA can apply

bigiain 3 days ago | parent | next [-]

The C in DMCA stand for Copyright. All (I think?) software licenses are underpinned and made legally enforceable by copyrights. EULAs are underpinned by licenses which are founded on copyright. Patents are the only one of those protections that are not based on copyright, and there are lots of very good arguments against at least most software patents (all software patents of the form "Do {well known and obvious thing} with a computer" should, in my opinion, be immediately revoked and potentially have every company who's enforced payments from such patents investigated for fraud).

giaour 3 days ago | parent | prev | next [-]

You may recall that the DMCA was originally written to protect music and movies. It does in fact apply to creative works. If you have ever purchased an MP3, eBook, or streaming movie, you will also be aware that you purchased a license to the underlying IP. This is also true of physical media, but the license agreement you have to accept when obtaining a digital work makes this explicit.

I agree that you can't patent a book, but I would point out that you can patent an idea, which may only appear in a book or journal article.

vel0city 3 days ago | parent [-]

You do patent ideas, but the actual words written in a book describing that idea would only be protected by copyright at best. FWIW, the exact words describing the idea being patented are technically public domain; that's the whole point. You're free to go look up that patent, print it out, make whatever copies of it you want. Take any of the drawings in patents, put them on t-shirts, and sell them. No problem. Implementing the ideas those words represent is a different story.

For example, a patent describing a chemical process. The actual idea of how to do it is public domain, go look up the patent. Print it out. Do whatever with those words. Its fine. Building a plant to go do that chemical process to make that same output chemical in that same way, that's IP infringement. Its not the words, its the idea.

wasfgwp 2 days ago | parent | prev | next [-]

How is “model output distillation” different to using outputs (which are legally copyrightable) for any other purpose?

queenkjuul 3 days ago | parent | prev [-]

Afaik (and ianal) there's nothing stopping anyone from attaching a EULA to a physical book

vel0city 3 days ago | parent | prev | next [-]

Let's assume model output can be claimed by copyright or some form IP. You can't really patent it, as the output isn't a novel idea or process, much like you don't patent a book or a movie. But for arguments sake, let's agree it is some kind of IP.

Who are you saying owns that IP? The people who trained the model? The people who ran the model? The people who wrote the prompt? The person who paid for all of that to happen?

If the model output is owned by the person prompting it and paying for the tokens, what's the problem here?

If the model output is owned by the trainer of the model, that's a big nasty can of worms.

preg_match 3 days ago | parent | prev | next [-]

Why would this be the case. Why would software output from a model magically have greater protection than the software the model trained on.

wasfgwp 2 days ago | parent | prev [-]

LLM outputs are not copyrightable. At least that’s the current established legal precedent in the US. The only question is whether the user owns the copyright without significantly transforming the output but that’s not really relevant in those specific situation.

I mean otherwise it’s a very slippery slope, effectively it would give Anthropic the ownership of any code generated by its models..

arbitrary_name 3 days ago | parent | prev | next [-]

there is a major god complex here.

MBAs and non technical managers = inept Catbert-type charlatans.

Software engineers, devs, etc = geniuses capable of mastering any domain, innate ability to be right on any topic.

joshuamorton 3 days ago | parent | prev | next [-]

I don't think that's what it's saying at all. It's saying that there's a level of creativity in model creation that isn't present in distillation.

skybrian 3 days ago | parent | next [-]

Maybe, but it's not like their AI is likely to repeat it back verbatim so it's unlikely to be a copyright violation. It seems like at most, they would be breaking Anthropic's terms of service?

Or maybe they're going through an intermediary "transfer station" that's breaking terms of service:

https://www.chinatalk.media/p/how-to-buy-cheap-claude-tokens...

perching_aix 3 days ago | parent [-]

Yes, it's just a ToS violation at present. Those are legally binding though, despite the common adage. What that really translates to here though, anyone's guess.

Anthropic's own copyright infringement could apparently be forgiven for 1.5B USD after all, so maybe there's a price that breaking the distillation clause for is acceptable too. Or some other arrangement.

Bratmon 3 days ago | parent [-]

But surely at least one of the websites Anthropic scraped to make Claude had a ToS forbidding automatic access?

Why is Anthropic's ToS any more binding than that of a rabidly-anti-ai literature blog with 50 readers?

skybrian 3 days ago | parent | next [-]

One reason is that they might not have scraped it themselves, so if there was a ToS, it was someone else who broke it. For example, see:

https://en.wikipedia.org/wiki/The_Pile_(dataset)

Another reason is that if you can download a web page without agreeing to a ToS, I'm not sure that counts as one?

queenkjuul 3 days ago | parent | prev | next [-]

> Why is Anthropic's ToS any more binding than that of a rabidly-anti-ai literature blog with 50 readers?

I mean i know you know the answer: anthropic is a corporation with lawyers on retainer, and that's really all that matters

perching_aix 3 days ago | parent | prev [-]

Do feel free to read the court documents to find out and let us know.

> Why is Anthropic's ToS any more binding than that of a rabidly-anti-ai literature blog with 50 readers?

Although I will say, this whole comparison stuff really doesn't seem to be your thing; might impede your analysis quite a lot: https://news.ycombinator.com/item?id=49013148

Maybe ask Claude?

trhway 3 days ago | parent | prev | next [-]

>a level of creativity in model creation that isn't present in distillation.

the same argument - a level of creativity in the world knowledge creation that ins't present in the model training on that knowledge.

Or in other words - model creation and training is just a distilling of the world knowledge.

joshuamorton 3 days ago | parent [-]

I don't disagree. I'm not sure why that's a relevant reply though.

If you think that the addition of a less creative process (model creation) to a more creative corpus ("art") is problematic, then it follows that you should think the addition of a less creative process (distillation) to a more creative corpus (a model) is also problematic.

trhway 3 days ago | parent [-]

I think both are natural and fine. Otherwise we'd have to outlaw analytical thinking.

remus 3 days ago | parent | prev [-]

Yes, this is what I was getting at.

gozucito 3 days ago | parent [-]

There is an even higher level of creativity in creating books, songs and all sorts of art used in model training though. That's your apparent blindspot.

There is no world in which me vacuuming the entirety of human knowledge to make a genai model is ok but hoovering my model answers is not. The hypocrisy is stunning and risible.

Now if you go and make a model based on purely synthetic data and not a single work made by humans, you would have a valid point.

remus 3 days ago | parent | next [-]

> There is an even higher level of creativity in creating books, songs and all sorts of art used in model training though.

No argument here, I completely agree.

> There is no world in which me vacuuming the entirety of human knowledge to make a genai model is ok but hoovering my model answers is not.

I disagree with this though. Clearly LLMs owe a huge debt to everything that has come before, but surely you'd agree that the models that are produced are something substantial and new and novel which didn't exist before and have lots of value in their own right. Let's be a bit reductive and pretend Moonshot had just outright stolen the weights from Fable somehow, clearly that wouldn't be contributing anything really new or novel. Now of course they've distilled rather than stolen, but the point is similar: how much value have they added along the way?

gozucito 2 days ago | parent [-]

Since this is HN Think of it like one of the GPL license for software.

It's ok for me to use your source code for free as long as I then let others also use my source code for free.

joshuamorton 3 days ago | parent | prev [-]

So, I'm not the person you were responding to. I'd like you to take a moment and suggest where anyone in the thread you're replying to, either me or Remus, has said anything that suggests disagreement with the statement

> There is an even higher level of creativity in creating books, songs and all sorts of art used in model training though.

He claimed there was more creativity in model training than in model distillation. That makes no claim about the relationship between the creativity in model creation and art. Why are you continuing to attack a claim that was never made, after a sub thread very explicitly clarifying that that claim was not made?

gozucito 3 days ago | parent [-]

This is the post Nemus was replying to:

>Perhaps even more importantly, the current frontier LLM models are self-admittedly the product of enormous quantities of copyright infringement and even less savory inputs, so calling them out for distilling the fruit of that tainted tree reads as highly hypocritical at best.

Context is important. And in this context, their argument only mentions creativity when it belongs to an AI lab. That omission is the blind spot I pointed out. Bottom line is whether or not Anthropic are being hypocritical and yes, they most definitely are, regardless of any attempted sophistry.

There is a reason courts want you to tell "The whole truth" and not just "the truth".

perching_aix 3 days ago | parent | prev | next [-]

No? They outright say the opposite!

Like look, I'm not a native speaker, sure. But I think when someone says "value add", that means there was value there (which you claim they're rhetorically erasing), and then that was added to. Under no interpretation of this phrase do I get an erasure of prior value.

So certainly, as long as words mean anything, no, they absolutely did not say or suggest what you claim they did, and what you extract a thus unreasonable amount of obnoxious schadenfreude from, while throwing in a cheap insult for funsies at the end.

It's the second time I feel compelled to reach for this just today: https://i.kym-cdn.com/photos/images/original/002/659/979/108...

bluegatty 3 days ago | parent | prev [-]

This is a misrepresentation though.

The LLM output, is not the same as the input - there is value add.

Of course works used as raw inputs to LLMs required work and are reasonably subject to IP concerns - but they are different.

It's possible that the LLM makers 'owe' the content creators that created the content they used to make their products - it's an interesting but separate question.

We could very well end up where content IP is protected, LLM output is not and visa versa with reasonable legal founding, doubtful but plausible.

blks 3 days ago | parent | next [-]

Lossly storing IP in LLM itself, and using IP for training (so it’s lossly stored in LLM), without licensing these works or otherwise following license agreements (eg GPL) is infringement. Using then this product for commercial activity is a smoking gun.

bluegatty 3 days ago | parent [-]

"Lossly storing IP in LLM itself, a" - that part I'm inclined to agree with.

But it's debatable if that's the case.

Google stores copyrighted content and produces in in their product.

Also - it's fair game to use snippets of things here and there, if the derived work is novel, which I think it is for LLMs, mostly.

I do agree though, that we ought to draw the line somehow.

jaggederest 3 days ago | parent | prev [-]

> but they are different.

How, and why?

> We could very well end up where content IP is protected, LLM output is not and visa versa with reasonable legal founding, doubtful but plausible.

That is the current state of legal rulings - LLM output is public domain, not copyrightable.

semiquaver 3 days ago | parent | next [-]

This misstates the small number of legal opinions and orders on this topic, none of which form binding precedent outside the districts where the cases happened. So even if a court had found that “LLM output is public domain” (none did) that wouldn’t make it “the law” until it went up the appellate system and was upheld.

Our current laws simply weren’t built for this and I expect the legal status of LLM output is not going to be resolved until Congress actually legislates on this topic.

bluegatty 3 days ago | parent | prev [-]

"> but they are different.

How, and why?"

How are they even remotely the same?

They're not even used the same way.

One is raw data input, the other is training content - designed to train LLMs.

One is a set of IP derived for other purposes entirely, and has esablished IP law - how you can use someone else's creative work or not ... for LLM outputs, less clear.