Remix.run Logo
madduci 3 days ago

So what is the issue here? Distilling is still fair, on the same level like Anthropic scraped copyright protected material for their training.

So here robbers are blaming robbers?

These claims are just pointless, everytime

dgellow 3 days ago | parent | next [-]

> on the same level like Anthropic scraped copyright protected material for their training.

I see no problem with distillation, on the other hand the complete dismissal of copyright by AI labs is pretty bad, I don’t think we should put them at the same level

SR2Z 3 days ago | parent | next [-]

> on the other hand the complete dismissal of copyright by AI labs

Courts keep ruling over and over that an LLM trained on copyrighted works qualifies as a transformative work and is therefore fair use. They don't have to dismiss copyright law, this has always been allowed.

The only thing they get in trouble for is pirating the works to get their hands on them.

dijksterhuis 3 days ago | parent | next [-]

> Courts keep ruling over and over that an LLM trained on copyrighted works qualifies as a transformative work and is therefore fair use. They don't have to dismiss copyright law, this has always been allowed.

*USA only.

the UK has fair dealing, which is more restrictive

https://www.gov.uk/guidance/exceptions-to-copyright#fair-dea...

https://www.britishcopyright.org/wp-content/uploads/BCC-Fair...

Diogenesian 3 days ago | parent | prev | next [-]

"Keep ruling over and over" is way too strong. There have maybe been two rulings, nothing nationally binding, and most of the litigation is still ongoing. In particular, last I checked OpenAI and Microsoft are still badly threatened by the NYT lawsuit: https://law.justia.com/cases/federal/district-courts/new-yor... https://www.cnet.com/tech/services-and-software/publishers-o...

This will have to wait for the Supreme Court. OpenAI and Microsoft 100% deserve to lose, even without OpenAI allegedly hiding evidence.

dijksterhuis 3 days ago | parent | next [-]

> "Keep ruling over and over" is way too strong. There have maybe been two rulings, nothing nationally binding, and most of the litigation is still ongoing.

yep

https://www.britishcopyright.org/wp-content/uploads/BCC-Fair...

> This ambiguity has resulted in extensive litigation on the limits of Fair Use to AI development. Currently, we only have 3 first instance decisions out of the 53 cases being tried. It will likely take a decade before we understand how Fair Use applies to any one step in AI training, let alone all.

> In the three lower court decisions so far, one held Fair Use did not apply (Thomson v Ross), one held Fair Use could apply (Kadrey v Meta) with the court suggesting more evidence was needed on the fourth factor ‘harm to the market’, and the third case held Fair Use may apply to some AI. As Fair Use is dependent on the specific facts at issue, none of these cases help educate the market or the public as to the limits of Fair Use in AI contexts.

asadotzler 3 days ago | parent | prev [-]

Not to mention the cases where the AI labs would have lost in court so bailed and settled for billions. Just this week, Anthropic agreed to pay $1.5B in a settlement to avoid losing a pretty cut and dry case.

Diogenesian 3 days ago | parent | next [-]

To be clear that was one of the few resolved cases where the judge agreed training was fair use. But the piracy was enough of a distraction that I don't consider that a particularly useful precedent. I am much more interested in the NYT case, which quite clearly shows GPT was trained on NYT articles and can spit them out verbatim (and has since been validated by academic research; all the commercial models are capable of mass plagiarism).

xienze 3 days ago | parent | prev [-]

> Anthropic agreed to pay $1.5B in a settlement to avoid losing a pretty cut and dry case.

It takes two parties to agree to a settlement. That the other party agreed to a settlement instead of taking it to court implies this was not the slam dunk you may think it was.

robocat 3 days ago | parent [-]

You're both reading tea leaves.

Settling just says that they expected the internal costs or risks to be more than 1.5 billion cashflow.

  In the $65B in Series H funding at $965B post-money valuation they said their run-rate revenue crossed $47B annualised.
With those numbers, there can be sound financial reasons for wanting to just get rid of the lawsuit.

Also if it ends up that other competitors also need to pay $1.5 billion, then maybe that does or doesn't have a competitive advantage.

Anthropic's business and legal strategies are not public. I would expect there to be multiple legs/reasons for settlement even for a decision below 1%. Trying to create a single narrative is what us spectators do.

xienze 3 days ago | parent [-]

> With those numbers, there can be sound financial reasons for wanting to just get rid of the lawsuit.

Yes, of course, on Anthropic's side. Why would the other side agree to a settlement?

robocat 3 days ago | parent | next [-]

Perhaps Anthropic is indirectly paying to have their competitors sued...

My narritive is that the terms of the settlement would be full and final.

It was a class action, with payment going to authors and publishers, and the legal team will get paid too.

My guess is that funding is a major issue for the legal team. Authors presumably can't pay for lawyers unless a percentage of winnings, although publishers may have invested.

But the legal team will ask the beneficiaries to use some of the warchest to fund different campaigns against every other AI company. I would assume the legal team wants to win again. They've now got a good story to sell to rights holders, who presumably like money and don't like risks.

I haven't even got to my armchair yet this morning.

SR2Z 2 days ago | parent | prev [-]

Because most authors don't make a ton of money from their books, and even a relatively low settlement from Anthropic is a large enough sum that they're OK with taking it.

Madmallard 3 days ago | parent | prev [-]

Man that's depressing to read someone defending this

SR2Z 2 days ago | parent [-]

I'm telling you what the law plainly says. I personally think that the exceptions for fair use made sense before the transformer was invented, and they still make sense now.

The two main problems with copyright also haven't changed: copyright lasts too long and is too expensive to defend.

whywhywhywhy 3 days ago | parent | prev | next [-]

I don't think anyone is really dismissing it, just pointing out the audacity of complaining about distillation after stealing so much themselves is comical.

mrtesthah 3 days ago | parent | prev [-]

The amount of original, copyrightable and trademarkable IP actually created by the AI labs themselves is dwarfed by their staggeringly vast infringement activities.

orangecat 3 days ago | parent | prev | next [-]

Distilling is still fair

I generally agree, in the same sense that it's "fair" for the US and China to spy on each other. It's not a moral outrage, but it is something that the targets can and should try to prevent.

archagon 3 days ago | parent [-]

Outrageous only to the died-in-wool corpocrats.

3 days ago | parent | prev | next [-]
[deleted]
Lalabadie 3 days ago | parent | prev | next [-]

"You are trying to kidnap what I have rightfully stolen!"

sciencesama 3 days ago | parent | next [-]

the whole AI is just internet distilled !!

azinman2 3 days ago | parent | prev [-]

Except it’s not just a dump of the internet, which Moonshot also did themselves (and probably used even more pirated content as laws in China are different without any recourse for the entire world). I don’t know why this is so unclear to folks.

trollbridge 3 days ago | parent [-]

Chinese IP law is actually quite solid. You do have to register your trademarks and copyrights properly in China, and then lawsuits have to be filed appropriately according to Chinese law. Is that a problem?

JKCalhoun 3 days ago | parent | prev | next [-]

I'm by no means taking the side of the AI companies, but it's possible that Anthropic "added value" to the data they harvested. Stealing that does seem kind of uncool.

Regardless, it was always inevitable—will continue to happen.

mrhottakes 3 days ago | parent | next [-]

So as long as Kimi added value to Fable, it's fine? Sounds good.

knollimar 3 days ago | parent [-]

Moonshot can prove they added value with their paper. Where's Anthropic's proof?

6gvONxR4sf7o 3 days ago | parent | prev | next [-]

I wonder how the "added value" argument can apply to Anthropic and not to the Kimi team. If Claude's value is that you don't have to pay a team of slow expensive subject-matter experts, and Kimi's value is that you don't have to pay Claude, it just seems like the same thing.

oliculipolicula 3 days ago | parent | prev [-]

Valuation is hard to perform when it's deep inside a black box. Ther "API" may be easier to evaluate. The problem with this angle is that Moonshot is actually producing _better_ value from Anthropic's blackbox.

Technically, providing better value from your competitor's private holdings could be theft (of trade secrets), but might it also be fair use? "Schrodinger's IP" be damned.

I don't think the 1.5B settlement has resolved this. The 2 cases need to be merged!

Matl 3 days ago | parent | prev | next [-]

> So what is the issue here?

The issue seems to be the US only likes competition when it is winning. Markets in Asia are meant for cheap labor and resources, they're not meant to actually compete. /s

matheusmoreira 3 days ago | parent | next [-]

> The issue seems to be the US only likes competition when it is winning.

This. Free markets for everyone when they're the dominant economic force. Protectionism, tariffs and import/export controls when they're not.

It's so disgusting.

catigula 3 days ago | parent | prev [-]

[flagged]

ceejayoz 3 days ago | parent | next [-]

Everything you just said describes the major American AI providers.

Anthropic just settled a $1.5B suit over it!

fwip 3 days ago | parent [-]

[flagged]

matheusmoreira 3 days ago | parent | prev | next [-]

> Stealing IP in a way that destroys the economic incentives

Like the US did when it "stole" the textiles IP from the UK in order to kickstart its own industry?

> The industry cannot sustain itself if that’s the model

Then let it fall apart.

soperj 3 days ago | parent | prev | next [-]

> a bad actor that leverages Ip theft wholesale.

It's like they've read the history of the US and how it got to where it is in the first place.

amanaplanacanal 3 days ago | parent | prev | next [-]

What IP is being stolen here? So far, the courts have ruled that anything generated by an LLM is not copyrightable.

rickydroll 3 days ago | parent | prev | next [-]

Stealing IP is how American industry got started. Goose: gander, pot: kettle.

It is what built and sustains the movie and music industries. See: work for hire and 100+year copyright length

The tech industry: see: copyright and patent assignment from discoverer to corporation.

I know that corporations forcing me to assign patents and copyright to them was an incentive to take published works from "software practice and experience" and other technical journals, use them as the core of my work, and disclose that source to the company I worked for. Didn't stop them from applying for patents, however.

I think the discussion of copyright needs more refinement. We need to separate the discoverer's need for acknowledgment of development effort from the rent-seeking core of copyright.

ux266478 3 days ago | parent [-]

The things you listed are very far downstream of the start of American industry, which is in primary resource extraction and processing. Which is you know, what actually built the country. The media industry has always been materially irrelevant, and what we think of as the tech industry is extremely new.

You're right to call out the nasty environment surrounding intellectual property in the US and the exploitation of ideation in general, you just needed a correction on that. Someone else in this chain said virtually the same thing, which is a weird coincidence of historical ignorance. Not too weird, people tend to forget the 18th and 19th centuries happened, and much of the causally important wheels of the world are in the unsexy grease pits nobody wants to think about.

rickydroll 3 days ago | parent [-]

You're right, I didn't include stuff at the beginning, for example, the theft of IP in textile manufacturing in the late 1700s. The US government didn't recognize copyrights on foreign literature which let US publishers reprint things such as Gilbert, Sullivan's operettas and Dickens novels

Then there is Alexander Hamilton's advocacy for importing foreign technicians that bring back IP and reproduce it here in the states. Best of all was the patent act of 1793 which like with the literature copyright ignoring, let us citizens patent inventions from the other side of the pond.

The founding fathers definitely had the right idea on IP.

ux266478 3 days ago | parent [-]

Which is to say they didn't have much of an idea at all, because it really didn't exist in much the same way. In fact, this is still a inaccurate characterization for exactly that reason. On the basis of copyright, take for instance the idea of exclusive rights to print a work. This actually wasn't implemented as a method of protection for the author, but a political reaction to the printing press being "misused" in the eyes of the anglo-colonialist entity controlling the British isles, and so was a means of preventing the mass creation of undesirable literature.

On the basis of patents, it didn't quite have nearly as much of a history of mutation culturally, but did experience massive whiplash in purpose and application following the implementation of globalism. What was once a system to protect technical innovation on an individual level, would find new purpose as a means to provide structure to an increasingly complicated and internationalized dynamic market. Another means of bureaucratic organization. Then, once again, the context and purpose would change when the world developed digital globalism. The entire engine of IP as a legal fiction became a significant geopolitical tool in an increasingly cramped and fragile world, a necessary gimmick holding up the sky.

Unfortunately, not much to do at this point. It'll likely only become even more nonsensically important as time wears on, until the globalist system collapses. It's certainly possible it'll even be the confounding factor that causes the great unraveling, though the problems hardly begin and end with IP. It was just a useful legal fiction in the wrong place at the wrong time.

rickydroll 2 days ago | parent [-]

I guess I should thank you for adding yet another load of deeper reading into American history. :)

nickphx 3 days ago | parent | prev [-]

Oh, ok. How would you describe how the "frontier us companies" acquired the data used to form their models?

petilon 3 days ago | parent | prev | next [-]

[dead]

xnoto 3 days ago | parent | prev | next [-]

++

softwaredoug 3 days ago | parent | prev | next [-]

If they did this in the US they would almost certainly be sued.

Meta, for examples, doesn’t want employees to use Claude Code due to distillation risk.

trollbridge 3 days ago | parent [-]

It turns out U.S. law doesn’t have jurisdiction across the entire world, nor does Anthropic and OAI’s rather blatant attempt to buy government influence.

softwaredoug 3 days ago | parent [-]

Well its not law. Its more terms of service.

For example, if OpenAI / Anthropic were actually open, other US labs could be building near-frontier open weights models by distilling off OpenAI / Anthropic. But because US companies don't want to be sued, US labs who obey terms of service, will be at a disadvantage to Chinese peers.

Maybe US labs need to just not care and distill from OpenAI / Anthropic anyways?

make3 3 days ago | parent | prev [-]

It's about the claim of whether these companies could develop a similarly powerful model without larger companies building their own first, which is an important point, and it's likely not the case.

It's also about the larger companies explaining why they can't be as efficient, of course they can't, they're not just ripping the outputs of another model that someone else invested billions to train.

IncreasePosts 3 days ago | parent | next [-]

Why would that matter? OpenAI or whatever frontier lab couldn't have built their frontier models without the entirety of humanity unknowingly developing their training set for 5000 years.

It would be one thing if Moonshot was breaking into OpenAI servers and stealing trade secrets, but the only thing they are doing is looking at the output of the program, which is exactly the service that OpenAI offers. So, at best, this is a ToS violation. Sucks for the frontier labs I suppose, but live by the sword - die by the sword.

mrhottakes 3 days ago | parent | prev | next [-]

> they're not just ripping the outputs of another model that someone else invested billions to train.

True, they're simply ripping the inputs that humanity invested thousands of years and trillions of dollars to produce.

make3 3 days ago | parent [-]

no argument from me here

doctoboggan 3 days ago | parent | prev | next [-]

Yeah agreed, from one standpoint I couldn't care less that they did a "distillation attack", but I am interested in knowing if China is able to develop open weight frontier models without the prior existence of a huge model to distill from.

PaulHoule 3 days ago | parent | prev | next [-]

Simply knowing it is possible to do something makes it easier to do.

cindyllm 3 days ago | parent [-]

[dead]

xcf_seetan 3 days ago | parent | prev [-]

> they're not just ripping the outputs of another model that someone else invested billions to train.

If they payed for inference, doesn't they own the output? So if I pay for a model to generate code, isn't that code mine to do with it whatever I want? Just curious.

make3 3 days ago | parent [-]

Not arguing for the morality of it, but if we're going by the law because that's what you're using in your comment ("don't I own" which only matters wrt the law), then you explicitly accepted a Terms of Use which excludes distillation as a use case.

Now of course they themselves trained on the whole Internet for free, etc.