Remix.run Logo
dgellow 3 days ago

> on the same level like Anthropic scraped copyright protected material for their training.

I see no problem with distillation, on the other hand the complete dismissal of copyright by AI labs is pretty bad, I don’t think we should put them at the same level

SR2Z 3 days ago | parent | next [-]

> on the other hand the complete dismissal of copyright by AI labs

Courts keep ruling over and over that an LLM trained on copyrighted works qualifies as a transformative work and is therefore fair use. They don't have to dismiss copyright law, this has always been allowed.

The only thing they get in trouble for is pirating the works to get their hands on them.

dijksterhuis 3 days ago | parent | next [-]

> Courts keep ruling over and over that an LLM trained on copyrighted works qualifies as a transformative work and is therefore fair use. They don't have to dismiss copyright law, this has always been allowed.

*USA only.

the UK has fair dealing, which is more restrictive

https://www.gov.uk/guidance/exceptions-to-copyright#fair-dea...

https://www.britishcopyright.org/wp-content/uploads/BCC-Fair...

Diogenesian 3 days ago | parent | prev | next [-]

"Keep ruling over and over" is way too strong. There have maybe been two rulings, nothing nationally binding, and most of the litigation is still ongoing. In particular, last I checked OpenAI and Microsoft are still badly threatened by the NYT lawsuit: https://law.justia.com/cases/federal/district-courts/new-yor... https://www.cnet.com/tech/services-and-software/publishers-o...

This will have to wait for the Supreme Court. OpenAI and Microsoft 100% deserve to lose, even without OpenAI allegedly hiding evidence.

dijksterhuis 3 days ago | parent | next [-]

> "Keep ruling over and over" is way too strong. There have maybe been two rulings, nothing nationally binding, and most of the litigation is still ongoing.

yep

https://www.britishcopyright.org/wp-content/uploads/BCC-Fair...

> This ambiguity has resulted in extensive litigation on the limits of Fair Use to AI development. Currently, we only have 3 first instance decisions out of the 53 cases being tried. It will likely take a decade before we understand how Fair Use applies to any one step in AI training, let alone all.

> In the three lower court decisions so far, one held Fair Use did not apply (Thomson v Ross), one held Fair Use could apply (Kadrey v Meta) with the court suggesting more evidence was needed on the fourth factor ‘harm to the market’, and the third case held Fair Use may apply to some AI. As Fair Use is dependent on the specific facts at issue, none of these cases help educate the market or the public as to the limits of Fair Use in AI contexts.

asadotzler 3 days ago | parent | prev [-]

Not to mention the cases where the AI labs would have lost in court so bailed and settled for billions. Just this week, Anthropic agreed to pay $1.5B in a settlement to avoid losing a pretty cut and dry case.

Diogenesian 3 days ago | parent | next [-]

To be clear that was one of the few resolved cases where the judge agreed training was fair use. But the piracy was enough of a distraction that I don't consider that a particularly useful precedent. I am much more interested in the NYT case, which quite clearly shows GPT was trained on NYT articles and can spit them out verbatim (and has since been validated by academic research; all the commercial models are capable of mass plagiarism).

xienze 3 days ago | parent | prev [-]

> Anthropic agreed to pay $1.5B in a settlement to avoid losing a pretty cut and dry case.

It takes two parties to agree to a settlement. That the other party agreed to a settlement instead of taking it to court implies this was not the slam dunk you may think it was.

robocat 3 days ago | parent [-]

You're both reading tea leaves.

Settling just says that they expected the internal costs or risks to be more than 1.5 billion cashflow.

  In the $65B in Series H funding at $965B post-money valuation they said their run-rate revenue crossed $47B annualised.
With those numbers, there can be sound financial reasons for wanting to just get rid of the lawsuit.

Also if it ends up that other competitors also need to pay $1.5 billion, then maybe that does or doesn't have a competitive advantage.

Anthropic's business and legal strategies are not public. I would expect there to be multiple legs/reasons for settlement even for a decision below 1%. Trying to create a single narrative is what us spectators do.

xienze 3 days ago | parent [-]

> With those numbers, there can be sound financial reasons for wanting to just get rid of the lawsuit.

Yes, of course, on Anthropic's side. Why would the other side agree to a settlement?

robocat 3 days ago | parent | next [-]

Perhaps Anthropic is indirectly paying to have their competitors sued...

My narritive is that the terms of the settlement would be full and final.

It was a class action, with payment going to authors and publishers, and the legal team will get paid too.

My guess is that funding is a major issue for the legal team. Authors presumably can't pay for lawyers unless a percentage of winnings, although publishers may have invested.

But the legal team will ask the beneficiaries to use some of the warchest to fund different campaigns against every other AI company. I would assume the legal team wants to win again. They've now got a good story to sell to rights holders, who presumably like money and don't like risks.

I haven't even got to my armchair yet this morning.

SR2Z 2 days ago | parent | prev [-]

Because most authors don't make a ton of money from their books, and even a relatively low settlement from Anthropic is a large enough sum that they're OK with taking it.

Madmallard 3 days ago | parent | prev [-]

Man that's depressing to read someone defending this

SR2Z 2 days ago | parent [-]

I'm telling you what the law plainly says. I personally think that the exceptions for fair use made sense before the transformer was invented, and they still make sense now.

The two main problems with copyright also haven't changed: copyright lasts too long and is too expensive to defend.

whywhywhywhy 3 days ago | parent | prev | next [-]

I don't think anyone is really dismissing it, just pointing out the audacity of complaining about distillation after stealing so much themselves is comical.

mrtesthah 3 days ago | parent | prev [-]

The amount of original, copyrightable and trademarkable IP actually created by the AI labs themselves is dwarfed by their staggeringly vast infringement activities.