Remix.run Logo
palmotea 13 hours ago

> you don’t address the ridiculous “copyright forever, nothing goes public domain” policies that got us to this place

1. You're mischaracterizing and exaggerating the policies to falsely support your point. Things go into the public domain all the time: https://web.law.duke.edu/cspd/publicdomainday/2026/.

2. It's not like the model companies had the attitude "oh copyright is too long, we disagree on the term." Their attitude was precisely: "We want it and we don't give a fuck about you. Published yesterday, published 50 years ago? We will take it, we won't pay for it, and we will use it to replace you and make ourselves rich. Don't like it? Suck my data center."

bloppe 13 hours ago | parent | next [-]

They're attitude is more "this is fair use" which, according to all precedent, is probably true in most cases (unless the models actually start regurgitating huge parts of the Copyrighted material without a license).

Of course, distillation is also fair use under Copyright law.

Copyright was never meant to be a moral framework. It was always a practical framework designed to incentivize publishing that would ultimately pass into the public domain. Everybody seems to want to attribute some sort of moral weight to it though; the idea that people are naturally entitled to certain rights over things they've published. That idea would be totally alien to the people who designed the Copyright system in the first place.

adrian_b 12 hours ago | parent | next [-]

"This is fair use" is just their public defense, but obviously they have never given a thought about this when hoarding data, as shown by the modest 1.5B fine of Anthropic, which they can now write off as a normal business cost.

I am completely willing to accept that "this is fair use" for any company that publishes the LLM weights, i.e. the result of processing all the copyrighted work, because they have performed a public service with this.

But when the so-called "fair use" was a method to transform public data into private data that they guard and claim that any access to it would now be IP theft and which they use to obtain huge profits, that does not look like fair use to me.

dTal 12 hours ago | parent | prev | next [-]

Was it really "designed to incentivize" anything? Or was it introduced to protect a powerful, influential business model? Looking at how laws are passed now, I know which explanation I find more congruent.

gruez 12 hours ago | parent | next [-]

It's literally in the constitution:

> [the United States Congress shall have power] To promote the Progress of Science and useful Arts, by securing for limited Times to Authors and Inventors the exclusive Right to their respective Writings and Discoveries.

https://en.wikipedia.org/wiki/Copyright_Clause

dTal 5 hours ago | parent [-]

It literally predates the United States Constitution by hundreds of years: https://en.wikipedia.org/wiki/History_of_copyright

Also, I'd be careful at taking the reasoning of political documents at face value. Many items in the Constitution are post-hoc, Lockeian/liberal justifications for a social order that was in fact largely copied over wholesale from English parliamentary monarchy, with surprisingly few tweaks.

bloppe 12 hours ago | parent | prev [-]

I think the expansion of terms over time is a bit damning, but originally the term was very pro-public-domain (sometimes as low as 7 years): https://en.wikipedia.org/wiki/History_of_copyright#/media/Fi...

dTal 5 hours ago | parent [-]

I think we have to consider the origin of it as a concept, which predates the United States entirely and is quite a lot more damning. From the same Wikipedia page you linked:

"The first copyright privilege in England bears date 1518 and was issued to Richard Pynson, King's Printer, the successor to William Caxton. The privilege gives a monopoly for the term of two years. The date is 15 years later than that of the first privilege issued in France. Early copyright privileges were called "monopolies," particularly during the reign of Queen Elizabeth, who frequently gave grants of monopolies in articles of common use, such as salt, leather, coal, soap, cards, beer, and wine. The practice was continued until the Statute of Monopolies was enacted in 1623, ending most monopolies, with certain exceptions, such as patents; after 1623, grants of letters patent to publishers became common...

As the "menace" of printing spread, governments established centralized control mechanisms,[19] and in 1557 the English Crown thought to stem the flow of seditious and heretical books by chartering the Stationers' Company. The right to print was limited to the members of that guild, and thirty years later the Star Chamber was chartered to curtail the "greate enormities and abuses" of "dyvers contentyous and disorderlye persons professinge the arte or mystere of pryntinge or selling of books." The right to print was restricted to two universities and to the 21 existing printers in the city of London, which had 53 printing presses. The French crown also repressed printing, and printer Etienne Dolet was burned at the stake in 1546. As the English took control of type founding in 1637, printers fled to the Netherlands. Confrontation with authority made printers radical and rebellious, and 800 authors, printers and book dealers were incarcerated in the Bastille before it was stormed in 1789.[19]"

So, to summarize: the principle of copyright came from monarchic economic protectionism and censorship. I will freely admit I didn't know this piece of history before this thread - I simply predicted it, correctly, from first principles.

palmotea 9 hours ago | parent | prev | next [-]

> Copyright was never meant to be a moral framework. It was always a practical framework designed to incentivize publishing that would ultimately pass into the public domain.

A practical framework that model training breaks. Why publish a resource if it'll just get ingested by a model, and the model maker will get your customers/users instead of you? You're already seeing that with Google, which uses AI Overviews to cannibalize more and more traffic that would have once passed to a website.

Ekaros 13 hours ago | parent | prev | next [-]

Certainly in Europe there is view that author has moral rights over their work. And the view has affected how international copyright framework operates.

SimianSci 12 hours ago | parent | prev [-]

What a soulless take. Morality is determined socially and does not exist in a vacuum. Stealing the livlihood of artists and creators so that you can offer competing products is not an act of neutrality. It is a deeply immoral act akin to theft. Stealing does not magically become "distillation" once you've stolen from enough people that it becomes difficult to match provenance.

The law may see this as no issue, as the law cares more about protecting power, but that does not mean it isnt immoral.

gruez 13 hours ago | parent | prev [-]

>We will take it, we won't pay for it, and we will use it to replace you and make ourselves rich

The whole point of fair use (which courts have so far ruled AI training is) is that you don't have to ask for permission or pay them for it.

palmotea 13 hours ago | parent | next [-]

> The whole point of fair use (which courts have so far ruled AI training is) is that you don't have to ask for permission or pay them for it.

Is that settled law? I doubt it.

And IMHO, AI training violates the spirit of fair use. It's not really a method of criticism or commentary. It's a system to use people's own work to undermine their ability to economically subsist on that work.

Though how about this for a proposed exception: you can AI train on anything as fair use: only if release your model and weights public domain open source.

dragonwriter 12 hours ago | parent | next [-]

> Is that settled law?

It seems to be the fairly consistent approach of trial courts addressing the question under different soecific fact patterns in different contexts; its not “settled law” in the sense of nationally binding precedent (which would take either a Supreme Court ruling kr separate appellate rulings in every circuit).

> And IMHO, AI training violates the spirit of fair use. It's not really a method of criticism or commentary.

Plenty of transformative uses that have been held to be fair use are not criticism or commentary, and AI training is a transformative use where, for any individual work used, the end product is both a very different class of work and the used work indiviudally has very small impact on the final work.

> It's a system to use people's own work to undermine their ability to economically subsist on that work.

Courts seem to disagree that this is generally the case with AI training as such. (And AI training being fair use would not make the use of models to create works that would otherwise be infringing copies with that function through inference any less infringing.)

> Though how about this for a proposed exception: you can AI train on anything as fair use: only if release your model and weights public domain open source.

You are, of course, free to try to convince Congress to amend copyright law to apply that rule (though since the current statutory form of the fair use rule is itself a legislative adoption that follows pre-existing court rulings on fair use as a Constitutional limit on the copyright power stemming from the First Amendment, Congress may not actually have the power to narrow it that way.)

gruez 12 hours ago | parent | prev [-]

>Is that settled law? I doubt it.

That just seems like a cope unless you have actual evidence that the lower/appellate courts have misruled. And no, "I don't like the ruling because [all the reasons AI is bad]" doesn't count, you need actual legal justifications, preferably from legal experts. Not to mention that even if the supreme court ruled on it, it's not really "settled", eg. Roe. v. Wade and Humphrey's Executor v. United States being overturned

palmotea 12 hours ago | parent [-]

> That just seems like a cope unless you have actual evidence that the lower/appellate courts have misruled.

No, it means lower court judges get overruled all the time and it's not like the courts and law always functions as some dispassionate applicators of some fixed framework. It's not settled until the process gets worked much farther than it probably has.

selectodude 13 hours ago | parent | prev [-]

When you realize that LLMs are “just” extremely efficient lossy data compression, it’s hard for me to see how it’s anything other than taking people’s shit, putting it into a gigantic zip file, and letting people search against it.

gruez 13 hours ago | parent [-]

Wait till you hear about Perfect 10, Inc. v. Amazon.com, Inc. (2007) and Authors Guild, Inc. v. Google, Inc. (2015), both of which ruled that lossy and verbatim copies (respectively) are allowed for for-profit use.

selectodude 12 hours ago | parent [-]

Too late. Authors Guild, Inc. v. Google, Inc. is a good one too because Internet Archive got the exact opposite outcome in court for doing the exact same thing. I recognize the bullshit, I just call it out to keep myself sane.

gruez 12 hours ago | parent [-]

>Internet Archive got the exact opposite outcome in court for doing the exact same thing

No, it's not the same thing. Contrary to what many people think, "fair use" isn't something you can invoke to do whatever copyright infringement you want. The judge is supposed to consider several factors, one of which is whether the work was "transformative". In google's case it was offering search results. Internet archive was operating a "digital library" (aka. a filesharing site). Whatever you hate about AI companies sucking up electricity and displacing jobs, they're certainly more transformative (and arguably more transformative than even google search) than whatever the internet archive was doing.

selectodude 12 hours ago | parent [-]

That’s not true. 1. Libraries have used Authors Guild as legal cover to lend out ebooks for paper books that they own. 2. Google provided access to the whole book, that’s why they got sued.

If I run a book through AES, that’s pretty transformative too!

gruez 12 hours ago | parent [-]

>1. Libraries have used Authors Guild as legal cover to lend out ebooks for paper books that they own

And has this been tested in court? After all, you see people uploading tv shows on youtube, then pasting a snippet of fair use in the description. That doesn't make it true. If anything, the unfavorable ruling for internet archive suggests libraries were incorrect with their interpretation of the law.

>2. Google provided access to the whole book, that’s why they got sued.

No it didn't. From wikipedia:

"For works still under copyright, Google scanned and entered the whole work into their searchable database, but only provided "snippet views" of the scanned pages in search results to users."