Remix.run Logo
brookst 13 hours ago

The thing about screaming and waving your arms is it makes it hard for people to connect with the argument.

A quick scan shows a lack of rigor that usually points towards zealotry rather than a reasoned position: you imply that models are only trained on copyrighted work, you don’t address the ridiculous “copyright forever, nothing goes public domain” policies that got us to this place, you don’t distinguish between training on an 80 year old copyrighted work whose author is dead and rights are just in a company portfolio, and you assert that neither OpenAI nor Anthropic has ever done anything altruistic.

If you enjoy the screaming, have at it. If you want to persuade people, dialing it down a few notches and at least giving token acknowledgement that nuance exists would help.

binoct 13 hours ago | parent | next [-]

I don’t see any screaming and waving arms in the parent. It’s a passionate but well reasoned and carefully presented case. If that is to be dismissed as zealotry we’ll all need to stick to writing journal papers to avoid the label.

If anything your “copyright forever” drumbeat and incorrect read of “only being trained on copyrighted works” are more hyperbolic than anything you’re responding to.

Also, the models from China are only being released open weights as competitive positioning against the globally more used competition. If the Chinese companies had a dominant market position they’d be falling prey to the same forces that have caused OpenAI and Anthropic to behave the way they do.

palmotea 13 hours ago | parent | prev | next [-]

> you don’t address the ridiculous “copyright forever, nothing goes public domain” policies that got us to this place

1. You're mischaracterizing and exaggerating the policies to falsely support your point. Things go into the public domain all the time: https://web.law.duke.edu/cspd/publicdomainday/2026/.

2. It's not like the model companies had the attitude "oh copyright is too long, we disagree on the term." Their attitude was precisely: "We want it and we don't give a fuck about you. Published yesterday, published 50 years ago? We will take it, we won't pay for it, and we will use it to replace you and make ourselves rich. Don't like it? Suck my data center."

bloppe 13 hours ago | parent | next [-]

They're attitude is more "this is fair use" which, according to all precedent, is probably true in most cases (unless the models actually start regurgitating huge parts of the Copyrighted material without a license).

Of course, distillation is also fair use under Copyright law.

Copyright was never meant to be a moral framework. It was always a practical framework designed to incentivize publishing that would ultimately pass into the public domain. Everybody seems to want to attribute some sort of moral weight to it though; the idea that people are naturally entitled to certain rights over things they've published. That idea would be totally alien to the people who designed the Copyright system in the first place.

adrian_b 12 hours ago | parent | next [-]

"This is fair use" is just their public defense, but obviously they have never given a thought about this when hoarding data, as shown by the modest 1.5B fine of Anthropic, which they can now write off as a normal business cost.

I am completely willing to accept that "this is fair use" for any company that publishes the LLM weights, i.e. the result of processing all the copyrighted work, because they have performed a public service with this.

But when the so-called "fair use" was a method to transform public data into private data that they guard and claim that any access to it would now be IP theft and which they use to obtain huge profits, that does not look like fair use to me.

dTal 12 hours ago | parent | prev | next [-]

Was it really "designed to incentivize" anything? Or was it introduced to protect a powerful, influential business model? Looking at how laws are passed now, I know which explanation I find more congruent.

gruez 12 hours ago | parent | next [-]

It's literally in the constitution:

> [the United States Congress shall have power] To promote the Progress of Science and useful Arts, by securing for limited Times to Authors and Inventors the exclusive Right to their respective Writings and Discoveries.

https://en.wikipedia.org/wiki/Copyright_Clause

dTal 5 hours ago | parent [-]

It literally predates the United States Constitution by hundreds of years: https://en.wikipedia.org/wiki/History_of_copyright

Also, I'd be careful at taking the reasoning of political documents at face value. Many items in the Constitution are post-hoc, Lockeian/liberal justifications for a social order that was in fact largely copied over wholesale from English parliamentary monarchy, with surprisingly few tweaks.

bloppe 12 hours ago | parent | prev [-]

I think the expansion of terms over time is a bit damning, but originally the term was very pro-public-domain (sometimes as low as 7 years): https://en.wikipedia.org/wiki/History_of_copyright#/media/Fi...

dTal 5 hours ago | parent [-]

I think we have to consider the origin of it as a concept, which predates the United States entirely and is quite a lot more damning. From the same Wikipedia page you linked:

"The first copyright privilege in England bears date 1518 and was issued to Richard Pynson, King's Printer, the successor to William Caxton. The privilege gives a monopoly for the term of two years. The date is 15 years later than that of the first privilege issued in France. Early copyright privileges were called "monopolies," particularly during the reign of Queen Elizabeth, who frequently gave grants of monopolies in articles of common use, such as salt, leather, coal, soap, cards, beer, and wine. The practice was continued until the Statute of Monopolies was enacted in 1623, ending most monopolies, with certain exceptions, such as patents; after 1623, grants of letters patent to publishers became common...

As the "menace" of printing spread, governments established centralized control mechanisms,[19] and in 1557 the English Crown thought to stem the flow of seditious and heretical books by chartering the Stationers' Company. The right to print was limited to the members of that guild, and thirty years later the Star Chamber was chartered to curtail the "greate enormities and abuses" of "dyvers contentyous and disorderlye persons professinge the arte or mystere of pryntinge or selling of books." The right to print was restricted to two universities and to the 21 existing printers in the city of London, which had 53 printing presses. The French crown also repressed printing, and printer Etienne Dolet was burned at the stake in 1546. As the English took control of type founding in 1637, printers fled to the Netherlands. Confrontation with authority made printers radical and rebellious, and 800 authors, printers and book dealers were incarcerated in the Bastille before it was stormed in 1789.[19]"

So, to summarize: the principle of copyright came from monarchic economic protectionism and censorship. I will freely admit I didn't know this piece of history before this thread - I simply predicted it, correctly, from first principles.

palmotea 9 hours ago | parent | prev | next [-]

> Copyright was never meant to be a moral framework. It was always a practical framework designed to incentivize publishing that would ultimately pass into the public domain.

A practical framework that model training breaks. Why publish a resource if it'll just get ingested by a model, and the model maker will get your customers/users instead of you? You're already seeing that with Google, which uses AI Overviews to cannibalize more and more traffic that would have once passed to a website.

Ekaros 13 hours ago | parent | prev | next [-]

Certainly in Europe there is view that author has moral rights over their work. And the view has affected how international copyright framework operates.

SimianSci 12 hours ago | parent | prev [-]

What a soulless take. Morality is determined socially and does not exist in a vacuum. Stealing the livlihood of artists and creators so that you can offer competing products is not an act of neutrality. It is a deeply immoral act akin to theft. Stealing does not magically become "distillation" once you've stolen from enough people that it becomes difficult to match provenance.

The law may see this as no issue, as the law cares more about protecting power, but that does not mean it isnt immoral.

gruez 13 hours ago | parent | prev [-]

>We will take it, we won't pay for it, and we will use it to replace you and make ourselves rich

The whole point of fair use (which courts have so far ruled AI training is) is that you don't have to ask for permission or pay them for it.

palmotea 13 hours ago | parent | next [-]

> The whole point of fair use (which courts have so far ruled AI training is) is that you don't have to ask for permission or pay them for it.

Is that settled law? I doubt it.

And IMHO, AI training violates the spirit of fair use. It's not really a method of criticism or commentary. It's a system to use people's own work to undermine their ability to economically subsist on that work.

Though how about this for a proposed exception: you can AI train on anything as fair use: only if release your model and weights public domain open source.

dragonwriter 12 hours ago | parent | next [-]

> Is that settled law?

It seems to be the fairly consistent approach of trial courts addressing the question under different soecific fact patterns in different contexts; its not “settled law” in the sense of nationally binding precedent (which would take either a Supreme Court ruling kr separate appellate rulings in every circuit).

> And IMHO, AI training violates the spirit of fair use. It's not really a method of criticism or commentary.

Plenty of transformative uses that have been held to be fair use are not criticism or commentary, and AI training is a transformative use where, for any individual work used, the end product is both a very different class of work and the used work indiviudally has very small impact on the final work.

> It's a system to use people's own work to undermine their ability to economically subsist on that work.

Courts seem to disagree that this is generally the case with AI training as such. (And AI training being fair use would not make the use of models to create works that would otherwise be infringing copies with that function through inference any less infringing.)

> Though how about this for a proposed exception: you can AI train on anything as fair use: only if release your model and weights public domain open source.

You are, of course, free to try to convince Congress to amend copyright law to apply that rule (though since the current statutory form of the fair use rule is itself a legislative adoption that follows pre-existing court rulings on fair use as a Constitutional limit on the copyright power stemming from the First Amendment, Congress may not actually have the power to narrow it that way.)

gruez 12 hours ago | parent | prev [-]

>Is that settled law? I doubt it.

That just seems like a cope unless you have actual evidence that the lower/appellate courts have misruled. And no, "I don't like the ruling because [all the reasons AI is bad]" doesn't count, you need actual legal justifications, preferably from legal experts. Not to mention that even if the supreme court ruled on it, it's not really "settled", eg. Roe. v. Wade and Humphrey's Executor v. United States being overturned

palmotea 12 hours ago | parent [-]

> That just seems like a cope unless you have actual evidence that the lower/appellate courts have misruled.

No, it means lower court judges get overruled all the time and it's not like the courts and law always functions as some dispassionate applicators of some fixed framework. It's not settled until the process gets worked much farther than it probably has.

selectodude 13 hours ago | parent | prev [-]

When you realize that LLMs are “just” extremely efficient lossy data compression, it’s hard for me to see how it’s anything other than taking people’s shit, putting it into a gigantic zip file, and letting people search against it.

gruez 13 hours ago | parent [-]

Wait till you hear about Perfect 10, Inc. v. Amazon.com, Inc. (2007) and Authors Guild, Inc. v. Google, Inc. (2015), both of which ruled that lossy and verbatim copies (respectively) are allowed for for-profit use.

selectodude 12 hours ago | parent [-]

Too late. Authors Guild, Inc. v. Google, Inc. is a good one too because Internet Archive got the exact opposite outcome in court for doing the exact same thing. I recognize the bullshit, I just call it out to keep myself sane.

gruez 12 hours ago | parent [-]

>Internet Archive got the exact opposite outcome in court for doing the exact same thing

No, it's not the same thing. Contrary to what many people think, "fair use" isn't something you can invoke to do whatever copyright infringement you want. The judge is supposed to consider several factors, one of which is whether the work was "transformative". In google's case it was offering search results. Internet archive was operating a "digital library" (aka. a filesharing site). Whatever you hate about AI companies sucking up electricity and displacing jobs, they're certainly more transformative (and arguably more transformative than even google search) than whatever the internet archive was doing.

selectodude 12 hours ago | parent [-]

That’s not true. 1. Libraries have used Authors Guild as legal cover to lend out ebooks for paper books that they own. 2. Google provided access to the whole book, that’s why they got sued.

If I run a book through AES, that’s pretty transformative too!

gruez 12 hours ago | parent [-]

>1. Libraries have used Authors Guild as legal cover to lend out ebooks for paper books that they own

And has this been tested in court? After all, you see people uploading tv shows on youtube, then pasting a snippet of fair use in the description. That doesn't make it true. If anything, the unfavorable ruling for internet archive suggests libraries were incorrect with their interpretation of the law.

>2. Google provided access to the whole book, that’s why they got sued.

No it didn't. From wikipedia:

"For works still under copyright, Google scanned and entered the whole work into their searchable database, but only provided "snippet views" of the scanned pages in search results to users."

phailhaus 13 hours ago | parent | prev | next [-]

"Quick scan"? It's a few paragraphs. You can't read a few paragraphs? What's the point of replying when you don't even know what you're replying to?

krs_ 13 hours ago | parent | prev | next [-]

> you don’t address the ridiculous “copyright forever, nothing goes public domain” policies that got us to this place, you don’t distinguish between training on an 80 year old copyrighted work whose author is dead and rights are just in a company portfolio

You're describing what Pirate Party proponents have been championing since the start. Not that I disagree with that notion. Copyright law needs to be changed for modern times. But the laws as currently exist are what they are.

Yokohiii 13 hours ago | parent | prev | next [-]

> The thing about screaming and waving your arms is it makes it hard for people to connect with the argument.

The question is whether people like Altman, Amodei, Musk are part of the common case in Hanlon's razor or they want to trigger emotional responses to gain agency in the mess.

Avicebron 13 hours ago | parent | prev | next [-]

If you would like to discuss the nuances you're welcome to elaborate. If the end result of this nuanced is they get to stay kings of the crab bucket and reap greater and greater rewards it's reasoning at both ends to get the middle you want.

Ensorceled 13 hours ago | parent | prev | next [-]

> A quick scan shows a lack of rigor that usually points towards zealotry rather than a reasoned position

You're doing the same thing:

> you imply that models are only trained on copyrighted work

Actually, they said "The technology is built off wholesale theft of protected works without compensation". The word "only" never appears here, nor is it really implied.

> you don’t address the ridiculous “copyright forever, nothing goes public domain” policies that got us to this place

This is pure "whataboutism". Both things can be bad.

I mean, you didn't address wealth inequality or global warning.

twister2920 13 hours ago | parent | prev | next [-]

have you considered that you're not the person this argument is trying to convince?

gruez 13 hours ago | parent [-]

So who is this trying to convince? All of this is bog-standard anti-AI argument that you can find on reddit. It's closer to preaching to the choir than actually trying to convince anyone.

OptionX 13 hours ago | parent | prev | next [-]

Such investment in technicalities to undermine the above post could itself speak of zealotry to a certain side in itself if one was prone to cynicism.

Not to mention that those nuances, as you state, mean very little. To say it was trained solely on copyrighted material or to say it was trained in large part on copyrighted material doesn't much move the needle over the fact it was done, a lot. Making a distinction between an "80 year old copyrighted work" or a recent one is also moot as it a matter of public record they themselves didn't make that distinction and happily fed on both kinds.

On to not being able to say they didn't ever do anything altruistic, well animal rights and welfare had great strides under the Nazi regime, not sure it evens out the rest. The same logic can be applied to the current state of affairs, doing a little good does not even out a lot of bad.

In closing, nuance is relevant when its relevant, after that its just being pedantic.

jzebedee 13 hours ago | parent | prev | next [-]

I think you can realize how ludicrous it is to invoke 80 year old copyrights for a product whose main use case is coding. How much code from 1946 was it lawfully ingesting?

rowanG077 13 hours ago | parent | prev | next [-]

Get real, you think LLMs got to where they are by training on 80 year old works? Your point is technically true, just basically irrelevant for LLM training.

uncivilized 12 hours ago | parent | prev | next [-]

If you really believe OpenAI, Anthropic, or any corporation for that matter has done something truly altruistic then I have a bridge to sell you.

12 hours ago | parent | prev | next [-]
[deleted]
zzzeek 12 hours ago | parent | prev | next [-]

there's persuasion, and then there's inspiring people already on your side. the upvotes/downvotes will give you a clue who the audience turns out to be. it was an upvote from me.

stego-tech 13 hours ago | parent | prev | next [-]

> If you enjoy the screaming, have at it. If you want to persuade people, dialing it down a few notches and at least giving token acknowledgement that nuance exists would help.

You’re not going sufficiently far back in comment histories (or using AI summaries of it from existing search tooling), or viewing the entire body of work I’ve written, shared, or discussed about all of those topics for the past twenty years.

I scream and wave my arms because nobody cared for decorum. I cut to the meat of the argument because nobody cares for context. Hell, this is the comments section for crying out loud; nobody is expecting a nuanced treatise about systemic interactions of multiple market and policy forces that resulted in this specific outcome in the comments section, even on HN.

Further, I suspect most folks here don’t actually care. I scream and wave my arms because it’s actually had the best results in outreach thus far - people see the flailing, read the bullet point, check the history, and somehow decide to reach out to me for more context (which I am always happy to give). When greeted with a wall of plaintext on an algorithmic site, the trained human response is to basically fuck off and go find a dopamine hit elsewhere, not stick around and read someone’s shoulda-been-a-blog-post on a social media comment page.

For all the crowing about “decorum” and “professionalism” and “intelligent debate”, the grim reality is that such concepts can actually detract from arguments by making it seem like we have time to discuss these further, that things aren’t that bad yet, that we can really think this through.

To which I say: no, we cannot. This language is a dire threat, it is a crimson red flag of intent to deliberately cause harm. You do not make an impassioned and polite argument about the nature of market forces or copyright policy with a mugger brandishing a gun, you do something about it.

aeon_ai 13 hours ago | parent | prev [-]

The latent space must have property boundaries as real and respected as those in the physical domain, lest we lose our ability to capitalize on them.

For now, and into the future, I hope that we protect that unique and novel idea of an anthropomorphic mouse, which should be expanded to cover all forms of anthropomorphism, if we're being honest with ourselves.

And, to that point, the protection of that idea should also extend to the (clearly infringing) hit 90s YA book series Animorphs, as it somewhat sounds like anthropomorphism.

Protect their rights! Anything produced within the conceptual acreage that has been clearly marked and claimed is theft!