Remix.run Logo
magarnicle 13 hours ago

Why would reading copyrighted material ever be an issue anyway? Wouldn't copyright law only apply to what you create and publish using the model? Training on every comic book should already be perfectly legal, as long as you accessed them legally, right? But publishing your own Batman comic using that training is copyright infringement.

What I'm saying is, doesn't the law already cover 1?

_aavaa_ 12 hours ago | parent [-]

Fair use requires more than you accessing the material legally.

In the US one of the factors is “ the effect of the use upon the potential market for or value of the copyrighted work”.

If anthropic Hoovers up the world’s books and trains on them, and then spits them out verbatim on command, then it will clearly impact the value of the work; nobody will buy the original, they’ll just ask Claude.

Others also argue that even if it’s not reproducing it exactly that the training runs afoul of that factor, specifically the “market for” portion. A rights holder can no longer license their book for training of LLMs if Anthropic goes ahead and just trains on it anyway.

magarnicle 11 hours ago | parent [-]

> If anthropic Hoovers up the world’s books and trains on them, and then spits them out verbatim on command, then it will clearly impact the value of the work; nobody will buy the original, they’ll just ask Claude.

Ah, right. So if we want models to be capable we need them to be trained on as much as possible, yet we also want to stop what you described. So what can be done?

_aavaa_ 3 hours ago | parent [-]

I mean the choice is: 1) we pass laws that explicitly say training models like this is legal (the original quote, 2) say it’s illegal and requires licenses for the data and ability to opt out, 3) we ignore it and continue because the companies are too big to jail.