Remix.run Logo
like_any_other an hour ago

> As mentioned elsewhere, this is also IP theft on a massive scale.

Scanning a legitimately purchased book is IP theft? How can he hold such a copyright-maximalist view, and at the same time defend the Internet Archive?

sebastiennight an hour ago | parent | next [-]

I imagine the argument could go like this.

Your granpa once wrote an obscure book about this amazing way he'd found to cure diabetes.

Corporation A buys all existing copies of the book, scans them, destroys the originals, and sets up a commercial business offering a monthly subscription to alleviate diabetes pains with this new method they claim they discovered.

Person B borrowed the book from a municipal library, Xerox'd it, and keeps a free ledger, open to all who want to read old books, as a way to safeguard free access to the world's knowledge.

Do you think there could exist any possible logic by which some people would defend person B and try to stop Corporation A?

like_any_other 42 minutes ago | parent [-]

In this hypothetical, my grandpa didn't patent this way to cure diabetes (otherwise the company would owe him royalties, assuming the patent hadn't expired).

So it reduces to a corporation selling services based on public-domain knowledge (I am ignoring the part where a miraculous advance is confined to a single unknown book). Lots of corporations do this, and there's nothing wrong with it.

I'm sure people would prefer if Amazon also offered library-like access to the book, but how is it theft? I'm not asking about "any possible logic" - he called it theft.

hattmall 32 minutes ago | parent | next [-]

Words are subjective. The person is saying that it is theft and if the law doesn't define it as such then the law should change. The implication is that taking from the public domain is theft if you then make it unavailable by destructive processes which render the public unable to access that which was previously available.

lmz a minute ago | parent | next [-]

The physical copy was in private ownership before, otherwise how can they sell it to Amazon?

sebastiennight 29 minutes ago | parent | prev [-]

> The implication is that taking from the public domain is theft if you then make it unavailable by destructive processes which render the public unable to access that which was previously available.

That is also my understanding of the claim, and I would intuitively agree with it.

voidhorse 29 minutes ago | parent | prev [-]

> So it reduces to a corporation selling services based on public-domain knowledge (I am ignoring the part where a miraculous advance is confined to a single unknown book). Lots of corporations do this, and there's nothing wrong with it.

Ok, but it's a bit different because they are not selling "the book" they are selling a fundamentally different thing for which the book is consumable input. Let's consider a batch of 5000 books. For kicks let's assume the known copies are N=1 for those books. Today, though only a few people can do this at a time, a person can purchase one of those books once read exactly the text in that book, and keep doing so to their heart's content. Or maybe a library buys it and now a rotating legion of people can do that for free. Now let's say amazon buys and scans them, destroying them in the process and refusing to release scans to avoid giving competitors a training edge. Now:

- You can never get the information as written again. At best you'll get an LLM output approximation/mutation of it. If this was the expression of a real human beings lived experience, that's kind of sad and goes against one of the spirited aspects of human existence, to leave a legacy, doesn't it? - Amazon is going to charge you per token every time you want to access that information. - Price demand for the information is now tied up with general demand for LLMs rather than the actual book, either reducing or greatly inflating the cost to you in addition to the now recurring charges.

So, now you (a) can't actually ever access that book as it was written (b) need to pay continually to access an approximation of its contents (c) and possibly more than the book is worth since now it's "value" in a price sense has been absorbed into general llm inference costs.

Not to mention you may have eradicated the last extant copy of a person's memoirs, but I guess it's pretty clear people in this industry don't care at this point. I hope one day your entire life story is ground up and consumed in some data farming operation and you are all summarily forgotten.

left-struck 31 minutes ago | parent | prev | next [-]

Scanning is not the ip theft, it never was.

kmeisthax 27 minutes ago | parent | prev | next [-]

They didn't say that the Internet Archive wasn't "IP theft" - nor did they say that IP theft was inherently wrong.

"AI training is infringement" is not exactly a copyright-maximalist view. The explicit training task used for pre-training is reproducing the content of the trained-on books; and models trained on such books are able to reproduce significant infringing chunks of them[0] unless specifically post-trained to refuse to do so.

Additionally, they might have thought that Controlled Digital Lending was OK (it wasn't, but that's a different issue to AI training). As I've mentioned elsewhere in this thread, there's a common misconception that copyright is concerned with the number of copies in circulation as opposed to individual acts of copying.

Or they don't care about any of that and just wanted to highlight the hypocrisy.

[0] Which, under the "compression is intelligence" point of view, is entirely expected and not surprising in the slightest.

noosphr 38 minutes ago | parent | prev [-]

AI psychosis cuts both ways.