Remix.run Logo
card_zero 19 hours ago

Quite possibly not many, and no copy held in any form by the copyright owner either. Say a few hundred copies of some obscure book from 40 years ago. They probably won't be erased from the face of the earth by the judicious and proportionate actions of, of a few, AI companies? Hmm.

scarmig 19 hours ago | parent | next [-]

The hypothetical "heroic figure goes and buys last copy of a 1962 guide to Ford cars to carefully maintain it in an appropriately climate controlled library" is vanishingly unlikely. A ten or a hundred or a thousand times to one, it just goes to the trash. At least here it gets scanned by the AI company.

asaddhamani 18 hours ago | parent | next [-]

But that scan is never made available to us in its original form. So it getting scanned by the AI company does nothing to preserve the book.

scarmig 18 hours ago | parent | next [-]

Dumpsters also don't typically come equipped with a robot scanner and network uplink built in.

Like, I really don't know what people objecting to this imagine typically happens to old, unwanted books. They don't get sent to some magical library in the countryside if unpurchased where they are carefully maintained forever (next to where Rover spends the rest of his days). They are very literally thrown into the trash.

That said, I'd be thrilled if the US government required AI companies to make them available to the public. I'd even settle for the US government making it legal for them to.

fmajid 16 hours ago | parent | next [-]

The Internet Archive tries to be that magical library, but they can only scan and physically archive what is sent to them.

ralferoo 14 hours ago | parent | prev | next [-]

> I really don't know what people objecting to this imagine typically happens to old, unwanted books.

In the UK at least, people usually take them to a second hand / charity shop, who sort through them and send the valuable ones to auction (typically early editions, 100+ years old) and then either sell them themselves (for recent books that are easy to get rid of) or sell them to specialised second-hand bookshops.

Most of the specialised second-hand bookshops rarely throw books away, usually if nobody buys them after a couple of years they end up in the extreme discount piles (20p, 50p etc) and probably only trashed if they still don't sell from there.

theshrike79 14 hours ago | parent [-]

So trashed, but with a bunch of extra steps then?

tentacleuno 3 hours ago | parent [-]

I would presume that the extra steps incrementally diminish the possibility of the book remaining unsold, and thus destroyed or sent elsewhere.

theshrike79 2 hours ago | parent [-]

And it also adds costs in every step. Someone needs to move thousands of unwanted books from high end stores to lower and lower end stores. Someone needs to store them in the proper environment etc.

I do get the _idea_ of preserving books, but... people don't care. I just threw out well over a thousand books from my grandparents house this spring.

There were ~6-10 "valuable" books there. Two because I personally knew someone who wanted old war-time books and a few 100+ year old bibles. And maybe two dozen books worth saving, mostly because they were from big-name authors or had stuff that nobody would print anymore (I have detailed instructions how to make laughing gas and how to build an underground chemical lab - hobby books in the 50s were ... interesting :D )

I literally couldn't give away the rest. And I tried. It was all just "interesting, but..." - no way to justify using the shelf space for books that, realistically, nobody will actually ever read again.

soperj 18 hours ago | parent | prev [-]

They're buying the books from resellers, not rescuing these books from dumpsters. Stop being an apologist.

skeledrew 17 hours ago | parent | next [-]

Dumpster is where they go when the resellers fail to complete sales.

scarmig 18 hours ago | parent | prev [-]

The magical library in the countryside, to be painfully explicit, does not exist.

Ekaros 17 hours ago | parent | next [-]

And they should not even be needed. In many places the issue is solved at start. Copy or copies of each commercially produced book is send to national library. Which with tax payer money keeps an archive. Meaning that at least one copy exist for research purposes if needed.

scarmig 17 hours ago | parent [-]

Unfortunately, that's not the case in the United States. The LOC only selects around half of published books to be permanently held. The rest are disposed of (usually returning them to the publisher, donating them to a library, or destroying them).

fmajid 16 hours ago | parent [-]

They should send them to The Internet Archive instead.

FeloniousHam 9 hours ago | parent [-]

Why aren't we storming the Library of Congress? They are the real villains here.

soperj 17 hours ago | parent | prev | next [-]

How many times can you post the same thing in a thread?

exe34 16 hours ago | parent | prev [-]

The internet archive

dukeyukey 17 hours ago | parent | prev | next [-]

If it were legal they may well do that as a public branding exercise. Google already tried and got punished for it!

skeledrew 17 hours ago | parent | prev | next [-]

It was never available to you/us in the original form either.

red75prime 17 hours ago | parent | prev [-]

...because it is illegal to copy copyrighted material. 70 years later they might do it.

rhdunn 17 hours ago | parent | next [-]

95 years after publication. Many other countries also have an X years after the author's death clause where X varies between countries but is at least 70.

There are also other weird issues such as the UK having a clause protecting Peter Pan (so a children's hospital gets royalties) and the King James translation of the bible (under Crown copyright) that extend the copyright even further.

In short, it's a mess.

mrweasel 16 hours ago | parent | prev [-]

The thing I find most hypocritical though is that they are probably never share their libraries with anyone. After scanning, downloading, stealing, overloading websites and everything in between, to acquire enough data for their stupid machine, they're not going to share their data? I get that most of it can't be shared, but a lot can. There's no reason why you need to destroy multiple copies of a book from 1880, when it's free to share.

At the same time I can understand keeping track of when each books enters public domain might also be an absolute nightmare, and I wouldn't blame the AI companies for not wanting to deal with that. For the stuff they absolutely know is clear, they should provide dumps for everyone to download.

novok 16 hours ago | parent | next [-]

This is solved by law, which is solved by 'we the people' and I bet many AI companies would be fine with something like the equivalent to patent law with bankruptcy escrow to the library of congress, where they must release the scans in 10 years for books that the vast majority will not give a flying shit about. By then the advantage is long gone in data moat.

fmajid 16 hours ago | parent | prev [-]

Since they seem to leapfrog each others’ models every few months, the training data is one of the few ways they can build competitive advantage, and that explains why they don’t share, even if we don’t have to like this.

svachalek 9 hours ago | parent [-]

Would it even be legal to share? I don't think it would be.

mejutoco 14 hours ago | parent | prev [-]

In my opinion this is one of the reasons why libraries should accept any book, even if all they do is examine it and throw it in the trash. This way they would have a chance at finding any treasures that could be regularly dumped in that way.

bulbar 19 hours ago | parent | prev | next [-]

> They probably won't be erased from the face of the earth by the judicious and proportionate actions of, of a few, AI companies?

I don't see why not. Pretty sure it's gonna happen. Doesn't matter if a hundred copies still exist somewhere, if access or discoverbility falls below a certain threshold, it doesn't matter, because those books become practically inaccessible to the world.

margalabargala 19 hours ago | parent | next [-]

Right, but if an AI company buys some vanishingly uncommon book, digitizes it, shreds it, and adds the information it contains to their permanent digital library and digests its contents into an AI that is then publicly accessible...are they making that book less accessible, or more?

scarmig 18 hours ago | parent | next [-]

You've got to compare it to the alternative. Books have a half-life, and the vast majority of these books being purchased are grody, moldering ex-lib copies of books that no one has read in decades. Their other likely outcome is mulching.

margalabargala 18 hours ago | parent [-]

Right, that's my point.

These generally are not books people care about. The information contained therein was doomed.

Now the information has been digitally preserved and a digestion of the information will be made publicly available.

michaelmrose 17 hours ago | parent [-]

[dead]

halsafar 18 hours ago | parent | prev | next [-]

Can you get the exact text back out with a prompt or not? Having or not having a book isn't fuzzy.

skeledrew 17 hours ago | parent | next [-]

Funnily the argument made just a few months ago by many rights holders who wanted their pound of flesh was that, if prompted a certain way, exact text could be retrieved.

margalabargala 18 hours ago | parent | prev | next [-]

Having or not having a book is absolutely fuzzy. If you have a translation, do you have the book? Even if, like the Odyssey, there are hundreds of wildly varying translations? What about an abridged copy? What about the Sparknotes version? If you have a copy of Pride And Prejudice And Zombies, do you have a copy of Pride And Prejudice? Certainly more so than if you have neither.

fluoridation 16 hours ago | parent [-]

>If you have a translation, do you have the book?

No, you have a translation.

>Even if, like the Odyssey, there are hundreds of wildly varying translations?

Precisely why translations are not considered equivalent to the original text.

>What about an abridged copy? What about the Sparknotes version?

An abridged copy is not a copy of the unabridged version.

>If you have a copy of Pride And Prejudice And Zombies, do you have a copy of Pride And Prejudice?

No.

I'm honestly surprised these were the questions you chose to ask, when you could have asked what if you have 90% of the pages, or what if most of the pages are missing pieces because the book was shot with a shotgun, or what if the book was scanned and OCRed and all the "rn"s were replaced with "m"s and all the lower case Ls with ones. Hell, is a scan of the book close enough to having the book, or is it far enough that one can no longer be said to have the book anymore?

margalabargala 9 hours ago | parent [-]

My opinion is different from yours.

If I have a translation of a book, I think I have more of that book than if I had nothing at all. It's fuzzy.

fluoridation 9 hours ago | parent [-]

You don't have the book, you have someone else's interpretation of the book's contents, re-expressed into a language you can read. Both steps can involve a loss or distortion of information, either because the translator doesn't fully grasp the original language or context, or because in the re-expression they chose to leave out details that were relevant to you. The more distant the original language to yours, the more translation involves interpretation, too. A translation is really not too different from a commentary; it's just a different text.

margalabargala 8 hours ago | parent [-]

That sounds to me like fuzzily having a version of the book. That is, it's more like having the book, than having nothing would be.

You seem to be arguing that a translation, etc is not literally having the book, which is something that has always been my stance as well.

fluoridation 8 hours ago | parent [-]

You're trying to use the Socratic method to show that the havingness of the book is a spectrum, and I'm taking the position of a hardliner who considers that "having the book" means having the original string of symbols from beginning to end, and anything besides that is not having the book. I'm trying to show you that your line of argumentation is uncompelling to such a person.

margalabargala 8 hours ago | parent [-]

I was never using the Socratic method. I was asking rhetorical questions, and then gave the (my) answer at the end.

We just have different opinions about what it means to have a book. I think that having a copy of "Pride and Prejudice and Zombies" is more like having a copy of "Pride and Prejudice", than having no book at all is like having a copy of that book. You disagree and that's fine.

fluoridation 7 hours ago | parent [-]

Asking rhetorical questions as a form of argumentation is the Socratic method.

margalabargala 7 hours ago | parent [-]

If one is using the Socratic method, rhetorical questions are one tool they might employ.

The reverse is not true. Just because someone asks a rhetorical question, does not mean they are using the Socratic method.

close04 14 hours ago | parent | prev [-]

The scan and destroy method is what a judge allowed to do in order to have a copy of the book in the training dataset. With the physical copy destroyed there's still only 1 copy in circulation. Once "inside" an LLM I don't know if anyone decided unequivocally that it's copyright infringement or not, and if that counts as a second copy.

There's no technical reason why an LLM couldn't reproduce verbatim some of the training material. It's sort of a lossy statistical compression engine. Enough of the info will survive to the output in the original form. With the amount of data and the commercial nature it's hard to argue fair-use. But nobody tested this in court. I'm not even sure the US wants to ever test this. Why even attempt something that has a non-0 chance to sabotage your most promising industry/bubble in ages?

GPerson 18 hours ago | parent | prev | next [-]

They’re not supposed to be storing a copy. What they’re doing is destroying their copy after training a model on it.

derektank 18 hours ago | parent | next [-]

No, US copyright law allows them to keep a single digital copy. The hypothetical issue is with them maintaining two copies, one digital and one physical, when they only purchased one.

monocasa 18 hours ago | parent | prev [-]

They're absolutely keeping the digitized copies. They're not going to just train a single model.

Natsu 18 hours ago | parent | prev [-]

AIs are weirdly bad at quoting stuff in my experience.

But you'd think that the Library of Congress and such would actually prevent stuff from vanishing just by collecting it themselves.

margalabargala 18 hours ago | parent [-]

Bad at quoting, good at digesting and regurgitating. The concepts are preserved even if quotes aren't.

I'd rather a digital copy exist in someone's hands than a rotting physical copy.

subscribed 16 hours ago | parent [-]

But you don't have access to this copy. The digital copy is removed and all you get is paraphrased content. Some frontier models have been explicitly forbidden from recalling exact quotes in system prompt.

It almost seems like you're suggesting that having Claude generate a paraphrased book is as good as having the original book but i don't think that could be your intention?

sharpshadow 18 hours ago | parent | prev [-]

On a similar topic are all those artifacts kept in museum storages for literally eternity. Maybe AI money can crack open access to it.

qingcharles 17 hours ago | parent | prev | next [-]

Many are just copyright "orphans", nobody knows who owns the copyright any longer. Maybe the author died and the copyright passed to their estate, but they're not even aware of it.

One book I'm hunting for a copy of right now was published in England in 1947 and in those days paper was rationed, so not many copies were made, and only a handful have survived. As soon as I find it I'll scan it and upload it to IA.

willy_k 19 hours ago | parent | prev [-]

Is there a specific book from 40 years ago you have in mind? Asking out of curiosity.

ipaddr 18 hours ago | parent [-]

Books by Zolar are interesting hard to find all editions. The Fearful Void by Geoffrey Moorhouse probably still has 100s of copies available but hate to see it lost.