Remix.run Logo
ezfe 20 hours ago

I dislike these AI companies but let's be clear here: the copyright holders are the ones locking these books up. If they don't want to print more copies, then they could release the copyright on them.

Instead, they enforce the copyright and force AI companies to shred books they want to ingest.

edit: Also, an AI company would only ever care to purchase, scan, destroy a book once. Presumably many books have more than one copy.

RajT88 20 hours ago | parent | next [-]

The articles I've read on this are not clear, but I strongly suspect "rare" is not the definition you and I probably use for the level of rarity of books actually being destroyed.

These are not going to be the kinds of books "The Ninth Gate" resolved around - truly one of a kind. It's not good they are destroying books, but they are books which do have other copies. Just perhaps not many.

card_zero 19 hours ago | parent | next [-]

Quite possibly not many, and no copy held in any form by the copyright owner either. Say a few hundred copies of some obscure book from 40 years ago. They probably won't be erased from the face of the earth by the judicious and proportionate actions of, of a few, AI companies? Hmm.

scarmig 19 hours ago | parent | next [-]

The hypothetical "heroic figure goes and buys last copy of a 1962 guide to Ford cars to carefully maintain it in an appropriately climate controlled library" is vanishingly unlikely. A ten or a hundred or a thousand times to one, it just goes to the trash. At least here it gets scanned by the AI company.

asaddhamani 18 hours ago | parent | next [-]

But that scan is never made available to us in its original form. So it getting scanned by the AI company does nothing to preserve the book.

scarmig 18 hours ago | parent | next [-]

Dumpsters also don't typically come equipped with a robot scanner and network uplink built in.

Like, I really don't know what people objecting to this imagine typically happens to old, unwanted books. They don't get sent to some magical library in the countryside if unpurchased where they are carefully maintained forever (next to where Rover spends the rest of his days). They are very literally thrown into the trash.

That said, I'd be thrilled if the US government required AI companies to make them available to the public. I'd even settle for the US government making it legal for them to.

fmajid 16 hours ago | parent | next [-]

The Internet Archive tries to be that magical library, but they can only scan and physically archive what is sent to them.

ralferoo 14 hours ago | parent | prev | next [-]

> I really don't know what people objecting to this imagine typically happens to old, unwanted books.

In the UK at least, people usually take them to a second hand / charity shop, who sort through them and send the valuable ones to auction (typically early editions, 100+ years old) and then either sell them themselves (for recent books that are easy to get rid of) or sell them to specialised second-hand bookshops.

Most of the specialised second-hand bookshops rarely throw books away, usually if nobody buys them after a couple of years they end up in the extreme discount piles (20p, 50p etc) and probably only trashed if they still don't sell from there.

theshrike79 14 hours ago | parent [-]

So trashed, but with a bunch of extra steps then?

tentacleuno 3 hours ago | parent [-]

I would presume that the extra steps incrementally diminish the possibility of the book remaining unsold, and thus destroyed or sent elsewhere.

theshrike79 2 hours ago | parent [-]

And it also adds costs in every step. Someone needs to move thousands of unwanted books from high end stores to lower and lower end stores. Someone needs to store them in the proper environment etc.

I do get the _idea_ of preserving books, but... people don't care. I just threw out well over a thousand books from my grandparents house this spring.

There were ~6-10 "valuable" books there. Two because I personally knew someone who wanted old war-time books and a few 100+ year old bibles. And maybe two dozen books worth saving, mostly because they were from big-name authors or had stuff that nobody would print anymore (I have detailed instructions how to make laughing gas and how to build an underground chemical lab - hobby books in the 50s were ... interesting :D )

I literally couldn't give away the rest. And I tried. It was all just "interesting, but..." - no way to justify using the shelf space for books that, realistically, nobody will actually ever read again.

soperj 18 hours ago | parent | prev [-]

They're buying the books from resellers, not rescuing these books from dumpsters. Stop being an apologist.

skeledrew 17 hours ago | parent | next [-]

Dumpster is where they go when the resellers fail to complete sales.

scarmig 18 hours ago | parent | prev [-]

The magical library in the countryside, to be painfully explicit, does not exist.

Ekaros 17 hours ago | parent | next [-]

And they should not even be needed. In many places the issue is solved at start. Copy or copies of each commercially produced book is send to national library. Which with tax payer money keeps an archive. Meaning that at least one copy exist for research purposes if needed.

scarmig 17 hours ago | parent [-]

Unfortunately, that's not the case in the United States. The LOC only selects around half of published books to be permanently held. The rest are disposed of (usually returning them to the publisher, donating them to a library, or destroying them).

fmajid 16 hours ago | parent [-]

They should send them to The Internet Archive instead.

FeloniousHam 9 hours ago | parent [-]

Why aren't we storming the Library of Congress? They are the real villains here.

soperj 17 hours ago | parent | prev | next [-]

How many times can you post the same thing in a thread?

exe34 16 hours ago | parent | prev [-]

The internet archive

dukeyukey 17 hours ago | parent | prev | next [-]

If it were legal they may well do that as a public branding exercise. Google already tried and got punished for it!

skeledrew 17 hours ago | parent | prev | next [-]

It was never available to you/us in the original form either.

red75prime 17 hours ago | parent | prev [-]

...because it is illegal to copy copyrighted material. 70 years later they might do it.

rhdunn 17 hours ago | parent | next [-]

95 years after publication. Many other countries also have an X years after the author's death clause where X varies between countries but is at least 70.

There are also other weird issues such as the UK having a clause protecting Peter Pan (so a children's hospital gets royalties) and the King James translation of the bible (under Crown copyright) that extend the copyright even further.

In short, it's a mess.

mrweasel 16 hours ago | parent | prev [-]

The thing I find most hypocritical though is that they are probably never share their libraries with anyone. After scanning, downloading, stealing, overloading websites and everything in between, to acquire enough data for their stupid machine, they're not going to share their data? I get that most of it can't be shared, but a lot can. There's no reason why you need to destroy multiple copies of a book from 1880, when it's free to share.

At the same time I can understand keeping track of when each books enters public domain might also be an absolute nightmare, and I wouldn't blame the AI companies for not wanting to deal with that. For the stuff they absolutely know is clear, they should provide dumps for everyone to download.

novok 16 hours ago | parent | next [-]

This is solved by law, which is solved by 'we the people' and I bet many AI companies would be fine with something like the equivalent to patent law with bankruptcy escrow to the library of congress, where they must release the scans in 10 years for books that the vast majority will not give a flying shit about. By then the advantage is long gone in data moat.

fmajid 16 hours ago | parent | prev [-]

Since they seem to leapfrog each others’ models every few months, the training data is one of the few ways they can build competitive advantage, and that explains why they don’t share, even if we don’t have to like this.

svachalek 9 hours ago | parent [-]

Would it even be legal to share? I don't think it would be.

mejutoco 14 hours ago | parent | prev [-]

In my opinion this is one of the reasons why libraries should accept any book, even if all they do is examine it and throw it in the trash. This way they would have a chance at finding any treasures that could be regularly dumped in that way.

bulbar 19 hours ago | parent | prev | next [-]

> They probably won't be erased from the face of the earth by the judicious and proportionate actions of, of a few, AI companies?

I don't see why not. Pretty sure it's gonna happen. Doesn't matter if a hundred copies still exist somewhere, if access or discoverbility falls below a certain threshold, it doesn't matter, because those books become practically inaccessible to the world.

margalabargala 19 hours ago | parent | next [-]

Right, but if an AI company buys some vanishingly uncommon book, digitizes it, shreds it, and adds the information it contains to their permanent digital library and digests its contents into an AI that is then publicly accessible...are they making that book less accessible, or more?

scarmig 18 hours ago | parent | next [-]

You've got to compare it to the alternative. Books have a half-life, and the vast majority of these books being purchased are grody, moldering ex-lib copies of books that no one has read in decades. Their other likely outcome is mulching.

margalabargala 18 hours ago | parent [-]

Right, that's my point.

These generally are not books people care about. The information contained therein was doomed.

Now the information has been digitally preserved and a digestion of the information will be made publicly available.

michaelmrose 17 hours ago | parent [-]

[dead]

halsafar 18 hours ago | parent | prev | next [-]

Can you get the exact text back out with a prompt or not? Having or not having a book isn't fuzzy.

skeledrew 17 hours ago | parent | next [-]

Funnily the argument made just a few months ago by many rights holders who wanted their pound of flesh was that, if prompted a certain way, exact text could be retrieved.

margalabargala 18 hours ago | parent | prev | next [-]

Having or not having a book is absolutely fuzzy. If you have a translation, do you have the book? Even if, like the Odyssey, there are hundreds of wildly varying translations? What about an abridged copy? What about the Sparknotes version? If you have a copy of Pride And Prejudice And Zombies, do you have a copy of Pride And Prejudice? Certainly more so than if you have neither.

fluoridation 16 hours ago | parent [-]

>If you have a translation, do you have the book?

No, you have a translation.

>Even if, like the Odyssey, there are hundreds of wildly varying translations?

Precisely why translations are not considered equivalent to the original text.

>What about an abridged copy? What about the Sparknotes version?

An abridged copy is not a copy of the unabridged version.

>If you have a copy of Pride And Prejudice And Zombies, do you have a copy of Pride And Prejudice?

No.

I'm honestly surprised these were the questions you chose to ask, when you could have asked what if you have 90% of the pages, or what if most of the pages are missing pieces because the book was shot with a shotgun, or what if the book was scanned and OCRed and all the "rn"s were replaced with "m"s and all the lower case Ls with ones. Hell, is a scan of the book close enough to having the book, or is it far enough that one can no longer be said to have the book anymore?

margalabargala 9 hours ago | parent [-]

My opinion is different from yours.

If I have a translation of a book, I think I have more of that book than if I had nothing at all. It's fuzzy.

fluoridation 9 hours ago | parent [-]

You don't have the book, you have someone else's interpretation of the book's contents, re-expressed into a language you can read. Both steps can involve a loss or distortion of information, either because the translator doesn't fully grasp the original language or context, or because in the re-expression they chose to leave out details that were relevant to you. The more distant the original language to yours, the more translation involves interpretation, too. A translation is really not too different from a commentary; it's just a different text.

margalabargala 8 hours ago | parent [-]

That sounds to me like fuzzily having a version of the book. That is, it's more like having the book, than having nothing would be.

You seem to be arguing that a translation, etc is not literally having the book, which is something that has always been my stance as well.

fluoridation 8 hours ago | parent [-]

You're trying to use the Socratic method to show that the havingness of the book is a spectrum, and I'm taking the position of a hardliner who considers that "having the book" means having the original string of symbols from beginning to end, and anything besides that is not having the book. I'm trying to show you that your line of argumentation is uncompelling to such a person.

margalabargala 8 hours ago | parent [-]

I was never using the Socratic method. I was asking rhetorical questions, and then gave the (my) answer at the end.

We just have different opinions about what it means to have a book. I think that having a copy of "Pride and Prejudice and Zombies" is more like having a copy of "Pride and Prejudice", than having no book at all is like having a copy of that book. You disagree and that's fine.

fluoridation 7 hours ago | parent [-]

Asking rhetorical questions as a form of argumentation is the Socratic method.

margalabargala 7 hours ago | parent [-]

If one is using the Socratic method, rhetorical questions are one tool they might employ.

The reverse is not true. Just because someone asks a rhetorical question, does not mean they are using the Socratic method.

close04 14 hours ago | parent | prev [-]

The scan and destroy method is what a judge allowed to do in order to have a copy of the book in the training dataset. With the physical copy destroyed there's still only 1 copy in circulation. Once "inside" an LLM I don't know if anyone decided unequivocally that it's copyright infringement or not, and if that counts as a second copy.

There's no technical reason why an LLM couldn't reproduce verbatim some of the training material. It's sort of a lossy statistical compression engine. Enough of the info will survive to the output in the original form. With the amount of data and the commercial nature it's hard to argue fair-use. But nobody tested this in court. I'm not even sure the US wants to ever test this. Why even attempt something that has a non-0 chance to sabotage your most promising industry/bubble in ages?

GPerson 18 hours ago | parent | prev | next [-]

They’re not supposed to be storing a copy. What they’re doing is destroying their copy after training a model on it.

derektank 18 hours ago | parent | next [-]

No, US copyright law allows them to keep a single digital copy. The hypothetical issue is with them maintaining two copies, one digital and one physical, when they only purchased one.

monocasa 18 hours ago | parent | prev [-]

They're absolutely keeping the digitized copies. They're not going to just train a single model.

Natsu 18 hours ago | parent | prev [-]

AIs are weirdly bad at quoting stuff in my experience.

But you'd think that the Library of Congress and such would actually prevent stuff from vanishing just by collecting it themselves.

margalabargala 18 hours ago | parent [-]

Bad at quoting, good at digesting and regurgitating. The concepts are preserved even if quotes aren't.

I'd rather a digital copy exist in someone's hands than a rotting physical copy.

subscribed 16 hours ago | parent [-]

But you don't have access to this copy. The digital copy is removed and all you get is paraphrased content. Some frontier models have been explicitly forbidden from recalling exact quotes in system prompt.

It almost seems like you're suggesting that having Claude generate a paraphrased book is as good as having the original book but i don't think that could be your intention?

sharpshadow 18 hours ago | parent | prev [-]

On a similar topic are all those artifacts kept in museum storages for literally eternity. Maybe AI money can crack open access to it.

qingcharles 17 hours ago | parent | prev | next [-]

Many are just copyright "orphans", nobody knows who owns the copyright any longer. Maybe the author died and the copyright passed to their estate, but they're not even aware of it.

One book I'm hunting for a copy of right now was published in England in 1947 and in those days paper was rationed, so not many copies were made, and only a handful have survived. As soon as I find it I'll scan it and upload it to IA.

willy_k 19 hours ago | parent | prev [-]

Is there a specific book from 40 years ago you have in mind? Asking out of curiosity.

ipaddr 18 hours ago | parent [-]

Books by Zolar are interesting hard to find all editions. The Fearful Void by Geoffrey Moorhouse probably still has 100s of copies available but hate to see it lost.

ErigmolCt 16 hours ago | parent | prev | next [-]

I suspect rare here often means out of print or commercially obscure, not unique

alightsoul 19 hours ago | parent | prev | next [-]

There's just a few copies in a single library worldwide which is probably a national or a university library

tptacek 19 hours ago | parent [-]

The 404 story suggested that these are largely vanity press books and instruction manuals for things no longer sold. Implying that these books were almost certainly headed for the recycling center had the AI companies not snatched them up.

alightsoul 19 hours ago | parent [-]

They are useful as a source of non synthetic data, replacing synthetic data consisting of rephrasings of common topics I assume?

runarberg 20 hours ago | parent | prev | next [-]

At this scale, there are no guarantees of anything. There very likely will be unique copies in there. If these were expert archivists a lot of damage could be prevented, but given the malice and indifference of AI companies, there very likely will not be an expert archivist involved, and unique copies will be destroyed unceremoniously.

eru 19 hours ago | parent [-]

My personal wastebook at home is so rare, it's unique. That doesn't mean it needs preservation.

card_zero 19 hours ago | parent | next [-]

Your what now? Made from your personal waste? That does sound unique.

https://en.wiktionary.org/wiki/wastebook

Oh right. But anyway, nobody knows what needs preservation, it's a basic problem of life, somebody usually mentions the BBC throwing out boring old Doctor Who tapes to save archive space because nobody liked it any more at that point in time. Some things should probably be thrown out now and then, I suppose.

eru 19 hours ago | parent [-]

The alternative for many of these un(der)appreciated books is that they will get unceremoniously dumped in the future anyway. The publishing industry and libraries etc dispose off lots and lots of books.

So at least with the AI companies they are scanning them and preserving them digitally. Not just in the trained weights, but also as raw training data for future runs.

P.S. I'm not sure why you need to make fun of your own ignorance? Just look up the word you don't know and don't mention it?

fmajid 16 hours ago | parent [-]

The AI companies’ working assumption is that if someone found it worth printing, it has enough information content to help train a model. That assumption might be invalid with some of the more rambling self-published books, however.

asdfsa32 19 hours ago | parent | prev | next [-]

https://en.wikipedia.org/wiki/Anecdotal_evidence

eru 14 hours ago | parent [-]

You are giving me too much credit: it's made up evidence. It's an illustration that rarity doesn't equal value, and doesn't depend on whether I actually own a wastebook or ten.

asdfsa32 14 hours ago | parent [-]

Your point is understood but the crux of the issue is a bit similar to capital punishment, the argument is that risk of losing even one innocent person or useful book isn't worth taking; specially considering the value created for society as part of such risky undertaking, whatever it is exercising capital punishment or scanning and destroying books.

eru 14 hours ago | parent [-]

Whenever you build a highway or a bridge or a power plant, your engineers have to put a money value on human life, or at least the worth of a statistical human life. Just to make ordinary engineering decisions.

And refusing to do this exercise just means that you behave as-if you put a really silly number on the value of human life, and probably not consistent between different parts of the project.

So I don't quite agree with these taboos in the absolute.

(I'm still against capital punishment on practical grounds.)

For books it's similar: if you taboo book destruction for the AI training folks, that doesn't rescue books from their ordinary pre-AI life cycle of getting destroyed all the time in the course of running a publisher or a library or a second-hand book store.

In fact, the AI craze is what's giving rare books _value_ and incentivises people to dig them up and preserve them. Or at least preserve them long enough to be scanned.

The scanning might destroy the physical copy of that book, but they save the contents. That's the whole point of scanning after all.

asdfsa32 12 hours ago | parent [-]

I am well aware of the Statistical Value of Life and how it impacts projects. But that is generally used for allocating preventative measures. No road is going to get approved if it is going to result in random deaths by design.

runarberg 19 hours ago | parent | prev [-]

Like I said, at this scale, there are no guarantees for anything. Very likely will there be a unique copy of an invaluable book or letter an the person feeding it to the scanner will not know and the book get destroyed.

Like did Icelandic author Þórbergur Þórðarson ever write an a book about Esperanto, and send the only copy of it to Halldór Laxness when he was in Los Angeles? I don‘t know, but it is certainly something he is likely to have done. If such a book exists it would be invaluable to both Icelandic culture and to Esprentists. It likely would have stayed in Los Angeles where nobody would know the significance of it until it ended up in an estate sale, a used book store, and then finally destroyed by an AI company never to be discovered.

My hypothetical is just one of trillions of possibilities. At this scale very likely several of these possibilities will unessiseraly remain unknown unknowns forever.

eru 19 hours ago | parent [-]

Well, at least afterwards the book is scanned and preserved digitally in their archives of training data.

If the book was just rotting away in some forgotten bookstore, it would more likely be unceremoniously disposed off in the future without anyone scanning it first.

enraged_camel 19 hours ago | parent | prev [-]

Also, a lot of these "rare books" are stuff like TV programming magazines from October 1994.

card_zero 19 hours ago | parent | next [-]

So, have you tried finding out what the programming was in October 1994? Or what cultural ephemera appeared in the TV guides of that era alongside the schedules? Either there's a copy for the week you want in an archive, or somebody's got one for sale, or most often neither. This can piss you off, if as it happened you had a reason to care.

tptacek 19 hours ago | parent | next [-]

That would make sense as an argument if the natural endpoint of these copies was preservation, and AI was disrupting that. But the natural next step for virtually all these books is to be recycled, not preserved. Books are generally not preserved. It is extremely normal for them to be pulped. Millions and millions of books are pulped every year.

mslt 18 hours ago | parent | next [-]

Might as well grind up the tablet of complaint to Ea-nāṣir and make cement out it, right? What use could there be in preserving the mundane facets of everyday existence?

svachalek 9 hours ago | parent | next [-]

Preservation would be cool but what does that have to do with anything? The system is that the remaining copies have private owners and the private owners can do anything they want with them, including sell them or shred them.

Most of these owners aren't doing anything particular to preserve them, they're stacked up with 10,000 other TV Guides in a hoarder's moldy basement. Anyone interested in keeping October 1994's TV Guide pristine has had over 30 years to procure and protect their copy.

These particular buyers are converting their copy to a digital one that's getting some kind of use, which is better than the fate of 99.9% of the other copies.

tptacek 6 hours ago | parent [-]

I don't even think ownership rights are the principal component of this situation. It's true that in a different copyright regime, Anthropic could simply publish an archive of every book they scanned (and I think they likely would do that, if they could). But the primary reason these books get destroyed is that nobody cares about them. Old, rare books are a garbage disposal problem, not a cultural one. To me the most important thing to know about this story is that the AI companies are a tiny, tiny blip in the big picture of what happens to old physical books.

Everybody imagines libraries hoping against hope to get their hands on all these rare books so they can shelve them and preserve them for generations to come. But if you donate a box of old books to a public library, there's a good chance the clerical staff handling that donation is going to roll its eyes and cart the books out to a dumpster. Large public library systems have stopped taking book donations, for this reason. "Bring them to a thrift store" is what they'll tell you.

dukeyukey 17 hours ago | parent | prev | next [-]

It's more that, we don't need to preserve a million copies of the same mundane book. Losing a few copies to AI training is fine.

13 hours ago | parent [-]
[deleted]
pfdietz 9 hours ago | parent | prev [-]

When applied to ordinary mundane objects, this is the mindset that leads to pathological hording behavior. Ultimately I think it's rooted in a fear of, an attempt to deny, mortality and the passage of time.

pfdietz 19 hours ago | parent | prev | next [-]

Around a million books are destroyed each day in the US.

19 hours ago | parent | prev [-]
[deleted]
mslt 18 hours ago | parent | prev | next [-]

To play devils advocate, completely on the terms of your argument, would it be better for that particular human artifact to be shredded and its contents melted into an anonymized data pool, or for it to exist in a museum archive, in its original form, such that future generations can better understand what it was like to be alive in 1994?

I’d personally choose the latter, especially given that the 1994 tv guide is not going to meaningfully improve the utility of the language models.

Direct access to pre-digital history is drying up rapidly, why accelerate that for incremental benchmark gains in a domain that isn’t even relevant to the most useful forms of a nascent technology?

tgsovlerkhgsel 8 hours ago | parent | next [-]

A reasonable opinion, but I'd personally strongly choose the former.

A physical book in a museum archive is useless for 99% of the worlds population even if they really wanted that specific book and were able to find it, as they'd have to arrange for access, then travel (at incredible expense) to access it.

Maybe they could ask the museum to digitize it, but that's still going to be days of delay and tens of dollars of cost to access parts of that book, if the museum even offers that service. If we go slightly beyond your "melted into" statement, the chances of the book becoming useful to the public are much higher in the AI company's digital archive, which might turn into something like what Google Books could have been, given the right incentives and copyright law changes.

And of course that presumes that the TV guide is going to stay in the museum rather than been thrown out as part of curation (or realistically, long before it makes it into a museum). Neither museums nor archives hoard everything, throwing stuff out is - as far as I know - one of the key jobs of an archivist. And a 1994 TV guide, while useful to understand what it was like to be alive in 1994, likely doesn't contain much unique information. You don't need that specific guide.

If there are 52 weekly editions, of 10 different guides, you would likely get most of what you want from any one of them. And for the parts that you wouldn't - there's a good chance that you'll have a much easier time getting the essence of this knowledge from the anonymized data pool that all the content was melted into, rather than chasing 10 different museums to find the original magazines.

skeledrew 17 hours ago | parent | prev [-]

A museum - or any other building - can only hold so much physical stuff. How much of it do you really want preserved? How do you choose what is preserved (it's an eventually inevitable choice)? Do you save the 1980s stuff but not the 90s? Or save every even/odd year? Some other method? How much direct access do you think people need to pre-digital history?

15 hours ago | parent [-]
[deleted]
throwaway219450 17 hours ago | parent | prev [-]

The BBC has both in some cases, but we know for sure what was broadcast when:

https://genome.ch.bbc.co.uk/about

Historic TV guides are also the sort of strange ephemera that people collect. They ought to be digitized like newspapers and other magazines, but this was always the purview of libraries anyway.

unleaded 18 hours ago | parent | prev [-]

Whose job is it to dictate what is and isn't worth saving?

azan_ 18 hours ago | parent [-]

Well for example yours - if you don’t pay for these rare books and don’t store them in good condition, then you have decided that they are not worth saving. Many of these rare books would run you like I don’t know, 1 buck?

unleaded 9 hours ago | parent [-]

I don't think I have enough money or room to buy all the books I think are worth saving, and my list no doubt has some overlap with someone else's. If I did buy them, what if someone else wants to read them? If I lose it or it gets destroyed in some way, the chance it will be gone forever increases (assuming there are multiple copies). Entrusting the availability of knowledge to individuals like that sounds like a bad idea. There could be some kind of publicly funded organisation that can take care of a big collection of books, afford to keep them safe, and make them accessible to anyone.

zmmmmm 20 hours ago | parent | prev | next [-]

It is also the case that the copyright holders are often putting restrictions around use of electronic forms that are driving the desire to use physical copies. I doubt AI companies would use a single physical book if they could avoid it - absent the legal cloud over electronic rights.

I have no evidence but I can't help suspecting in part the publicity around this is driven in part by rights holders that want to force AI companies back to e-books where they can force them into licensing deals.

hn_throwaway_99 20 hours ago | parent | next [-]

There is a whole legal saga here that is often misunderstood. Googling "Project Panama" should give more information.

The legal ruling from Judge William Alsup declared that if AI companies purchased the books legally and then copied them to their servers, it was fair use as a "transformative" operation, but the originals had to be destroyed in that case, because then there was only one copy still in existence (the one on Anthropic's servers):

From https://www.theguardian.com/commentisfree/2026/aug/05/anthro...

> Under US copyright law, the “fair use” doctrine allows you to make “transformative” use of copyrighted works without the owner’s permission. Anthropic took printed books and scanned them, “transforming” or remediating them into a new, electronic format. They then disposed of the original printed copy: the “destructive” part of destructive scanning. Along the way, Anthropic’s vendors had already sliced the spines and edges of the books, to scan them more easily before destroying them. “One replaced the other,” as Judge William Alsup wrote, noting: “There is no evidence that the new, digital copy was shown, shared, or sold outside the company.”

ivell 17 hours ago | parent [-]

Can they keep backup of the digital copy?

hn_throwaway_99 8 hours ago | parent [-]

It's a good question.

This site, https://copyrightalliance.org/education/copyright-law-explai..., states "It is important to note that this exception for backup copies only applies to computer programs and not to other copyrighted works, such as digital movies, music, or photographs or ebooks." But it seems unbelievable to me that they would have spent millions copying all these books and not have backups.

freejazz 19 hours ago | parent | prev | next [-]

> I doubt AI companies would use a single physical book if they could avoid it

They just don't want to pay what the copyright holders want to charge

alightsoul 20 hours ago | parent | prev [-]

Ai companies don't use ebooks, because they are more expensive than second hand books

breezybottom 20 hours ago | parent | next [-]

They absolutely do. Meta torrented 81 terabytes of ebooks. They just have no incentive to pay when the law looks the other way.

hn_throwaway_99 19 hours ago | parent | next [-]

The entire ironic thing here is that a huge part of those 81 terabytes of ebooks that Meta torrented were directly pirated books from Anna's Archive.

alightsoul 20 hours ago | parent | prev [-]

I meant paid ebooks. That's probably what the commenter refers to, because that's what publishers want. Obviously ai companies don't want to pay so they try to use pirated ebooks

warkdarrior 16 hours ago | parent | prev | next [-]

On Amazon right now, retail prices for e-book copies are higher than for the corresponding paperbacks.

steelframe 15 hours ago | parent [-]

This is exactly why Amazon has also been doing this acquisition and destructive scanning of millions of books for some time now.

bawolff 19 hours ago | parent | prev [-]

i imagine its because the doctrine of first sale does not apply to ebooks.

lkbm 2 hours ago | parent | prev | next [-]

With many old books, a big part of the problem is that it's non-trivial to determine who owns the copyright. Sometimes the contract would say the copyright reverts to the author after a certain amount of time out of print,but you have to go dig through old contracts to figure out whether that's the case for any given book.

cm2012 19 hours ago | parent | prev | next [-]

Yes. I dont understand at all what AA is worried about. One copy of a book is no big deal? good will and used book stores throw out a lot more than that.

wesleywt 16 hours ago | parent [-]

They are scanning "rare" books. I presume there are not a lot of copies left to throw out.

joshstrange 12 hours ago | parent [-]

Rare by whose definition?

I’m not aiming this at you directly by: ISBNs or STFU

Show me which “rare” books they are destroying and _maybe_ I’ll care but so far the pearl-clutching over this leads me to believe it’s people worked up about the idea of destroying (except it’s not destroying, it’s transforming, a fact often ignored) books, books that it’s not clear at all there is any strong demand for.

People want to invoke things like F451 but it doesn’t compare in the slightest. It’s like when people get mad about libraries throwing away or otherwise liquidating books that no one is reading in order to bring in books people want to read. People get all up in arms about that as if a book itself, in isolation, is inherently valuable or worth protecting. It’s not. If no one wants to read it then what value does it have? The impetus is on the people that think the book has value, it’s on them to carry the torch, to preserve what they think is worthy.

It would be like a company going to a yard sale and buying unsold/unwanted items to 3D scan them and destroy them in the process. This isn’t breaking into the Louvre and destroying one-of-a-kind artwork.

squidbeak 11 hours ago | parent [-]

> Show me which “rare” books they are destroying and _maybe_ I’ll care but so far the pearl-clutching over this

BBC good enough for you?

https://www.bbc.com/news/articles/cp3rprx2wl4o

"A recent academic text published in only 100 copies, 75 of which are already in libraries, may be very rare on the market - but it is perhaps not such a great loss if one copy is destroyed," says Derek Walker, owner of Edinburgh bookshop McNaughtan's.

"But we have, and have sold, books which are for example the only known surviving example of an edition from the 18th century.

"It would be a much more significant problem if one like that were to be bought for destruction, having survived this long."

joshstrange 11 hours ago | parent [-]

It would be if it made the point you think it is. All of this is more and more hand waving. 75 copies in museums? Then I think we’re good. As for the 18th century books, that’s pure speculation. It’s like a museum saying, “yes we have they prints for sale that they keep buying and destroying but wouldn’t it be a shame if someone destroyed the actual Mona Lisa?”.

And lastly, if these books are so important, then don’t sell them, hold onto them, digitize them without destroying them. This isn’t complicated. Amazon/etc aren’t breaking into museums and libraries, they are buying books on the open market.

If these books are so rare and important, then why has no one cared until now to actually preserve them?

squidbeak 10 hours ago | parent [-]

> It would be if it made the point you think it is. All of this is more and more hand waving. 75 copies in museums? Then I think we’re good.

You're repeating the seller's contextual point, as if it's a counterargument. Do I need to explain to you that what makes the lone surviving 18th century edition important is that there aren't 75 copies of it in museums?

> And lastly, if these books are so important, then don’t sell them, hold onto them, digitize them without destroying them.

It's good to see you agree any digitization of this category of book should be non-destructive.

> If these books are so rare and important, then why has no one cared until now to actually preserve them?

(Lastly for realz this time, eh?) Why has no-one cared to actually preserve the actually preserved book being sold by the bookseller... Bit of a strange question, that.

joshstrange 10 hours ago | parent [-]

You are conflating 2 parts of the article to make it say something it's not. No one has put forward any examples of a "lone surviving 18th century edition" being bought up my AI companies and destroyed. That was just an example of "wouldn't it be terrible if", not a "this has actually happened". It's a complete hypothetical, it's a made up scenario, it's a boogeyman. You've hung your entire 2 comments on something that has no evidence of happening.

> It's good to see you agree any digitization of this category of book should be non-destructive.

I don't. Digital or physical, it's the same, there is no difference in my mind. I have many paper books but they are art, not functional, I also have the ebooks which is what I actually read. The _ideas_ are what's important, not a dusty, decaying shell in which the ideas are contained. I wouldn't shed a tear over every library digitizing their books and destroying the physical versions, nothing is lost. More importantly, once you've bought something it's yours, yours to read, yours to display, yours to destroy. On hacker news, of all places, the people fighting _against_ first sale doctrine is appalling.

> Why has no-one cared to actually preserve the actually preserved book being sold by the bookseller... Bit of a strange question, that.

Our definitions probably differ here but preserving is not storing a book, preserving is ensuring that even if this copy is destroyed the ideas inside live on. I think that people that hoard (actually) rare books without a thought or care to making sure the text inside is preserved for future generations out of some desire to simply own something rare are the actually monsters here. And let's dispose with the notion that booksellers are "preserving" books, they are holding inventory, inventory they were happy to sell to Amazon/etc. If AI companies were raiding museums at gunpoint we'd be having a different discussion. They are buying books for sale, if they keep them in a library at corporate or scan and destroy them it makes no difference.

ctm92 14 hours ago | parent | prev | next [-]

They also only scan books that are easily and cheaply available, which means they are either not rare or have no significance.

Books that are rare of have historic significance will surely be in museums or libraries and not going away for pennies.

tgsovlerkhgsel 8 hours ago | parent | prev | next [-]

Not the copyright holders, "we the people": Copyright is an artificial legal construct that was repeatedly ratcheted up over and over again.

Unfortunately 50 years after the death of the author (or 50 years after publication for corporate owned works) has been locked in as a minimum term through international treaties, so it'll be somewhat hard to lower it beyond that, but many countries (including the US) enforce much longer terms, so that would be a first lever that could be applied quickly.

Maybe countries could could also establish an exception for out of print books offered to the public for free, or a general "library exemption" for public archives after a certain number of years?

I'm sure one of the AI companies would be willing to host a LibGen style library as a PR measure if legally allowed (with sign up required for rate limiting and as an extra benefit for the company to get daily active users).

asdefghyk 14 hours ago | parent | prev | next [-]

RE "...Instead, they enforce the copyright and force AI companies to shred books they want to ingest...."

Why are AI companies forced to shred books?

jeroenhd 14 hours ago | parent [-]

They want to have a digital copy and the judge ruled they can only keep one copy.

drtgh 12 hours ago | parent [-]

You do not need to shred books to scan them. This only happens if you don't care about preserving the integrity of the books and you want to scan more cheaply.

jeroenhd 11 hours ago | parent [-]

But their legal framework for being permitted to scan the books en masse (they are "transforming" the book from physical to digital) requires destruction of the original. Otherwise it wouldn't be transforming, it would be duplicating.

9 hours ago | parent [-]
[deleted]
ErigmolCt 16 hours ago | parent | prev | next [-]

Copyright holders are certainly responsible for keeping unavailable works inaccessible, but shredding is mostly an industrial scanning decision, not a copyright requirement

ranit 12 hours ago | parent | prev | next [-]

> Also, an AI company would only ever care to purchase, scan, destroy a book once. Presumably many books have more than one copy.

The OP article sounds quite opposite though - that AI companies are doing exactly this - destroying books so only they have the scanned content.

breezybottom 20 hours ago | parent | prev | next [-]

They don't "force" anything. Trillion dollar AI companies and their owners have as much agency as book publishers.

parineum 20 hours ago | parent | next [-]

They are "forced" to do this because that's what they have to do to abide by copyright law. They can't create a digital duplicate without destroying the original.

freejazz 19 hours ago | parent | next [-]

If that's true, then why did they pirate so many books?

scarmig 18 hours ago | parent | next [-]

Is your complaint that they follow copyright law, or that they don't follow copyright law?

freejazz 12 hours ago | parent | next [-]

I'm not complaining

Planktonne 14 hours ago | parent | prev [-]

A company with no respect for copyright law can't pretend they respect it when it is convenient for them; it's disingenuous, and shows that there is clearly another explanation.

parineum 9 hours ago | parent | prev [-]

Because they hadn't been sued for pirating so many books yet.

freejazz 8 hours ago | parent [-]

Well its not like copyright was invented last year...

breezybottom 20 hours ago | parent | prev [-]

Since when do AI companies care about copyright law? They're destroying them so their competitors can't use them.

gpt5 19 hours ago | parent | next [-]

A little meta - I want to point out demagogic/populist comments like these that try to clear all nuance and brush a topic in black and white tend to come from a really small portion of the users here, but the same user (whose account is only 4 months old), dominates posts like this by posting many many bait-like comments in the same post that deviate the discussion away from meaning and insight.

This is just one example, but it has become unfortunately common across all social media platforms.

alightsoul 19 hours ago | parent | prev | next [-]

Since they had to pay 1.5 billion for it

freejazz 8 hours ago | parent [-]

You'd think they could've negotiated something better than $3k per work if they had actually gone to the table first.

eru 19 hours ago | parent | prev [-]

> They're destroying them so their competitors can't use them.

That's pretty silly. My competitor can't ride my bike either, and I didn't have to destroy the bike for that to be true.

tptacek 20 hours ago | parent | prev [-]

To do what?

runarberg 20 hours ago | parent | next [-]

To not destroy rare books.

fenomas 19 hours ago | parent | next [-]

Not under US copyright law. The Bartz case ruled that if you scan a book and destroy the original it's considered format-shifting and you're likely fine. But not so for keeping the original and using the scan in its place - Internet Archive tried that (in an incredibly limited way), and publishers sued and won.

So companies scanning books already know they'll be sued, successfully, if they don't destroy the originals. So they destroy the originals.

tptacek 19 hours ago | parent | prev [-]

When you read "rare books", what are you thinking these are? The 404 article that spun this story up goes into more detail. These are like vanity press books. They're rare because nobody cares about them. The book industry already destroys these books.

runarberg 19 hours ago | parent [-]

I am thinking about a long essay Icelandic author Þórbergur Þórðarson wrote to his pals abroad, and were left abroad. I am thinking about a photo book by an Indonesian naturalist who is famous on Bali, and took amazing photos of wildlife on Sulawesi in 1926, and colored in, and somehow ended up in New Jersey in the 1980s. I am thinking about a collection of essays written by a teenage J.D. Salinger who he left unsigned at a café thinking nobody would want to read them but just leaving it up to chance. Or maybe a Jackson Pollock sketchbook he lost at a party which ended in the host’s bookshelf, and finally at an estate sale.

Plenty of such unknown unknown exist, and the AI machine will inevitably destroy a bunch of them at this scale.

tptacek 19 hours ago | parent | next [-]

You're trying to imagine rare valuable books and then fantasizing about AI companies destroying them, but what's actually happening here is that AI companies are acquiring, digitizing, and then pulping the instruction manuals to 1983-vintage copy machines.

This is all such a special-pleading argument. You know what other institution snatches up books and destroys them at huge scale? Public library systems. People clean out their attics and basements and drop off huge boxes full of books at libraries; libraries take the things they know will circulate, and destroy the rest. Take a guess as to how Þórbergur Þórðarson fares at the Newark Public Library. Wait, bad example, they stopped accepting book donations because nobody wants your old books. They tell you to give the books to thrift stores instead. Guess what the thrift stores do with them?

You know how many times I've read stories about the grave damage libraries are doing to human culture? Zero, zero times.

frm88 10 hours ago | parent | next [-]

You're trying to imagine rare valuable books and then fantasizing about AI companies destroying them, but what's actually happening here is that AI companies are acquiring, digitizing, and then pulping the instruction manuals to 1983-vintage copy machines.

Source? You state that in a tone that implies you have verifiable knowlege of this. All the information I found says that the exact number, titles and authors are under NDA.

svachalek 8 hours ago | parent | next [-]

Given no information, would you assume companies that are buying "huge quantities" of "rare" books are getting first edition Jules Verne by the crate, or 1989 Highlights magazines recovered from dentist offices?

runarberg 7 hours ago | parent [-]

Both. I think they are just buying whatever they can get their hands on most of it is going to be junk, but one in a million is going to be a rare and valuable book which we didn’t even know existed.

After scanning and destroying 10 million books, AI companies will have destroyed 10 such books.

(I am obviously assuming a Poisson distribution here where I pulled the parameter p = 1/1000000 out of thin air; point is p is non-zero; and at this scale the undesirable event is bound to happen a bunch of times).

tptacek 8 hours ago | parent | prev [-]

It's in the 404 story.

runarberg 9 hours ago | parent | prev [-]

I assume public libraries know what they are doing and know which book they are destroying, that they keep an accurate inventory and hire professionals maintaining said inventory and marking which books can safely be destroyed.

I assume no such things of AI companies.

tptacek 8 hours ago | parent [-]

Obviously that's not true. Newark is literally telling people to talk all their old rare books to thrift stores, which will immediately throw them in dumpsters.

runarberg 7 hours ago | parent [-]

This is a bad comparison. Public libraries are trying their best to preserve rare copies, despite being overwhelmed and underfunded. If they come across a rare copy chances are it will be preserved. When an AI company comes across a rare book, it will not recognize it as such.

There is also a fundamental difference in intention. AI companies are seeking out old books and will destroy them. Public libraries are trying their best but simply don‘t have the resources to find rare books in every collection.

And finally there is a fundamental difference in scale. Public libraries are not buying millions of copies to destroy, completely eliminating the chance they will ever be discovered.

---

PS. I don‘t like the anti-academic tone of your post. Librarians know what they are doing, their expertise is valuable, and when they act according to their specialized knowledge it does have a positive effect on the world.

tptacek 6 hours ago | parent [-]

What comparison? I'm talking about libraries as they actually exist and you're talking about what you think libraries should be doing. I'm arguing positively, and you're arguing normatively.

I'm not "anti-academic". I'm just very involved in my local community and I read the annual reports from our library system. We accept book donations, like a lot of suburban library systems do (from our extremely book-y community), and we are open about the fact that most of those books get trashed. We do not have the personnel on staff you claim libraries generally do.

I would say that between the two of us, I'm the one arguing more respectfully about libraries. I see them as real institutions with real pressures and constraints that do an important public service, and you see them as an instrumentality in your argument against AI.

runarberg 5 hours ago | parent [-]

You were the one who brought public libraries into this conversation, in an attempt to legitimize this practice by AI companies. I assumed this attempt was you saying the two behaviors are comparable.

I pushed back on that comparison. These behaviors are in fact not comparable. What the AI companies are doing is bad actually, and what libraries are doing is, while unfortunate, acceptable, given the limited resources they have.

tptacek 5 hours ago | parent [-]

Yes, I brought public libraries into this conversation, because they daily destroy hundreds of thousands of books, most of them without any professional evaluation whatsoever. I don't have a problem with this, because I understand a little how the book trade works, and whatever romantic idea people have about the preservation of books, it has little to do with what books actually are.

What I don't understand is your response to this. An AI company digitizes a book before destroying it: "bad actually". A public library takes a cartload of books to a dumpster without so much as opening the front cover of any of the books: just fine.

dukeyukey 17 hours ago | parent | prev [-]

Can you explain why AI companies destroying these is worse than a library or bookseller destroying these?

runarberg 9 hours ago | parent | next [-]

Public libraries staff professional archivist, they take careful inventory and know exactly which book they are destroying.

AI companies buy boxes and boxes of books with unknown content staff anybody who can operate a scanner and have no idea which books they are destroying.

If a public library comes across a book they didn’t know they had, it is very likely that somebody will notice and know how to continue, to find out if this book is worth saving etc. AI companies will treat this book exactly like any other and destroy it to feed the plagiarist machine.

kasey_junk 7 hours ago | parent | next [-]

This is not what most public libraries do with donation books. They mostly just recycle them.

The very act of cataloging donation books takes a huge amount of resources that the collection managers don’t have.

Many public libraries don’t have a collection management department at all! They outsource that function entirely to vendors and those vendors mostly source new books and spend resources preparing them for the hard use of a library (changing bindings, uploading and cleaning metadata, tagging, etc).

Decommissioning is done by volunteers who just look at the condition of books and chuck the bad looking ones in a bin for recycling.

My wife worked in this industry on the vendor side, your model of public libraries is closer to a small subset of certain big city libraries narrowed to their rare and research departments. The median book bought by a library is a Danielle Steele romance novel packaged for library consumption and sent to a small town library system that will be lucky to have a single professional librarian for the whole system.

tptacek 8 hours ago | parent | prev | next [-]

At this point I believe you're just sort of making up a fantasy of what you think a public library should do, not what they actually do.

dukeyukey 7 hours ago | parent | prev [-]

Rare and valuable books are not going to be in the giant pallets AI companies are buying. Most public libraries do not have professional archivists or anything like that, not so charity shops, or booksellers.

runarberg 7 hours ago | parent [-]

No, but they are going to be intermixed in boxes from an old estate, or a used books store which is closing up, etc.

> Most public libraries do not have professional archivists

Most reasonably sized library systems do in fact. Even small ones have trained librarians with an some degree in library science who took at least an introductory class in archiving as a part of their degree, and very likely has some idea how to prevent rare books from being destroyed.

tptacek 5 hours ago | parent [-]

Right, and what happens to those boxes from old estates are that they get put into dumpsters. The only difference here are:

* It's happening on a much smaller scale

* They're actually digitizing the books before they destroy them

This obviously isn't about the books. It's about people want reasons not to like AI companies.

runarberg 5 hours ago | parent [-]

The scale matters. It is disingenuous to just remove it from the equation as if it doesn’t.

As does the intention matter. The AI companies are doing this because they want to profit off of it. They don’t have to do this, and the fact that they do is part of what makes them bad for humanity.

tptacek 5 hours ago | parent [-]

Right: the scale matters, which is why it's very weird to take AI companies to task for something the public library system does on a far greater scale. And, in fact, the AI companies do have to do this: they're required by law to destroy these books.

breezybottom 11 hours ago | parent | prev [-]

"Can you explain why Soviets persecuting Jews is any worse than Nazis persecuting Jews?"

dukeyukey 7 hours ago | parent [-]

If someone was ok with Nazis killing Jews but not Soviets, I am going to think they don't believe it's the killing that is the problem.

20 hours ago | parent | prev [-]
[deleted]
HedonicEscal8r 20 hours ago | parent | prev | next [-]

If only this complaint was being posted by an organization ideologically opposed to copyright itself!

raincole 20 hours ago | parent | prev | next [-]

> Instead, they enforce the copyright and force AI companies to shred books they want to ingest.

What? Even if there are no copyright holders, the AI companies will still do scan'n'destroy because it's just cheap.

Are you expecting the authors/publishers to send digital copies to AI companies directly? Or expecting AI companies to preserve the physical copies indefinitely? Both are not gonna happen, copyrighted or not.

remus 12 hours ago | parent [-]

> ...AI companies will still do scan'n'destroy because it's just cheap.

There is also a legal element. If they kept the physical copy around after scanning the argument is that they're making copies of the book which puts them on tricky legal ground. By destroying the physical copy they can argue that there is only one version of the book that now exists solely in digital form, so this usage is better protected under fair use.

bondarchuk 15 hours ago | parent | prev | next [-]

We the people in the society who have the power to make laws through democratic means are the ones locking these books up.

wotamess 18 hours ago | parent | prev | next [-]

"want to ingest"

Not "need to ingest"

Copyright holders are capitalizing on laws on the books just like Jeff Bezos companies buying their own copies to shred

So in the end it's really a Congress problem as usual

sophacles 19 hours ago | parent | prev | next [-]

If you're buying second hadn books by the lot, you'll get a lot of duplicates and its eaiser to scan wholesale and dedupe in the computers than it is to try to run a sorting operataion on "things".

signa11 15 hours ago | parent | prev | next [-]

mr. vernor-vinge's "Rainbows End" is oddly prescient ! Highly recommended nevertheless.

customguy 18 hours ago | parent | prev | next [-]

> force AI companies to shred books they want to ingest.

Nothing forces them to shred books, they do it because it's slightly cheaper that way.

hparadiz 18 hours ago | parent | next [-]

There was a court case where they said that if they copied the books it's not fair use because they didn't own it but if they bought physical copies and then destroyed them then somehow it was fair use because it fell into the niche of personal backups. I forget the details but basically they buy one time prints and destroy them immediately.

customguy 12 hours ago | parent [-]

I had no idea about that, or how fucking bad this actually is:

https://en.wikipedia.org/wiki/Project_Panama

So I stand corrected: at least some don't do it because it's cheaper (than to buy a license, or simply forego some things), but because they're fucking evil, or so stupid that it effectively is the same as being extremely evil.

skeledrew 17 hours ago | parent | prev | next [-]

They do it because, for each work, they bought one copy, which they scan and no longer need the physical version of, and would be in copyright violation if they keep more copies than they bought.

annapanna 15 hours ago | parent [-]

>and would be in copyright violation if they keep more copies than they bought.

They can contact the copyright holder and ask/buy a license to make multiple copies.

hparadiz 15 hours ago | parent | next [-]

yes and they can ask for a pony and then a unicorn too

skeledrew 15 hours ago | parent | prev [-]

For what?

postepowanieadm 18 hours ago | parent | prev [-]

By destroying them they don't copy only convert them into another format.

20 hours ago | parent | prev | next [-]
[deleted]
ajsnigrutin 14 hours ago | parent | prev | next [-]

It's also the regulation, where most systems still look at "one pirate copy" = "one sale of lost profits", especially when pirates end up in court. If the book (or game or whatever) is not sold anymore in any way where you could give the copyright holder money in an easy accessible way (eg. buy it on amazon, or a local bookstore), they shouldn't be able to claim losses from piracy, since they clearly don't want your money.

On the other hand, there are grey zones here, the lord of the rings books (still copyrighted and easily obtained pretty much everywhere) have been translated into my language many decades ago, and many of us read and liked those translations, but when the movies came out, a new translator did a new translation, where they changed a lot of things, including the last names of bilbo and frodo (Bogataj->Bisagin) and the Shire (Grofija->Šajerska), and the old version is sadly available only in paper form on second hand markets. On one hand, copying that if you only want this specific version would not cause a lost sale, on the other, you can get new translations (or english originals) pretty much everywhere.

watwut 16 hours ago | parent | prev | next [-]

Copyright allows you to sell book you have and does not force you to shread it.

They are not forced to shread them by copyright.

wesleywt 16 hours ago | parent | prev | next [-]

I was wondering what the pro book shredding take was going to be. Why destroy the book after scanning? You can create a beautiful library of rare books with all the AI debt bubble.

ls-a 19 hours ago | parent | prev | next [-]

[dead]

jacobo37 20 hours ago | parent | prev | next [-]

this is plainly stupid ... many of these books are likely to have no current publisher nor any way to "reprint" the book. "ai" companies are simply burning our cultural context ...

rpdillon 20 hours ago | parent | next [-]

Wait: the entire premise of copyright is to prevent someone from publishing a book, and a competitor buys a copy, clones it, and sells copies way cheaper because they don't have to pay the author.

Now, in 2026, we're acting like cloning a published book is not technically feasible? That doesn't track. With publishing on-demand, it's easy to imagine a business with digital copies of all these works that they make available for print-on-demand.

The uncomfortable reality is that most of these books are nothing anyone cares about. Even the book sellers in the 404 story call them dead inventory.

Can we get some actual book titles into the discussion so we can focus on facts rather than speculation?

alightsoul 19 hours ago | parent [-]

This is not a technical problem at all. This is a copyright problem. Anthropic thought it was just a technical problem until they had to pay 1.5 billion after they lost a copyright court case

Op means a lot of those books were made before computers were used for that purpose and the publishers and probably authors no longer exist, so there is no digital copy to just reprint, unless someone scans it themselves and publishes it, risking copyright violation when done at large scale due to possible exceptions to this rule

rpdillon 12 hours ago | parent | next [-]

> publishers and probably authors no longer exist

Who is going to claim a copyright violation?

Sha1rholder 18 hours ago | parent | prev [-]

> Anthropic thought it was just a technical problem

Did they? Then why did they "don't want anyone to know about this"? Or do you think their lawyers are dumb?

dukeyukey 17 hours ago | parent [-]

I imagine they know the public would react like the public is reacting. Like obviously this is not worse than what libraries and thrift shops do daily, but it is bad optics.

Sha1rholder 13 hours ago | parent [-]

I don't think so. Libraries and thrift shops usually don't consider it a good idea to damage those precious, hard-to-reprint books. None would say anything if Anthropic only destroyed 1 million Harry Potter or Foundations

footydude 14 hours ago | parent | prev | next [-]

> "ai" companies are simply burning our cultural context ...

Most developed countries have a 'legal deposit' system with a national archive/national library that requires publishers to send a copy of their works to them. They've existed in some form for centuries in some countries.

Example for the UK - British Library guidance: https://www.bl.uk/services/legal-deposit

Ekaros 17 hours ago | parent | prev [-]

Many of them don't have current publisher because no one wants to buy them. I would bet that vast majority of these books have no commercial market.

jscd 18 hours ago | parent | prev [-]

Sorry, is your stance seriously that authors and publishers should digitize and freely distribute their work, at their own expense?

Also, who’s forcing AI companies to “ingest” books in such a destructive way?

Also also, if there’s one thing I’ve learned from AI scrapers, it’s that they’d never scan the exact same thing multiple times at the expense of public access to the resource.

scarmig 18 hours ago | parent | next [-]

Relinquishing copyright does not imply any of the labor you're suggesting. It's the opposite: you're just committing not to perform the labor of pursuing legal action against someone who does digitize and freely distribute the work.

Anna's Archive, for one, would be more than happy to host at no cost to the author.

jeroenhd 14 hours ago | parent | prev [-]

AI companies are buying the physical books, they can turn them into confetti if that's what they want to do. If the physical books are running out, the authors can print and sell more. Or they can sell digital copies so the information is not lost.

The law is currently forcing these companies to destroy the books after scanning them.

jscd 10 hours ago | parent [-]

> The law is currently forcing these companies to destroy the books after scanning them.

No, the law is stopping them from digitally sharing their scans. They are perfectly capable of reselling or donating or storing the books they buy. (Wasn’t Amazon originally a book seller?)