Remix.run Logo
AI companies are shredding rare books(xcancel.com)
376 points by anon373839 2 hours ago | 219 comments
squidbeak an hour ago | parent | next [-]

I've limited sympathy for the publishers.

It pisses me off to reflect that they can sit on works until copyright expires, keeping them out of print. There's no real need for any of these so-called rare books to be rare while they're under copyright.

And related to this, the books that are in print are mostly only in print in the shittiest way. I often see well-made books from the 17th or 18th centuries which are still in good nick. It's ridiculous that in the 21st century, publication standards have fallen to the point where for most works a disposable format is the only type available - where no amount of money could buy a truly decent hardback copy.

If we have to have copyright laws, I'd like to see two changes to them.

When a publisher has no incentive to keep an edition in print, it should be available to any other publisher to print, without compensation to the original publisher, and with renegotiated royalties for the author.

And if the publisher keeps a book in print - but only in bestseller-grade materials, bogroll paper that furrows in any humidity and perfect binding that molts its pages a couple of dry seasons later - and if it refuses to print a durable hardback copy with signatures, good paper and decent print - something that will still be readable in several generations' time - any other publisher keen to have a crack at it should be able to, again without any compensation for the original publisher, though perhaps in this case, with matching royalties for the author.

kasey_junk an hour ago | parent | next [-]

Books in the 17th and 18th century often didnt get bound by the publisher. They were done by independent binders for _custom_ orders. You’d see whole libraries with the owners binding/cover standards rather than per book.

Books in that time were _luxury_ goods. Most people could not afford them. One of the ways that was changed was to introduce cheap, mass produced bindings that were lower quality than the bespoke artisianal bindings done by specialist craftsmen.

You can still get custom bindings done. There exists whole niches on the internet of crafters that will take a production run book and strip its binding and make you extremely high quality and custom bindings and covers.

DrewADesign an hour ago | parent | next [-]

I worked in an academic library that rebound pretty much every book they acquired. One of the biggest in the world, too.

kasey_junk an hour ago | parent | next [-]

My wife worked for a company that specialized in rebinding paperbacks for schools and libraries. It makes more economic sense to do that for niche use cases rather than make all print runs more expensive.

pfdietz 43 minutes ago | parent | prev [-]

Nowadays, academic libraries are moving their books to archival storage. New publications are electronic only. You need to be a formal member of the university community to access it, unlike in times past when any member of the public could stroll the aisles of physical books and journals.

ryanmcbride 31 minutes ago | parent | prev | next [-]

Learning this fact is what got me interested into binding my own books!

With how the quality of things seems to have been degrading over the years (either real or just me getting older and experiencing the impermanence of all things) I've been trying to adopt an attitude of "if this practice existed before the industrial revolution, I can _probably_ do it" and it's been really great to learn how things were made before they had to be mass produced as cheaply as possible.

Lutzb 31 minutes ago | parent | prev [-]

One of my relatives is one of those micro publishers. Selected works are printed and bound to extremely high standards and materials in a way that their customers are willing to pay 1k-25k+ per book. These editions only have a couple of prints and are mostly made to customer requests.

There is a market for these type of books, albeit a very small one.

tookmund an hour ago | parent | prev | next [-]

> I often see well-made books from the 17th or 18th centuries which are still in good nick

At the risk of stating the obvious, any poorly made books from then wouldn’t have lasted this long and so you would never see them.

ortusdux 10 minutes ago | parent [-]

https://en.wikipedia.org/wiki/Survivorship_bias

nyeah 3 minutes ago | parent | prev | next [-]

[delayed]

kevin_thibedeau an hour ago | parent | prev | next [-]

> I often see well-made books from the 17th or 18th centuries which are still in good nick.

Those books predate the development of wood pulp paper. It isn't the publisher's fault they can't economically print on rag paper anymore.

ShinyLeftPad an hour ago | parent | prev | next [-]

I have even less sympathy for IP stealing LLM operators

phoghed 40 minutes ago | parent [-]

It’s been determined that training on lawfully acquired works is fair use. Presumably in this discussion of shredding physical books Dario and Sam are not pulling heists at the local library.

I’m sure there’s ongoing litigation, and better sources than this, but fair use was determined in June 2025 in a sf federal district court https://www.goodwinlaw.com/en/insights/publications/2025/06/...

Similar conclusion vs meta https://www.jw.com/news/insights-kadrey-meta-bartz-anthropic...

And the more recent $1.5B settlement did not overturn it https://www.reuters.com/world/us-judge-approves-anthropics-1...

shimman 7 minutes ago | parent | next [-]

Might surprise people to learn but the law isn't the final arbiter on what is moral or just. The law only decides what is legal and as we've seen over our lived history as humans, many "legal" things may not be those we want society to uphold.

phoghed 2 minutes ago | parent [-]

Yes IP theft and fair use are discussing legality, not morality. Whether or not it’s moral for someone to make money off somebody else’s labor or not, I don’t want to get into a discussion about personally.

sofixa 21 minutes ago | parent | prev | next [-]

No, courts so far in the jurisdictions which have heard such cases, have ruled it's fair use. There is plenty of ongoing litigation in many jurisdictions, so it's way too early to just decree "it's been determined". It likely won't be for years to come.

bscphil 16 minutes ago | parent [-]

Why did Anthropic settle with authors for 1.5 billion then? Surely their lawyers must have decided there's a pretty good chance of judges ultimately deciding that it is copyright infringement?

rileymat2 8 minutes ago | parent | next [-]

They had illegally obtained books. 1.5 billion is an incredible deal compared to the per book infringement fine. https://fortune.com/2026/07/21/anthropic-copyright-settlemen...

hn_acker 9 minutes ago | parent | prev [-]

Anthropic settled because even though the training is fair use, Anthropic did not acquire all of the training material through legal means.

ButlerianJihad 35 minutes ago | parent | prev | next [-]

You don’t know what “Fair Use” is.

Fair Use is not an activity that you engage in. Fair Use is not a category with criteria that you meet. Fair Use is not a precedent that paves the way for everything afterwards.

Fair Use is a defense that can be used in court when you’re named in a copyright lawsuit. Fair Use is how you justify your actions before the court finds infringement.

hn_acker 17 minutes ago | parent | next [-]

Fair use is a legal defense with a specific test to demonstrate whether a particular instance of use of a copyrighted work is not infringement. Although fair use is always evaluated on a case by case basis, fair use does produce some precedent (example at [1], not related to TFA), and fair use is not mere justification [2]:

> Notwithstanding the provisions of sections 106 and 106A, the fair use of a copyrighted work... is not an infringement of copyright.

[1] https://en.wikipedia.org/wiki/Google_LLC_v._Oracle_America,_....

[2] https://www.law.cornell.edu/uscode/text/17/107

gruez 30 minutes ago | parent | prev | next [-]

>>It’s been determined that training on lawfully acquired works is fair use

>You don’t know what “Fair Use” is.

This isn't the opinion of some armchair HN commenter. Actual judges have affirmed this, as other commenters in this thread has pointed out.

parineum 33 minutes ago | parent | prev [-]

s/fair use/ip theft

The comment uses the same language as it's parent.

parineum 34 minutes ago | parent | prev [-]

Being informed is much too high a barrier for the people who just want to pretend that everyone with money is lex luther.

arduanika 10 minutes ago | parent [-]

If you're going to call people ignorant, you should probably know how to spell the name of the character you're referencing.

streetfighter64 an hour ago | parent | prev [-]

> When a publisher has no incentive to keep an edition in print, it should be available to any other publisher to print

How do you adjudicate that? And wouldn't it just lead to loopholes such as "Ghost Printings" (cf. Ghost Flights https://en.wikipedia.org/wiki/Ghost_flight_(commercial_aviat... ) where the books are technically printed in the required volume but practically unavailable to customers through one method or another. Because, the cost of wastefully printing a few books to warehouse, is less than the potential losses of the IP rights, probably.

trollbridge an hour ago | parent | prev | next [-]

We reprint old books after checking out copyrights (for all books, this means pre-1930, but for some (I'd actually say most) it also means ones published up to 1964 and 1973, depending on how the rightsholders did (or didn't) do the renewals).

We use a special guillotine type cutter to cut off the binding and then store the pages in a sealed plastic bag which goes in the archives; they're stored there indefinitely in case the book needs rescanned for some reason. We also keep the original, uncompressed copies of the books on magnetic disks.

We also go out of our way to try to find rare books published in 1931, 1932, etc. so they are ready to go once the copyright expires.

And no, no AI company has ever come to us and asked to run training on all of our scanned copies.

palmotea an hour ago | parent | next [-]

> We reprint old books after checking out copyrights

Who is we?

> then store the pages in a sealed plastic bag which goes in the archives; they're stored there indefinitely

Is that the best thing for archival storage? Like could things like chemical breakdown increase the humidity in the sealed bag or concentrate corrosive chemical vapors? I was under the impression the best environment was an actively climate-controlled environment.

vander_elst an hour ago | parent | prev | next [-]

Some source or citation or context is needed here, is this the work of a 2 person no profit or a trillion valued pre IPO company?

htrp 42 minutes ago | parent | prev | next [-]

> And no, no AI company has ever come to us and asked to run training on all of our scanned copies

Yet

butlike an hour ago | parent | prev | next [-]

Why cut off the spines? Isn't that how you end up getting unattributed 'dead sea scrolls'?

remus an hour ago | parent | next [-]

It's easier to get good quality scans from individual pages than it is from pages in a complete book. Imagine laying a book flat, then the page is distorted in the area around the spine. You can work around this (either by trying to correct for the distortion in software or with clever scanners that position the book more advantageously) but it adds complexity compared to chopping off the spine and just dealing with flat sheets of paper.

ryukoposting an hour ago | parent | prev | next [-]

Speed and cost. You can run the book through a typical sheet-fed scanner instead of using a contraption like this: https://linearbookscanner.org/

As for consumer-grade solutions, look for the Fujitsu SV600.

phasefactor an hour ago | parent | prev [-]

Easier to send them through a duplex scanner (or put them in a flatbed one if they are fragile). Cheaper than buying the automated ones with the page turning robot arm.

I have done it at home for my books since the mid-00s.

open-paren an hour ago | parent [-]

Do you purchase every book two times, or do you have a home devoid of physical books that you enjoy? Genuinely curious

yorwba an hour ago | parent | prev | next [-]

Most likely they don't know you exist. If you contacted them first, maybe something could be arranged...

freejazz an hour ago | parent [-]

Geez - their LLMs couldn't figure it out?

phoghed 38 minutes ago | parent [-]

Can you?

ck2 an hour ago | parent | prev [-]

or just send the books/scans to a country that doesn't recognize US copyright

est31 2 hours ago | parent | prev | next [-]

> You can reprint a bestseller. You can't replace the last three copies of an 18th-century botanical text once someone shreds them for training data. And the judge said it's legal. So it's going to accelerate.

Aren't they shredding only the books still under copyright protection? How is an 18th century botanical text still under copyright?

IDK about the shredding, it's not nice, but it's more a problem with copyright law than AI companies.

Scanning books you own should be legal from a copyright point of view, and not require shredding.

Second, one should think about abandoned property provisions for copyright works published more than 50 years ago and in danger of being forgotten: once challenged, either you as the owner have to prove that the work is preserved for future generations (e.g. in various libraries around the world), or you have to authorize further copies, or you give up copyright on the work.

ACCount37 2 hours ago | parent | next [-]

Scanning books by taking them apart into singular pages and scanning those pages is faster and cheaper. AI training is a numbers game, so they want faster and cheaper.

What happens to the pages after? No one needs them anymore, so they get mulched and recycled.

That would be the dominant scanning method even if copyright wasn't a thing. But then again - if copyright wasn't a thing, there would be much less need to scan any physical media.

The reason why OpenAI can't just go on Amazon, buy a "digital edition" of a 2018 book and use that is that it would violate the license in ten ways, and then the DMCA laws that forbid breaking DRM on top of it.

edoloughlin 2 minutes ago | parent | next [-]

> What happens to the pages after? No one needs them anymore, so they get mulched and recycled.

Strictly speaking, no one needs the Sistine Chapel or the Pietà etc. It would be a shame if they were mulched and recycled, though.

voakbasda an hour ago | parent | prev | next [-]

Think about that last point for a moment. Our “rights to read” are diminished significantly with digital works as compared to printed works. Right of resale. Right to lend.

In the end, digital publishing just isn’t right and will lead to massive gap in our historical records. They require active curation and cannot be preserved simply by resting on a dusty shelf.

butlike an hour ago | parent [-]

Every innovation since the microprocessor isn't worth saving in the grand scheme of things.

When today's algae evolve enough into tomorrow's sentient creatures, they're really only going to need up to the industrial revolution and should probably stop right before that.

dragonwriter an hour ago | parent | prev | next [-]

> if copyright wasn't a thing, there would be much less need to scan any physical media.

Because there’d be much less content created in any media to capture in the first place.

eru an hour ago | parent [-]

Empirically, probably not. We had lots and lots of content before copyright, and people seem to produce lots of content even in jurisdictions with weaker copyright.

classified an hour ago | parent | prev [-]

It's the law, logic doesn't enter into it.

sethops1 2 hours ago | parent | prev | next [-]

> Aren't they shredding only the books still under copyright protection? How is an 18th century botanical text still under copyright?

It's cheaper to scan the books if you do it destructively. Cost. That's why they're shredding irreplaceable texts. Nothing to do with copyright.

https://www.404media.co/ai-companies-are-buying-tons-of-old-...

graemep an hour ago | parent | prev | next [-]

> You can't replace the last three copies of an 18th-century botanical text once someone shreds them for training data. And the judge said it's legal.

An 18th century book would be out of copyright so why would it be illegal to keep the original and scan it?

mc32 2 hours ago | parent | prev | next [-]

Old rare books where there are single digit copies should enjoy some sort of patrimonial protection just like museum pieces. You can own them but have the state have the option to buy it if you’re about to significantly deface it or destroy it.

soco an hour ago | parent [-]

The only issue I see is, how could you tell which are those books?

9dev an hour ago | parent [-]

We manage to do this for endangered wildlife too without anyone counting every single specimen; why shouldn’t we be able to estimate how rare a book is?

phoghed 35 minutes ago | parent | next [-]

And if you instituted this, the commenters of this very website would surely decry it as a prime example of government overstep and waste.

shimman 6 minutes ago | parent [-]

Commentators on this web site work for some of the most evil organizations on the planet and have beliefs that 95% of the population rejects. You can safely ignore the YC cohort of devs and be fine.

eru an hour ago | parent | prev [-]

It's pretty expensive for the wildlife.

Most rare books are rare because no one cared enough about them. Ie most rare books are rubbish.

croes 2 hours ago | parent | prev [-]

Books that are shredded can’t be scanned by competitors.

lousken 2 hours ago | parent | prev | next [-]

That's why archive.org should have never been sued for lending books they had physical copy of. This is the result. Publishers should be more careful what they wish for.

kingstnap 2 hours ago | parent | next [-]

The archive.org story was more nuanced than that. If I recall correctly the full story was that they used to lend digital versions of books they physically bought and scanned with DRM to enforce a sort of one to one at a time restriction.

But during covid archive.org decided to just remove the limit and lend unlimited copies concurrently which started the debacle with the publishers.

cvadict an hour ago | parent | next [-]

> But during covid archive.org decided to just remove the limit and lend unlimited copies concurrently

IIRC, this was 100% it. Lending one digital version of one physical asset was likely already a violation copyright. Lending UNLIMITED digital versions of one physical copy was DEFINITELY a blatant violation of copyright.

phasefactor an hour ago | parent | next [-]

Correct, it was switching to unlimited lending instead of one lend per physical book that got them in trouble.

butlike an hour ago | parent | prev [-]

How is lending one digital version of one physical asset a violation of copyright? Since I MAY be able to lend out the physical as well?

ndiddy an hour ago | parent [-]

From the court decision:

"IA maintains that it delivers each Work “only to one already entitled to view [it]”―i.e., the one person who would be entitled to check out the physical copy of each Work. But this characterization confuses IA’s practices with traditional library lending of print books. IA does not perform the traditional functions of a library; it prepares derivatives of Publishers’ Works and delivers those derivatives to its users in full. That Section 108 allows libraries to make a small number of copies for preservation and replacement purposes does not mean that IA can prepare and distribute derivative works en masse and assert that it is simply performing the traditional functions of a library. 17 U.S.C. § 108; see also, e.g., ReDigi, 910 F.3d at 658 (“We are not free to disregard the terms of the statute merely because the entity performing an unauthorized reproduction makes efforts to nullify its consequences by the counterbalancing destruction of the preexisting phonorecords.”)."

Cthulhu_ an hour ago | parent | prev [-]

Yeah that was it; if I got this right, US libraries got the right to lend out one digital version of a book that they had in their inventory. Archive.org combined those digital versions so that people could check out a digital book if any library in the US had it (digitally) available. But during the 'rona they removed this limit and just lent out books regardless of it being "checked out" digitally from a library.

This wasn't a very smart move of them. I get why they did it but they put themselves at a huge legal risk.

Incipient 2 hours ago | parent | prev | next [-]

Publishers don't care if rare books get shredded?

the-grump 2 hours ago | parent | next [-]

And, regrettably, The Archive lent books regardless of physical possession.

Publishers had accepted the prior arrangement before The Archive decided to push it, if not explicitly then implicitly by not suing.

I'm a believer in The Archive's mission, and I wish they had treated the goodwill they'd accumulated as something worth preserving and not a currency to be spent.

It has been stated by many before me: lending books should have been handled by a separate entity, especially when they removed the physical backing requirement.

kmeisthax an hour ago | parent [-]

Just to be clear, publishers hadn't accepted the "controlled digital lending" (CDL) premise, not even with the one-to-one ratio. Their position was always "first sale ends when the atoms do". There was even controlling precedent: a few years before IA tried their online lending library thing, there was an "MP3 resale" company called ReDigi that had lost on very similar grounds. The publishers suing IA even made sure to sue in the same venue that had decided the ReDigi case so it'd be controlling precedent.

Furthermore, in the discovery for the Internet Archive case, publishers had already found a case where IA had lent out books despite knowing their partner libraries wasn't actually withdrawing loaned-out copies from circulation. The CDL premise was always just a suggestion, and IA would have still lost their case if they hadn't done the National Emergency Library (NEL) stunt or if they'd been sued in another venue that hadn't had the ReDigi case as precedent.

It's important to note that whenever a company decides to sue for copyright, it is often late, because the company is banking infringements up to the 3-year statute of limitations and because building a meritorious case takes time. The lack of a timely lawsuit proves almost nothing about the intent of a publisher with a valid case against you.

The thing is, I don't even think the whole stunt damaged much of the IA's goodwill? I know of a few people who withheld donations to IA, but that was mainly under the assumption that publishers would be getting a billion-dollar damage award that would immediately bankrupt IA and result in it's archives being sold off to Lexis-Nexis or something. The funny thing is, IA wound up settling for a sum so small they had to promise never to reveal it, and the danger is gone, so the only thing people complain about now is just that the NEL stunt maybe pushed them "above the radar" or something.

It's still insane that shredding books for AI training is legal, but this isn't.

azan_ 2 hours ago | parent | prev | next [-]

Yeah, why would it be bad for publishers? If anything they'd most likely encourage more book shredding!

infinite_spin 2 hours ago | parent | prev [-]

of the rare books, which was the rarest of them all? What year was it published?

JumpCrisscross an hour ago | parent | prev [-]

> That's why archive.org should have never been sued for lending books they had physical copy of. This is the result

How are these things remotely related? If anything, Archive.org’s callous, thoughtless approach nuked the hands of legitimate archival efforts.

ACCount37 2 hours ago | parent | prev | next [-]

The publishers sued AI companies for training on shadow library data, hoping to negotiate content deals for big $$$ down the line. Instead, they got analog hole'd.

Turns out that buying an old book for $5 and destructively scanning it for $25 is way cheaper than paying extortion fees to the copyright-mongers.

What I don't buy is it being "rare, precious books". First, they're not after ancient texts - they're after the books that there's still copyright on. Second, when it comes to books, "old" doesn't mean "valuable" - plenty of libraries destroy old books because there's no demand for them, and storage costs you. This is how those scanning companies get books for so cheap.

tencentshill 2 hours ago | parent | next [-]

So they're not valuable... except to AI companies. They should pay a fair amount.

ACCount37 2 hours ago | parent | next [-]

They are paying a fair amount. In the ballpark of $5 per book.

You know, piracy online is nice and simple - but it's kind of hard to get physical media without paying what the previous owner considers "a fair amount" to part with it.

chii 2 hours ago | parent | prev [-]

> They should pay a fair amount.

they should pay the marginal value that the next buyer would buy.

Do you also think that a person dying of thirst ought to pay the maximum price they could possibly pay for water?

soco an hour ago | parent [-]

It might be me but I fail to see the equivalence of AI companies with thirst deaths. That', or maybe because it's a strawman.

Minor49er 15 minutes ago | parent [-]

The point is that you can buy a book for a few bucks, read it, and have its insights completely available to you. But people are demanding that LLM companies pay far more to publishers for the same copies

SirFatty 2 hours ago | parent | prev | next [-]

I see.. so the various AI companies are in the right on this?

infinite_spin 2 hours ago | parent | next [-]

I think they are in the legal sense of right, and I think they only discarded the remains of these dissected books because previous rulings (e.g. archive.org's lending practices of digital copies of books they physically owned) gave rise to a situation where destruction bore less legal risk. As for the moral case, I don't have much to say on that, we all have our own lines in that sand.

jerf an hour ago | parent [-]

I would say that especially for books out of copyright, they are unquestionably legally in the right. There is no legal standing for "I like books and it makes me feel squicky when someone disrespects them". Nothing stops you from buying old books and using them as decoupage[1] fodder, using them to start a campfire, or making paper airplanes out of their paper.

Moreover, while it is emotionally appealing to some people to want to add some sort of "responsibility to society" to people who own the old books, it's a very emotional plea that can't really be manifested in the real world. As already pointed out in other places, "an old book" itself doesn't really mean much in terms of what its value is in any particular dimension. Plus, I am always deeply suspicious of anything that expands to "Other people, who are not me, should expend vast quantities of resources so that in the next five or ten times I think about this issue for the rest of my life I feel slightly better about this issue" which is what this really amounts to. I think that as superficially appealing as that may be, it's really a very hostile and demanding position to take.

Personally, to the extent that I would want to lay a "social responsibility" on the AI companies, I'd like to see something like they are either obligated, or ideally, just do it of their own free will, to make the scans of the books that are out of copyright available for some reasonable fee (ideally, "free because we like the PR", but given the scope demanding it be free is not reasonable), and without them trying to lay any further claims on the public-domain results. Trading "one old book somewhere, inaccessible to the world" for "a scan of the book and an OCR of it" I would judge a net win for society for rather a lot of these old books, which are by no means "worthless" sitting in some old collection somewhere but would be a lot more useful for being available.

[1]: https://en.wikipedia.org/wiki/Decoupage , since I imagine a number of people won't know what that is.

eru an hour ago | parent [-]

I agree. As for making the scans available:

They could give a copy of the data once to some third party organisation that then seeds it in bittorrent or something like that. Basically, what I want to say is that this doesn't need to be an ongoing obligation for the scanner to be worthwhile for society.

master_crab 2 hours ago | parent | prev | next [-]

It can be the case that everyone in “a fight” is wrong.

yeoyeo42 an hour ago | parent | prev [-]

yes. in the sense that they probably wouldnt want to do this but the law forces them to do an extremely dumb thing.

they're paying for the books, no shady things going on there. whether the publishers should deserve more than a single copy's worth is a separate question.

having the law such that its illegal to scan a book and then keep it, but legal to scan it and destroy it - gg no re there, law people retardmaxxed themselves as they tend to do with anything related to digital data.

margalabargala an hour ago | parent [-]

Law aside, destructive scanning is cheaper and easier. Cutting the book into individual pages would be happening regardless, because that is the easiest, fastest, cheapest way to scan it.

butlike an hour ago | parent | prev | next [-]

Modern texts become ancient texts; given enough time.

wwweston an hour ago | parent | prev [-]

Buying up old books is legal. Digitizing and distilling them arguably is too. And also:

> paying extortion fees to the copyright-mongers

Yes yes, greedy fat cat publishing oligarchs treading on the poor put-upon scrappy AI underdogs. /s

Back in reality, the fraction of people who got into publishing books to get rich collecting rents is… not large. There’s so many other fields that are likely to reward participants with more wealth that it’s absurd — even with all the passion for the work in tech it’s probably relatively less pure.

And whatever the excesses of copyright have been, the whole bargain has always been on more pro-social foundations and stronger intellectual foundations than “extortion” sneers. It recognizes that incentives matter and work that’s valuable should be rewarded and incentivized.

A culture that takes a Robin Hood approach to low marginal cost billing points but fawns over the hypercapitalized distribution King Johns isn’t creating a freer or richer society or fighting the real cartel center, it’s indulging resentment and caricature.

sherr 2 hours ago | parent | prev | next [-]

I see mentions of Bradbury's "Fahrenheit 451" in that thread but what this really seems to be mostly like is Vernor Vinge's "shred and scan" factory in his novel "Rainbows End".

vessenes 2 hours ago | parent | next [-]

Perhaps the last great near-term predictor. I often wish he'd written more. To remind us all, he predicted shred and scan would be a short stop over done by villains on the way to nondestructive scanning.

That said, supporting Anna's archive is one of the best things you could do for humanity long term in my opinion.

dwohnitmok 39 minutes ago | parent | next [-]

Vinge also coined the term "the Singularity" (https://accelerating.org/articles/comingtechsingularity) what he describes as "an opaque wall across the future" once superhuman artificial intelligence comes on the scene.

orthoxerox 2 hours ago | parent | prev | next [-]

Were they the villains? I remember the rogue three-letter-agency executive being the BBEG.

vessenes 2 hours ago | parent [-]

I think some villainy is implied by "billionaire-backed-woodchipping of a library," but it's just my interpretation, no actual knowledge of Vinge's perspective.

clickety_clack 2 hours ago | parent | prev [-]

That’s digital though, so it requires the continued survival of readers for the data that is stored. The best thing you could for the long term is probably to buy a few hundred physical books to keep in a bookcase in your home.

trollbridge an hour ago | parent [-]

Speaking as someone with dozens of bookshelves and tens of thousands of books... I kind of prefer the idea that continued survival means getting a bunch of 14TB drives and, you know, hosting certain files obtained from certain places. The reality is that most people's book collections are simply going into dumpsters, speaking as one of the people who go and try to buy these book collections at estate sales. (We can't do anything with the sheer volume of these books so most of them go into the dumpster. I cannot store hundreds of thousands of books, or millions, and no libraries want them.)

Also, holy cow, hard disks (as in the magnetic oxide kind) got a lot more expensive.

vessenes an hour ago | parent [-]

Yeah agreed. I have on my long-term project list a hardware 'oracle' that would have everything and a local good model as a librarian/assistant and be solar powered in a pinch.

cyberrock an hour ago | parent | prev [-]

Well it's close to the author's intention for Fahrenheit 451, but just not what everyone wants it to mean.

pyrophane 2 minutes ago | parent | prev | next [-]

Edging a little closer to the Krazam video "rare data hunters."

JumpCrisscross an hour ago | parent | prev | next [-]

Maybe we can kill two birds with one stone: digitize rare books and reverse the damage from Authors Guild v. Google [1].

Let AI companies do this. But require them to make the digital copies public. Maybe with a multi-year delay, to give the original scanner advantage to doing it.

[1] https://en.wikipedia.org/wiki/Authors_Guild,_Inc._v._Google,....

sixdimensional 44 minutes ago | parent [-]

Yeah, I was thinking along these lines...

Let's say one of the books to be digitized and destroyed is the sole remaining copy of a book from 1850, which is now considered public domain.

On one hand, hoarding such a book, stealing its content from the public domain, locking its content behind a for-profit machine, and destroying the only remaining copy is clearly wrong. It's equivalent to stealing a public resource, just like mining minerals or oil on public lands without a permit or mineral rights. Pure extraction.

On the other hand, taking care to digitize the copy and making it available for free in perpetuity, as well as being required through regulation to provide access to that content through, let's say a public utility LLM/AI available for free through libraries and online... and perhaps after fair due diligence being required to preserve physical copies in a public archive of rare books of which there are no known remaining physical copies...

That seems much more reasonable to me at least. I can imagine there are many who would not see it that way though. Do we see it happening or gaining regulatory, moral and/or public support?

s1artibartfast 28 minutes ago | parent [-]

I think this misconstrus what public domain is.

It provides a freedom to circulate, but not access to the material. It is not a public owned resource.

Turning a copy over to the public or state might be an interesting requirement for obtaining a copyright, but instituting that fix for new works now would have a 70 year lag time.

Think of it this way, if I copyright a book and put it in my dresser for 70 years, that doesn't give the public the right to access it or come into my house and scan it after expiry

aerodexis an hour ago | parent | prev | next [-]

Reminds me of Blood Meridian where The Judge meticulously sketches the rock glyphs that he comes across, and then destroys the original.

Now that I think about it, The Judge is an apt metaphor for AI : "Whatever in creation exists without my knowledge exists without my consent."

Cider9986 19 minutes ago | parent | prev | next [-]

Can only blame AI companies to a limited extent. This is apparently the legal way to do things because of stupid copyright laws.

Support your local shadow library: https://annas-archive.pk/donate

_m_p an hour ago | parent | prev | next [-]

Librarians were already doing this at scale in a process euphemistically called "weeding":

https://www.ala.org/tools/challengesupport/selectionpolicyto...

oliwarner an hour ago | parent | next [-]

This is a bizarre comparison.

Weeding is the natural process of disposing of less-demand books. Like the rest of us, libraries operate in finite space, so if they want new books, they have to remove ones their users aren't using. Most libraries will try to sell books before disposing of them in any destructive way.

What similarities do you see here?

Jweb_Guru an hour ago | parent [-]

People are very desperate to try to claim that something that's somewhat obviously morally wrong is actually highly nuanced, because it makes them feel uncomfortable.

streetfighter64 20 minutes ago | parent [-]

I would disagree on the "obviously morally wrong". What's the alternative for the books if they were "saved" from the fate of being scanned and shredded? How many rare books have you personally brought in the last year? Not everything that's ever created needs to be preserved forever.

malfist an hour ago | parent | prev | next [-]

This is really dishonest framing, unless you really, honestly can't tell a difference between pulping a mass market paperback romance novel that there's 3 million of in circulation, and shredding an 18th century botanical text that there's only 2 copies of in existence.

dwedge an hour ago | parent [-]

You call of dishonest framing, but you're begging the question twice.

Once that AI companies are really shredding 200 year old rare books, and once that libraries are only weeding mass market pulp fiction.

malfist 26 minutes ago | parent [-]

That is literally what this article is about.

streetfighter64 17 minutes ago | parent [-]

The "article" in question is a tweet. The "rare botanical text" is just an example the author of the tweet made up in order to generate sympathy.

malfist 11 minutes ago | parent [-]

The tweet is about a 404 Media article about the practice. It's linked elsewhere here.

emddudley an hour ago | parent | prev [-]

Libraries are not archives, and weeding does not mean destruction.

pu_pe 2 hours ago | parent | prev | next [-]

From what I understand, the rare books in question are not some historically relevant medieval manuscripts, but rather some relatively recent books (still under copyright) for which there are few print copies available for purchase.

I am not sure that physically destructing one copy of this type of book to preserve its contents digitally is so bad. Pretty much anything that is still under copyright should be valuable only for its content, not for the physical medium it's printed on.

godshatter 27 minutes ago | parent | next [-]

Is slightly modifying some weights in a markov chain somewhere really preserving it's contents?

storus 16 minutes ago | parent [-]

It's not really Markov chain as you need full P(x_t|x_{t-1}, x_{t-2}... x_1) instead of just P(x_t|x_{t-1}).

eru an hour ago | parent | prev | next [-]

> Pretty much anything that is still under copyright should be valuable only for its content, not for the physical medium it's printed on.

Well, there are special collectors editions with the signature of the author and gold pages or whatnot. But the AI companies are probably not using those.

lejalv 2 hours ago | parent | prev [-]

"To preserve them digitally"

For whom? is the relevant question

extra-AI an hour ago | parent | prev | next [-]

And this is the beginning of the end for content creators.

Why would I spend hours creating original content if Google can extract it and present the answer directly in an AI Overview? What is the incentive to keep doing the work?

If creators stop producing high-quality original material, the information we get over the next few years will increasingly be based on recycled, low-quality garbage.

secretsatan 43 minutes ago | parent | next [-]

It’s even noted they’re looking for text pre 2022 as afterward, it’s tainted by their own shit, they don’t believe in the crap they’re making.

Traster an hour ago | parent | prev [-]

Take a look at most successful journalism today. It's behind a paywall. You get paid by the people who are interested and value your work.

Can Google steal it and present it in an AI overview? Well kinda. Today Google is doing a trick - they're saying "You can refuse to consent to being fed into the slop machine, but if you do we won't crawl you for Google so you'll get no search traffic. But you're not going to get search traffic anyway! So you might as well opt out of being fed into the slop machine. And companies are starting to do that [1]

It's really interesting, because essentially what it means is Google is turning into a walled garden, but there's nothing growing inside it so they have to continually import new plants to live in their walled garden and they're going to have to pay to do that. So soon Google will be paying news sites for the right to plumb their feed into the slop machine.

[1]: https://www.wsj.com/business/media/google-search-publishers-...

Springtime 2 hours ago | parent | prev | next [-]

It seems a key contention of theirs is the possibility that rare books are being destroyed this way, yet the things they cite don't seem to suggest this (based on their paraphrasing), they just throw the following at the end to make it seem like it's occurring to irreplaceable books:

> You can't replace the last three copies of an 18th-century botanical text once someone shreds them for training data. And the judge said it's legal.

Is there evidence of this? Since otherwise they could very well be describing what is only occurring to in-print or non-rare books. (This is a genuine question since their post doesn't shed any light on it.)

D13Fd 2 hours ago | parent [-]

There is zero reason to shred 18th century books. Any such books are out of copyright.

samastur an hour ago | parent [-]

They are, but using the same process for all books is simpler and cheaper.

Cider9986 6 minutes ago | parent | prev | next [-]

It's not like these books were available to everyone before. If they are destroying one physical copy that's not accessible to the public and replacing it with a scanned copy that isn't accessible to the public, that's not a huge change. Except that will probably last longer digitized. Ideally they wouldn't destroy the physical copies, but this isn't as anywhere bad as book burning in Nazi Germany.

I find going after shadow libraries to be much worse because law enforcement is trying to prevent discrimination of knowledge to the public.

The real blame here should be going onto copyright laws.

Anthropic could take more care by figuring out if the books are still affected by copyright.

But this is just a company trying its best in an unfortunate regulatory environment.

Support your local shadow library:

https://annas-archive.pk/donate

thechao 2 hours ago | parent | prev | next [-]

Which book that was rare was destroyed? I'm interested to know a few titles.

Cynddl 2 hours ago | parent | next [-]

The 404media article mentions notably https://nltimes.nl/2026/06/25/rare-book-dealers-fear-tech-fi... which says:

> The attachment contained 3,000 English-language titles organized by ISBN number, including books such as Distinct Element Modelling in Geomechanics by K.R. Saxena (1999); Barrett's Traditional Fairy Tales (2021), an academic study of Irish folklore; and Laser Shock Peening of Advanced Ceramics by Pratik Shukla (2018).

infinite_spin 2 hours ago | parent | next [-]

> Barrett's Traditional Fairy Tales (2021)

How is a book from 2021 considered rare in this context? There's almost certainly a digital copy of it in existence prior to Anthropic purchasing a print edition.

ACCount37 2 hours ago | parent [-]

Niche text. It's not impossible that there was only ever under a thousand of them printed and released into circulation.

A digital copy would exist somewhere, of course. But for us, that only matters if we can buy or download it. And for AI companies, that only matters if they can get a digital copy DRM-free and licensed permissively enough.

sidewndr46 an hour ago | parent [-]

This argument doesn't make any sense. All manner of AI companies just ingest whatever random text they can find on the internet to train their data, including copyrighted publications. Why would DRM on a digital copy of a book matter?

streetfighter64 13 minutes ago | parent | next [-]

> Why would DRM on a digital copy of a book matter?

Because DRM is just a way to make "breaking copyright" more practically cumbersome. What's easier, breaking digital DRM for each and every E-book you find, or just establishing a single pipeline for scanning physical books?

eru an hour ago | parent | prev [-]

There might be special rules around DRM that go beyond normal copyright?

jdub an hour ago | parent [-]

Ladies and gentlemen, the Digital Millennium Copyright Act

(which is terrible, but I would be delighted if they breached it and got thoroughly spanked)

tokai an hour ago | parent | prev [-]

Non of those are rare. All a available in libraries for ILL.

andrepd 33 minutes ago | parent | prev [-]

Rare or not, destroying books is in itself a morally repugnant act. I don't know man, it's not so long ago we used to view nazi book burnings as an archetype of evil. Today companies are offering book-burning-as-a-service and it hardly causes a stir.

Bearded German man was right.

999900000999 an hour ago | parent | prev | next [-]

They should be forced to publicly release the books as an Ebook.

Leave it up for anyone to download and then compensate the copyright holders later.

In fact if ingesting these books for LLMs is fair use, us commoners should be able to read them for free. Maybe restrict commercial redistribution though.

streetfighter64 32 minutes ago | parent [-]

There's a bunch of details regarding LLMs and copyright, but I don't see how

> They should be forced to publicly release the books as an Ebook.

would be reasonable in any way? The books aren't theirs to release publicly. If I brought a copy of any given movie on DVD, ripped it and used it to train my own "LLM" located at /dev/null, should I then be allowed (or even forced) to release the movie publicly for anyone to watch for free?

999900000999 6 minutes ago | parent | next [-]

We’re in a brave new world now. If theirs reason to believe your destroying the last copy of a book( or another copy isn’t easy to obtain) then you should make a copy of it available.

Set up a compensation fund for the rights holders. Anything is better than culture literally being sucked into the void.

cj 21 minutes ago | parent | prev [-]

The alternative (throwing books away for no good reason) isn't much more reasonable.

writtenone an hour ago | parent | prev | next [-]

From 1984: "Every record has been destroyed or falsified, every book rewritten, every picture has been repainted, every statue and street building has been renamed, every date has been altered. And the process is continuing day by day and minute by minute. History has stopped. Nothing exists except an endless present in which the Party is always right."

Cider9986 16 minutes ago | parent [-]

Support your local shadow library: https://annas-archive.pk/donate

D13Fd 2 hours ago | parent | prev | next [-]

This is the result of our copyright law in the United States, which is extremely tilted to favor authors and publishers. The judge made exactly the right call and the companies are following the law.

The fix here is to change the law to permit training AI without destroying the original materials. But that is going to be a heavy lift.

coffepot77 44 minutes ago | parent | prev | next [-]

This was known a long time ago? https://arstechnica.com/ai/2025/06/anthropic-destroyed-milli...

altcognito 2 hours ago | parent | prev | next [-]

Is there any proof of this at all beyond this random message?

Ratelman an hour ago | parent [-]

Yeah - it does feel a bit overly dramatic, books mentioned from a potentially related article are from like 2018 (https://nltimes.nl/2026/06/25/rare-book-dealers-fear-tech-fi...). Let's not pump the drama more than we need to, Sam Altman made another mention of the singularity over the weekend so enough of that going around.

storus 19 minutes ago | parent | prev | next [-]

That's one way to pull the ladder if regulatory capture fails...

HelloUsername 23 minutes ago | parent | prev | next [-]

Are the AI companies also destroying the digital scans they made?

DarkIye 2 hours ago | parent | prev | next [-]

This is the opposite of a book burning. These books which only a few would ever know the names of, let alone find, let alone read, are being digitised so they can be found in electronic searches.

dpark an hour ago | parent | next [-]

> being digitised so they can be found in electronic searches

You make it sound like they are running a second Project Gutenberg. They most definitely are not making these available for electronic searches. At least not searches the public can participate in.

coffeefirst an hour ago | parent [-]

Correct. If they were doing this with Gutenberg/Smithsonian/some library so there's both a public archive and training the language model, it wouldn't have the ick factor.

imhoguy an hour ago | parent | prev | next [-]

And knowledge from these books can be "imprinted" into a model to be used by much more people or even survive this planet once sent into space in a probe.

swed420 an hour ago | parent | prev | next [-]

> are being digitised so they can be found in electronic searches

What guarantee do we have that the book contents will be served unfiltered and unaltered?

left-struck an hour ago | parent | prev [-]

Are they being digitised? That word implies that the book would have a digital representation of the original, which is not what an LLM is. Can they be found in electronic searches? I mean I kinda get what you mean, the knowledge was potentially forgotten and now it might not be, but also the authors who created that knowledge get no credit, no reward.

krunck 21 minutes ago | parent | prev | next [-]

In a society that values science and knowledge, preservation of knowledge - and no, slurping text up into an LLM is not preservation - is more important than profits.

These companies are regressive book burners.

hackernudes 2 hours ago | parent | prev | next [-]

Also discussed here https://news.ycombinator.com/item?id=44381838 from June 2025.

graemep an hour ago | parent | prev | next [-]

Is this true? Rare books would very often be out of copyright for a start. What the the actual ruling that says you can scan if you destroy the original? ISBNdb is a database of book information, as you would guess from the name.

roywiggins an hour ago | parent [-]

It's cheaper to destructively scan.

Arshad-Talpur an hour ago | parent | prev | next [-]

May be I am old schools but rare books must remain rare and LLMs shouldnt have access to those books

tasuki an hour ago | parent | prev | next [-]

Maybe we should not have forced AI companies to shred rare books so they can use them for training?

simonw an hour ago | parent | prev | next [-]

Here's the original story from 404 Media that this tweet is paraphrasing (without linking to, of course, because X disincentivizes links): https://www.404media.co/ai-companies-are-buying-tons-of-old-...

The segment that talks about rare books:

> One professional bookseller who specializes in selling foreign language books on these marketplaces told me that, starting in April, he and other booksellers noticed a historic spike in sales. [...]

> This bookseller said his inventory is full of rare, foreign language, and low circulation books, meaning that if they are destroyed in the process of becoming training data, they’ll be even harder to obtain.

qsera 2 hours ago | parent | prev | next [-]

I will make a robot scanner for books. I will then scan all the books in my state libraries and make a digital copy of them (without destroying them) before these things come for them.

I wish...

musha68k 2 hours ago | parent | prev | next [-]

At the very least why not upload the scanned books to the internet archive while already at it?

This is highly disturbing news; is this standard practice? What did Google Books do before?

qsera an hour ago | parent [-]

Yes, this is very disturbing to me. I recently discovered how great old, out of circulation/print can be. Now old abandoned libraries are like a treasure trove to me. If these books are not digitaly preserved, that sounds very sad to me.

johnxianren 2 hours ago | parent | prev | next [-]

I have zero proof for this, but just a what if: what if Anthropic's strict anti-China stance actually means the Chinese training corpus is way more valuable than people realize?

iamsaitam 33 minutes ago | parent | prev | next [-]

If you consider that AI companies should pay the same for a book, as a person does, you should read about royalties.

dandelioness an hour ago | parent | prev | next [-]

They change the info then destroy the source. Now the lie is in the LLM.

skybrian 2 hours ago | parent | prev | next [-]

It’s unclear whether ISBNdb will scan books without ISBN’s, which were invented in the late 1960’s. Customers appear to be ordering books to be scanned by ISBN? Here is one book seller’s experience:

> Bulk purchases also usually reflect interest in a specific topic, whereas the recent, very large purchases were of books that had little in common, except for the fact that they all had ISBNs. This seller also sells rare books that do not have ISBNs, and none of those were part of the bulk purchases.

Article is paywalled, but I saved a few quotes here:

https://skybrian-links.exe.xyz/post/1026

hagen8 41 minutes ago | parent | prev | next [-]

Where are the sources for that?

petesergeant an hour ago | parent | prev | next [-]

So the premise here is that the books are valueless enough that they're being sold by weight, but they're also rare enough that someone will one day wonder where they went, and also that the AI companies aren't also saving the text somewhere much safer than physical media in a warehouse. k.

regnull an hour ago | parent | prev | next [-]

The source for this is "I've heard it from some guy". Before we freak out, perhaps make sure it's actually happening?

bryan_w 31 minutes ago | parent [-]

This is quite a reasonable take.

azan_ 2 hours ago | parent | prev | next [-]

The comments there are absolutely unhinged. There are some good reasons for being anti-AI, but why dilute it with this kind of bullshit:

> It is equivalent to book burning in the past. A form of thought control

nhinck3 an hour ago | parent | next [-]

You're right, it is arguably worse than book burning, not only are they seeking to deprive others of the books, they also want to profit off it.

swed420 2 hours ago | parent | prev [-]

> why dilute it with this kind of bullshit:

> > It is equivalent to book burning in the past. A form of thought control

That's only bullshit if you trust AI companies to serve the book contents without alteration.

dpark an hour ago | parent [-]

I don’t trust or expect AI companies to serve books at all. That’s not what they are scanning them for.

swed420 an hour ago | parent [-]

Not in a traditional sense, but obviously on the surface, they're using the info to regurgitate in some fashion and serve back.

The point is that even under the best intentions, hallucinations occur. Then there's the fact that most models have an ideological bias programmed into them.

The only expectation I have is for companies or anybody else to not destroy rare books. Is that such a tall order?

dpark an hour ago | parent [-]

> they're using the info to regurgitate in some fashion and serve back.

Sure, in the same sense that they regurgitate any other text they consume. LLMs by definition do not have the full training dataset available, though. It’s far larger than the resulting model. So they can’t reliably reproduce full text without an external source (or if it’s in the training data repeatedly). ChatGPT actually refused to give me a bible quote the other day, presumably because I ran into some general “book regurgitation” safety net.

> The only expectation I have is for companies or anybody else to not destroy rare books. Is that such a tall order?

Honestly, yeah. The idea that people or corporations should hold onto books forever because of a cultural “ick” about throwing out books is a bit ridiculous. Most books end up in landfills.

They aren’t feeding Da Vinci manuscripts into this pipeline. They are feeding still-in-copyright books.

swed420 23 minutes ago | parent [-]

> LLMs by definition do not have the full training dataset available, though.

That makes it even worse, then. This proves the original point.

> The idea that people or corporations should hold onto books forever because of a cultural “ick” about throwing out books is a bit ridiculous.

If we're building black and white straw man arguments, then sure, let's not archive anything.

tannerr_dev 23 minutes ago | parent | prev | next [-]

absolutely diabolical

paxys an hour ago | parent | prev | next [-]

People keep bringing up these rare books but never share what they actually are. What are their names? When were they written? How many copies were in existence? Were they in libraries, or locked up in vaults? Did people have access to them prior to being scanned for AI? How many such books have actually been destroyed?

Weird to see so many of these "trust me bro" twitter stories make it to the front page and cause outrage when no one has any real information.

qsera an hour ago | parent | prev | next [-]

All those AI powers, and they cannot find a way to do it without destroying the books? Pathetic.

cluckindan 36 minutes ago | parent | prev | next [-]

This is not that different from having a book burning march. The fervent neocon tech overlords and their funding circles want a monopoly; not only on truth, but on ideas and thought in general.

enaaem 2 hours ago | parent | prev | next [-]

This is the destruction of Western civilisation.

mindslight 42 minutes ago | parent | prev | next [-]

The design of copyright has always been to restrict the dissemination of knowledge so that somebody can turn a profit. The handwaved justification is that the profit encourages the creation of new knowledge whose dissemination can be restricted, but that still doesn't erase the fundamental dynamic.

This is merely the latest incarnation. We can imagine a slightly different process on a few fronts - AI companies pay to digitize books (still for their own purposes), but are prevented from destroying the physical copies and they have to openly shared the digitized results. We would view that situation much more favorably - perhaps even as ideal, right?

Those two dynamics could be backed up by court decisions or laws iff they weren't so plainly at odds with how copyright has been and is generally implemented and interpreted. For example, imagine them having to do this through some nonprofit library whose goals was preservation and dissemination. Instead, libraries have been sidelined as things that operate at the edge of the law rather than vital public institutions, whereas shredding books in secret is fully legally condoned.

Cider9986 12 minutes ago | parent [-]

This is the correct take.

Copyright has done more than anything else to prevent preservation and dissemination of knowledge. And it's forcing Anthropic's hand now. Although they could take more effort to preserve the books.

tokai an hour ago | parent | prev | next [-]

I don't buy this at all. Looking up some of the mentioned rare books, in other sources, on worldcat shows every single book available somewhere. Most of them in dozen or even hundreds of libraries.

api an hour ago | parent | prev | next [-]

People would have at least somewhat less of a problem with this if they also put up an archive of PDFs of all these rare books if they are out of copyright.

But that would help competitors with training data, which I assume is why they don’t do this.

freejazz an hour ago | parent | prev | next [-]

Why would you need to shred a book from the 1800s when it is in the public domain? Something is fishy here.

nailer 33 minutes ago | parent | prev | next [-]

Direct link: https://x.com/HedgieMarkets/status/2081534588485296565

1234letshaveatw an hour ago | parent | prev | next [-]

you wouldn't believe how much shredding your local library does in the name of space conservation and to address changing borrower preferences and demographics. I doubt any AI company comes close to the annual combined library turnover

pfdietz 30 minutes ago | parent [-]

Where I live there's a large library-associated biannual book sale where books are (effectively) reverse auctioned over a period of a few weeks. At the end, anything (with a few exceptions, like the Collector's Corner) that isn't sold is disposed of, with a large 18-wheel truck sized dumpster filled with items to be sent for pulping and recycling. The price at the end is $1 for a grocery bag full of books, so things that don't sell truly are perceived as worthless.

https://booksale.org/

Books are information delivery vehicles. We mostly shouldn't care about them any more than we care about a particular set of bits on a disk.

Publishers also pulp large numbers of books themselves. This is a consequence of the Supreme Court's Thor Power Tools ruling, which clarified tax rules in the US so that keeping large inventories of unsold books was less economical.

Good4boothee 2 hours ago | parent | prev | next [-]

> A federal judge ruled the practice is fair use because eliminating the original means only one copy exists at a time.

Really? That sure wasn't a thing when one startup got sued for streaming from its wall od dvd-players, and they adhered to 1 disc = maximum 1 stream at same time.

fmaccomber 2 hours ago | parent [-]

That's not an accurate characterization of the ruling

retinaros 2 hours ago | parent | prev | next [-]

if they do this you can foresee what else they can do

azan_ 2 hours ago | parent [-]

What else they can do based on this?

qsera 2 hours ago | parent [-]

If they are destroying old books, then it shows where their values lie...

azan_ 42 minutes ago | parent [-]

Where? Could you stop with the vague posting?

qsera 13 minutes ago | parent [-]

Could you stop playing dumb?

ChrisArchitect an hour ago | parent | prev | next [-]

Source: https://www.404media.co/ai-companies-are-buying-tons-of-old-...

impsunrise an hour ago | parent | prev | next [-]

ironically, this tweet and account appear to be complete AI slop

bebe8393jrir an hour ago | parent | prev | next [-]

What if dog pees on and shreds rare book? Like me home work? The only copy im existence!

vessenes 2 hours ago | parent | prev | next [-]

I don't see any proof of shredding here. Most book scanners I'm aware of are from Google's scanning days, and those had cameras plus page turning.

If we use 'shredding' to mean A book is laid flat, its cover is removed, and then a paper cutter cuts through the binding to create a flat stack of sheets, which are then fed to a sheet feeder, then I could maybe imagine this is better than a page-turning scanner. But, sheet feeding old paper sucks shit, bro. It's not fun.

Upshot, I think we'd like to hear from an anonymous frontier lab employee here to see what's going on -- there are a lot of books in Anna's archive available at considerably less difficulty.

Jolter an hour ago | parent [-]

The idea here is that the frontier labs are no longer willing to risk using the likes of Anna’s Archive, because they have already been held legally liable for that in the past.

vessenes an hour ago | parent [-]

Are you referring to the META lawsuit? I think the current landscape is not bad for the frontier labs -- it's settled law that it's legal to 'read' and ingest this data. Anthropic went ahead and just settled a licensing deal for content. The open issue in that META suit is whether or not any distribution happened, as I understand it. I'm certain they all have full backups of the archive somewhere in the org.

timcobb 2 hours ago | parent | prev | next [-]

This kind of reads like a blood libel. My guess is they're buying all those books that University libraries are throwing away these days (to see other HM threads for that), kinda sad but probably aren't "rare books" in the way people are thinking

iandanforth 44 minutes ago | parent | prev | next [-]

If true this is basically unforgivable. The wanton destruction of history is the stuff of the Third Reich and the Taliban. You simply cannot profess to care about knowledge, culture, or civilization while destroying the physical manifestations of the same.

pfdietz 39 minutes ago | parent [-]

Yes, actually you can. The fetish of physical book worship is a fossil of an age when information storage and retrieval was much harder.

BoxOfRain 35 minutes ago | parent [-]

I strongly disagree with this take, a book might remain readable thousands of years from now but very few if not zero of our digital data formats likely will. We shouldn't be so quick to throw away diversity in the way information which may be useful for future generations is stored for the long term.

pfdietz 25 minutes ago | parent [-]

Paper doesn't easily survive for thousands of years.

You know that pleasant used bookstore smell? It's paper slowly decomposing.

stuartjohnson12 2 hours ago | parent | prev [-]

I think people are, on the whole, too precious about old things. In the case of books produced after major commercial printing began, I don't believe it is the paper that imbues the book with historical value.

Indeed, I think there's a high chance that this process increases preservation of the most relevant part of the media - the actual content!

There are tons of old rolls of film slowly rotting away in warehouses that were never digitised. Even for beloved media, the BBC occasionally tracks down an old lost episode of Dr Who.

For now these books are in corpuses of training data, but eventually I trust they will make their way to the rest of us.

Cynddl 2 hours ago | parent | next [-]

> For now these books are in corpuses of training data, but eventually I trust they will make their way to the rest of us.

What makes you think they will? What would be the incentives for these companies to do so?

stuartjohnson12 2 hours ago | parent [-]

Well, you probably weren't going to go and find any rare, non-digitised books to physically go and read (unless you were going to, in which case, rock on), so we can start by benchmarking relative probability there.

1. At some level of critical information-withholding mass, a leak or disclosure similar to SciHub is inevitable because of the commonly held opposition to hiding knowledge.

2. Availability via Google Books or similar.

3. Availability via AI model reference.

4. Failing any of the above, better AI models that are more capable of doing more things, at the expense of books that were likely to go unread (revealed preference, rare books are often rare for a reason). This will obviously be a nonstarter if you don't want this to happen, but I think it would be good for the world if it did.

I think category of old books that were going to be read or otherwise become important parts of human knowledge that have not yet been digitised and now will never become so because they are instead being shredded and will never make their way into the light because of AI company data hoarding is a small category.

b3lvedere an hour ago | parent | prev | next [-]

https://www.youtube.com/watch?v=RH96S4rBZbI

breezybottom 2 hours ago | parent | prev | next [-]

That's only true if you think digital storage has more longevity than paper. It doesn't.

croes 2 hours ago | parent | prev [-]

So if the Mona Lisa is part of a model we can burn it?

stuartjohnson12 2 hours ago | parent | next [-]

I pre-empted this - my argument does not apply to texts where the physical object is a major part of the historical value of the thing. No, I'm not saying to destroy one of the four remaining Magna Cartas that were meticulously copied by hand. But even if I was, we're only dealing with texts here that are irrelevant enough to have never been digitised already - we tend to digitise most things of value and so the Mona Lisa and Magna Carta would never have been part of this discussion in the first place.

I am however OK with destroying one of the remaining 50 children's books of which only 300 copies were ever printed in a small town in Ohio in the 1970s as a test run for a failed book which was subsequently never commercialised.

croes an hour ago | parent [-]

> I am however OK with destroying one of the remaining 50 children's books of which only 300 copies …

What if the perception of those books change over time and are considered masterpieces later on?

Moby Dick was out of print when Melville died 1891 and not a huge success until it got a revival in the 1920s

stuartjohnson12 44 minutes ago | parent [-]

When was the last time a formerly undigitised book that was printed but fell out of print suddenly became a success after being discovered in the last, say, 50 years? Is once in 100 years the sort of level of frequency we're talking about?

This is the kind of rationalization hoarders use. The inability to get rid of things because it could turn out to be something we want in the future for reasons that we cannot currently describe.

It's a loss avertive instinct that I think is misplaced. Treating every printed book as priceless is intractable. It's not how we treat these books at the moment. Apparently, today we don't even care enough to spend a few hours per book nondestructively scanning them in.

Let's say we, instead of destructively scanning this books, nondestructively scanned them. What would you propose doing with the copies afterwards? Sell them? To who? They're valueless individually for the overwhelming part. Warehouse them? Why? For who? Do you want to go and look through them? Why haven't you done so already? Have you ever shown interest in consuming an undigitised book a single time in your life to date?

It feels like the anti-AI crowd here have to tie themselves in knots here to make the loss minimization work.

infinite_spin 2 hours ago | parent | prev [-]

If you purchased the Mona Lisa (or some rare book), in this hypothetical, you can burn it.

croes an hour ago | parent | next [-]

There is a difference between legal and right.

That’s the whole point because the book shredding is already declared legal.

infinite_spin 13 minutes ago | parent [-]

You asked a question about whether you can do something, not whether I found the behavior moral. I'm not interested in moral debates.

Invictus0 2 hours ago | parent | prev [-]

> For instance, a painter may insist on proper attribution of their painting, and in some instances may sue the owner of the physical painting for destroying the painting even if the owner of the painting lawfully owned it.[1]

https://en.wikipedia.org/wiki/Visual_Artists_Rights_Act

infinite_spin 2 hours ago | parent [-]

The Mona Lisa's painter isn't alive, they can't sue, and this act doesn't apply to printed books. You're expanding the hypothetical to say "also you're under a specific set of laws that forbid exactly the thing I questioned".. that's moving the goalpost.