| ▲ | cladopa 11 hours ago |
| It is not a big deal. Since the invention of the printing press any important
book has been duplicated by thousands, tens of thousands or even million of units. Just taking one of those and "destroying them"(it is not destroyed, a digital copy with the ability of doing millions of copies is stored somewhere) is not problematic for Humanity. By the way, I always search for second hand books. Most of the books there are garbage. Most people clean their shelves with the books they don't care about, but preserve the ones that are good. If they are young people that inherited a house and don't care about books, they pick and sell the good ones, giving away the bad books. If you go to a recycling centre, the garbage to quality ratio is over 100 or more. That is, for every 100 books that are garbage there is one good quality book. It is very rare to find a jewel there. |
|
| ▲ | nloomans 10 hours ago | parent | next [-] |
| > a digital copy with the ability of doing millions of copies is stored somewhere somewhere were we can't access it. as the article states: “permanently locking human knowledge inside private corporate servers” the issue isn't that the physical copy is gone, it's that they are preventing people from making digital copies that are actually accessible by destroying the physical copies. > If you go to a recycling centre, the garbage to quality ratio is over 100 or more. archivists keep everything, because we don't know right now what will be important 100 years from now. |
| |
| ▲ | radu_floricica 10 hours ago | parent | next [-] | | > somewhere were we can't access it By any metric imaginable, it's making the information more accessible, not less. First, it's taking a single copy of a 10k physical print and it's making it digital. Is it "locked"? Yes, by copyright laws, if you don't like that lobby to have them changed. But it's _closer_ to being widely available, not farther. Plus having the info part of a LLM makes it immediately available to literally billions. I happen to actually actively shop second hand bookstores, so I am potentially affected by this - as opposed to most people complaining because they don't like the idea. And I still absolutely support it. | | |
| ▲ | jjulius 10 hours ago | parent | next [-] | | >Plus having the info part of a LLM makes it immediately available to literally billions. Help me understand how. Not only are these LLMs expressly prohibited from specifically regurgitating copyright works if the users asks them to, but they habitually hallucinate or paraphrase things wrong. If they won't regurgitate the copyrighted text verbatim, and are known to be confidently incorrect and hallucinatory, I'm struggling to see how these texts are "immediately available to literally billions". And I ask this as someone who has had LLMs give me incorrect assertions about the contents of books. | | |
| ▲ | SkyBelow 10 hours ago | parent [-] | | Depends upon what you want. For example, knowledge about how the book smells when you open it, something that reader do talk about enough I don't think this should be a strawman, is lost. But, that is about the experience of reading the book, not the knowledge of the book. The exact text? Yeah, I think that is largely lost as well. This is a summary. And for rarer books, it will be a particularly bad summary. The basics of the book are being captured in a space that the right question that retrieve it, but worse than a sparksnote and any well read reader will tell you all the sorts of things a sparknotes already loses compared to reading the book directly. But, that little bit of data is a bit more data than existed before, and future LLMs should get better at giving the information. So, in that sense, the knowledge is better being spread compared to copyright where the book stays in a warehouse until it is disposed of. If it was between this and a sparksnote of the book being made, the sparksnote is far better, but between this and the book simply being disposed of, then the LLM is better but far, far from great. That's a lot of assumptions that goes into the judgment, which is probably why different people reach different conclusions. One person is imagine the alternate fate of this book being slowly rotting in a landfill, the other resting on a bookshelf where it is read at least once fully and then flipped through time to time, and neither are wrong. | | |
| ▲ | chefandy 7 hours ago | parent [-] | | > But, that little bit of data is a bit more data than existed before, No, it’s not. Undigitized data is still data. Wording, style, nuance, type, artwork, binding, metadata, attributions, citability… all of these things are permanently lost, which is a fucking tragedy because they don’t have to be. Even if you’re not willing to take the time to scan every page in a cradle scanner, as rare books should be, (and don’t tell me they don’t have the money to get a handful of library interns to do this,) you can disbind the books and store them as they did in the Caselaw Access Project at Harvard Law. They removed the pages from the binding, scanned them on a high speed conveyor belt scanner which yielded full color 600 DPI jp2 images, placed the pages back in the binding like a folio that could be re-bound if needed, vacuum sealed them, and stored them in a salt mine. It’s not like it was slow, either — we did 40k in 18 months and we did take the time to scan the rare ones with a cradle scanner. And we did it all in less than open AI probably spends in a day on inference. > So, in that sense, the knowledge is better being spread compared to copyright where the book stays in a warehouse until it is disposed of. That’s a false dichotomy. Libraries exist for this exact reason, and their not already having a copy does not make “you snooze you lose” a morally acceptable strategy. I’ve been pretty cool on the direction of SV for the past decade at least, but I am absolutely gobsmacked by the unbridled hubris of these companies over the past 5 years. | | |
| ▲ | nuancebydefault 6 hours ago | parent | next [-] | | > Undigitized data is still data. Wording, style, nuance, type, artwork, binding, metadata, attributions, citability… all of these things are permanently lost I understand what you mean but... "permanently lost" sounds dramatic. When I trow away old pictures, old drawings or pieces made by my son at school, they are lost as well. Not that I do that often, but it begs the question, should every 'ip' made by humans be preserved? | | |
| ▲ | chefandy 2 hours ago | parent [-] | | It sounds dramatic because it’s dramatic. You’re combining two things — whether something is, in fact, completely lost, and if it is something that should be kept. Something that should not be kept is still completely lost if it’s destroyed. It’s difficult to imagine they’d digitize it if it was worthless. |
| |
| ▲ | SkyBelow 6 hours ago | parent | prev [-] | | >all of these things are permanently lost A small fraction of them is saved in the model. Far more is saved in the digitized copy as long as they keep it which they have plenty of incentives to do so (future training of newer models). That's more than what happens if that book was burned or sent to a landfill, but less than if the book is giving a loving home. >They removed the pages from the binding, scanned them on a high speed conveyor belt scanner which yielded full color 600 DPI jp2 images, placed the pages back in the binding like a folio that could be re-bound if needed, vacuum sealed them, and stored them in a salt mine. My understanding is that this simply isn't legally allowed for these books. The original must be destroyed for the digital copy to not be copyright infringement. >That’s a false dichotomy. I pointed out there is a spread of possible outcomes and that different people are considering different outcomes and the comparison of if this is good or bad depends upon which outcome one considers. I even mention that both outcomes are sometimes right. That's about as far from a false dichotomy as I can see it. |
|
|
| |
| ▲ | choo-t 3 hours ago | parent | prev | next [-] | | > Is it "locked"? Yes, by copyright laws, if you don't like that lobby to have them changed.
Or be useful and scan book for Anna archive and other shadow libraries. Citizen's lobbying against megacorp is a mirage. >But it's _closer_ to being widely available, not farther. By what metric ? The copy is now guarded by a company instead of being on the second hand market. | |
| ▲ | winterismute 7 hours ago | parent | prev | next [-] | | > Is it "locked"? Yes, by copyright laws, if you don't like that lobby to have them changed. I don't if that is true: a lot of old books might still have copy or other rights associated to them, likely owned by author and/or publisher, directly or inherited, but often those who have the rights do not have digital or physical copies at hand anymore (some old books are, well, really old). Does Anthropic make sure to track down, contact and then share the digital copy they make with those who have rights on the work? If not, they are not making it in any way easier to re-print the books, while making their supply more scarce (they destroy existing embodiments). | |
| ▲ | darkwater 10 hours ago | parent | prev [-] | | > Is it "locked"? Yes, by copyright laws, > Plus having the info part of a LLM makes it immediately available to literally billions. Isn't this a contradiction?
I mean, maybe you can invoke a fair use policy if an LLM spits out some text from the scanned book, but then you are _not_ "making it available to literally billions". |
| |
| ▲ | gnfargbl 10 hours ago | parent | prev | next [-] | | > permanently locking human knowledge inside private corporate servers History tells us that very few "permanent" situations are truly permanent. Provided a set of information has value (which in this case it clearly does) then the overwhelming likelihood is that, eventually, through some method or other, the information will become public. | | |
| ▲ | darkwater 10 hours ago | parent | next [-] | | > History tells us that very few "permanent" situations are truly permanent. If you destroy the only copy of a physical artifact, the situation is as permanent as it can get. | | |
| ▲ | pixl97 4 hours ago | parent [-] | | I mean books get destroyed all the time, really they are a major pain in the ass to keep together, especially as they age. Paper loves to crumble. Insects think they are tasty. Floods and fire love destroying them too. So physical books are rather non-permanent themselves. |
| |
| ▲ | mohamedkoubaa 7 hours ago | parent | prev [-] | | Exactly. Calling anything on an SSD permanent is criminal. |
| |
| ▲ | mvlipwig 7 hours ago | parent | prev [-] | | I wonder if Anthropic could rent or sell access to their collection to the internet archive? It would probably be a good PR move (which they probably need right now), but I'm not sure what type of legal shenanigans they would need to do in order to not violate copyright law. |
|
|
| ▲ | TFNA 7 hours ago | parent | prev | next [-] |
| > any important book has been duplicated by thousands, tens of thousands or even million of units. Books in the former USSR display their print runs on the last page. "Important book" is a vague and arbitrary term, but rhere are works in whole fields (e.g. history, archaeology, linguistics, ethography) that any scholar would consider key references, and as few as 100 copies were printed. The shadow libraries have made a lot available to the whole world. It would suck if private corporations scan and shred remaining copies of these before the shadow libraries can get a scan. |
|
| ▲ | afpx 10 hours ago | parent | prev | next [-] |
| I think you may be greatly underestimating the long tail. Several times a year I read sources that reference older books that I can't find online. When I am able to locate them, they often cost at least several hundred dollars, sometimes into the 10s of thousands. |
| |
| ▲ | quietsegfault 5 hours ago | parent [-] | | Do you think that you are somehow special and unique in needing these books? If the books cost in the 10s of thousands, then there's obviously value to other people. I've seen no evidence that Amazon or others are buying $10k books to scan into their corpus. All evidence I've seen is that they're scanning cheap books with no current value and no clear use to people today. I have volunteered with a library, and probably threw hundreds of books over a couple week engagement from a university library into a shredder at the direction of a professional, academic librarian. Libraries are constantly culling books, the EXACT category books we're talking about here (old, never-read). This is happening at a much larger scale, so I would recommend railing against university librarians in addition to the AI juggernauts. | | |
| ▲ | hughlilly 14 minutes ago | parent [-] | | > I've seen no evidence that Amazon or others are buying $10k books to scan into their corpus. Have you seen evidence that they’re buying only readily available books that are plentiful on the market? |
|
|
|
| ▲ | mannyv 9 hours ago | parent | prev | next [-] |
| I have books, but they are just objects. They're nice objects, but just objects. Fetishizing books isn't going to help. In fact, most of those "rare" books don't sell because nobody wants them. The AI companies are making them even more rare, so the booksellers should be thankful. |
|
| ▲ | sajithdilshan 10 hours ago | parent | prev | next [-] |
| Exactly, also all those physical books would anyways get molded, eaten by moths or just naturally decay. It's not like the AI companies are obliterating every copy of every single book. |
|
| ▲ | NoMoreNicksLeft 18 minutes ago | parent | prev | next [-] |
| >Since the invention of the printing press any important book has been duplicated by thousands, tens of thousands or even million of units. No, every popular book has been duplicated thousands of times. This is not the same thing as important. They're orthogonal. When an important book is popular, it is safe. When it is not, it is in danger. Only fools assume that important books are recognized often enough to become popular. |
|
| ▲ | brightball 10 hours ago | parent | prev | next [-] |
| Whenever my wife wants to visit antique stores, I always look for old books. I have found several 100+ year old gems. |
|
| ▲ | shiandow 10 hours ago | parent | prev | next [-] |
| Somehow I don't think they're looking for the books that have been copied over and over. |
| |
| ▲ | quietsegfault 5 hours ago | parent [-] | | Why do you think that? Do you have evidence, or is this just a hunch? Why would Amazon waste money on uber rare books when there are thousands and thousands of not-so-rare books that could serve the exact same purpose? | | |
| ▲ | shiandow 4 hours ago | parent [-] | | For one they've already used the entire library genesis. Anything not in there is going to be obscure in some capacity. |
|
|
|
| ▲ | pibaker 7 hours ago | parent | prev | next [-] |
| > any important book has been duplicated by thousands, tens of thousands or even million of units. It is common for academic books to have publication runs in the low three digits. You may argue these books are not important. But how do we know if we fail to preserve it? |
| |
| ▲ | quietsegfault 5 hours ago | parent [-] | | Is there evidence that these mythical low-print-run books are being purchased by Amazon and friends for destructive scanning? I simply don't see why railing against the AI giants about this without also contributing significant effort into protecting these low-print-run books from getting jettisoned by university libraries makes any sense. |
|
|
| ▲ | arttaboi 6 hours ago | parent | prev | next [-] |
| With all due respect, I would say it wouldn’t hurt not to downplay this. |
|
| ▲ | jll29 10 hours ago | parent | prev | next [-] |
| Beware that the notion of "quality" is entirely different for AI companies: they don't seek entertainment, but sentences in a language to train an LLM. |
|
| ▲ | GreenLightGo 7 hours ago | parent | prev | next [-] |
| Honestly, it’s easier to find a good movie than a good book, because books are way cheaper to publish. These days, the quality of pretty much all kinds of content has become a problem... |
|
| ▲ | mtkd 6 hours ago | parent | prev | next [-] |
| >It is not a big deal have you ever held and read an old book? |
|
| ▲ | rvz 10 hours ago | parent | prev | next [-] |
| First of all, it IS destroyed and it is a big deal. Hardcover copies of books especially 1st - 2nd edition ones (even with mistakes) are rarer than digital scans. Maybe the Bodleian Library at Oxford University should give all their rare books to AI companies to scan and destroy them since it is not a "big deal" anyway. Except that when they did do a pilot with OpenAI to scan these rare books, [0] they did NOT destroy them. I wonder why? [0] https://www.bodleian.ox.ac.uk/services/research-partnerships... |
| |
| ▲ | kccqzy 9 hours ago | parent | next [-] | | Why wonder? The answer is abundantly clear if you follow the news. If a book is copyrighted under U.S. law, scanning and destroying counts as a format conversion which qualifies it as fair use, so there is no need to negotiate with copyright holders. See Judge William Alsup’s decision. If Anthropic did not destroy the books after scanning it would have not won the lawsuit, and scanning would be illegal. If a book is already out of copyright then of course they do not have to destroy it afterwards. | |
| ▲ | 8 hours ago | parent | prev | next [-] | | [deleted] | |
| ▲ | quietsegfault 5 hours ago | parent | prev [-] | | Why is it a big deal? Do you think there's no difference between the books curated at Oxford University and the crap that Amazon is buying? | | |
| ▲ | voidhorse 2 hours ago | parent [-] | | Why would amazon buy "crap"? Surely they want their model to succeed and they want to train it on valuable input, no? They have more than enough resources to determine whether or not the books are worth buying. They have been in book selling for a long time. |
|
|
|
| ▲ | pshirshov 10 hours ago | parent | prev [-] |
| Read more about this. Depends on the definition of the "big deal" but from what I can understand the problem is that they buy rare things - which exist in just several copies - and they tend to buy _all_ copies. |
| |
| ▲ | wasmperson 6 hours ago | parent | next [-] | | I was also skeptical of this claim but managed to find someone who explains it: https://downtownbrown.substack.com/p/five-fallacies-ai-and-d... It's not that individual companies buy all copies of a given book, but that there's more than one book scanning company, and they aren't sharing the scans with each other. The result: books that were rare but nevertheless easy to find for purchase (thanks to the internet) are now vanishing off of the market, becoming de facto no longer accessible to the public. | | |
| ▲ | demibabs 4 hours ago | parent | next [-] | | Good article, but I still feel unsatisfied because even it cannot find an example of a book that’s actually been lost because of the destructive scanning frenzy (it only lists books that hypothetically could be lost because there’s not many physical copies available for sale online.). If anyone has an example, I’d love to hear it. | | |
| ▲ | voidhorse 2 hours ago | parent [-] | | Since we don't know what was actually purchased and what was actually destroyed, how do you expect us to furnish an example? This would require the destroyers to admit it, and beyond that it would require all of them to admit it since more than one of them might have been responsible for the extinction of one text. Seeing as they were already keeping this operation under wraps, I don't see that happening. "possibly extinct because no copies available online" is probably the best we can do. The distributed nature of the problem and the utter lack of transparency are huge factors here too. |
| |
| ▲ | quietsegfault 5 hours ago | parent | prev [-] | | [dead] |
| |
| ▲ | pfdietz 8 hours ago | parent | prev | next [-] | | Here we have another entry in the long list of "things described on the Internet that never happened". | |
| ▲ | mistercow 10 hours ago | parent | prev | next [-] | | Most old books that are rare and unpreserved are so because their value is marginal, so nobody has bothered to collect and preserve them. But where did you hear that they’re buying “all copies”? And to what end? | | |
| ▲ | dbspin 10 hours ago | parent [-] | | This is a classic mistake. We have no way of estimating the future value of a given book. It's perceived current value (a large part of which is simply obscurity) may be low. But it's future value - to historians, ethnographers, to researchers seeking a specific fact or example of language use or a hundred other things - is literally inestimable. To take a crude example in a different medium - new york in the 90s - widely documented right? Yet, if you want to find high definition video of street life in a given burrough on a given day or year, you're faced with an enormously difficult task. There were some HD test videos done in Manhattan in the late 90s (which have been posted to Hackernews before), but there's no equivalent for the other burroughs. Your best best would be finding original negative out takes or location scouting footage from feature films, a very hard task. That's only 30 years ago. Outside of the focal points of the worlds attention - English language, rich countries, places in the news, contemporaneous sources for 'non notable' events (lifestyle, how people spoke dressed etc) is surprisingly poorly preserved. Hopefully you can infer how this tracks to the written world and primary sources for language, technical manuals etc etc. | | |
| ▲ | famouswaffles 9 hours ago | parent [-] | | >Hopefully you can infer how this tracks to the written world and primary sources for language, technical manuals etc etc. I honestly can't and I think you can't either or you would have used an example with books/printed media rather than film, an entirely different ballgame. | | |
| ▲ | dbspin 9 hours ago | parent | next [-] | | OK... I'm going to assume good faith even though your wording makes it somewhat unlikely. Similar textual examples would be any text containing actual language as it's spoken in a given place or time. Or any factual textbook detailing the buildings present in a given location. Or any text book detailing a now defunct construction process. Or any text book (generally small run) detailing a niche interest, now missing ecosystem or the state of a particular political situation at a given time. Essentially all textual primary sources for events which are not currently considered important - but which we have no way of estimating the future importance of. One can continue to create countless counterfactual examples in this vein. My overall point is we cannot know what may be useful or even essential in the future, and knowledge should be preserved under the assumption that it is likely to be. Historians frequently refer to this paradox - how everyday aspects of life are frequently not explicitly documented, since they're so obvious to the communities or communities of expertise that observe and carry them out. So it's actually incredibly important to preserve what seems like ephemera. Hell we couldn't have AI training at all if we lacked the corpus of existing written literature - but there was no way any author could have anticipated this future utility more than a couple of decades ago. | | |
| ▲ | famouswaffles 7 hours ago | parent | next [-] | | I think there are two issues here: 1. If someone is acquiring books in bulk for bargain-bin prices and shredding them, they're books whose physical copies have essentially no market value and which, absent this buyer, were overwhelmingly headed for pulping or landfill anyway. Millions of books are destroyed every day. Could one of these worthless looking books turn out to contain information historians care about in a 100 year? Sure. But that doesn't create an obligation for someone to pay to warehouse every extant copy forever. Physical Preservation has costs: space, cataloguing, handling, transportation etc. Archives and libraries have always had to make choices for this reason. 2. I'm not arguing that preservation has no value. The question is whether destroying a physical copy after digitizing it is a serious loss when talking about mass-produced printed material. Was this the last surviving copy ? Is the information unavavilable in libraries, archives, other editions, scans, citations, contemporary works etc ? If not, nothing has been lost except one physical instance of a reproducible object. Your film analogy worked a lot better because old film footage is often unique primary source material. A camera recording of a random brooklyn street in 1993 may literally be the only recording of those people, storefronts and circumstances. The nth printe dcopy of a technical manual is not analogous to that. | |
| ▲ | 8 hours ago | parent | prev | next [-] | | [deleted] | |
| ▲ | 8 hours ago | parent | prev [-] | | [deleted] |
| |
| ▲ | voidhorse an hour ago | parent | prev [-] | | How would film be any different in any capacity whatsoever? Believe it or not, what matters here is the message and access to the message, not the medium. |
|
|
| |
| ▲ | dataflow 10 hours ago | parent | prev | next [-] | | Where did you see they tend to buy all the copies? This comment is the first time I've heard of this. | | |
| ▲ | wmeredith 10 hours ago | parent | next [-] | | I'd also be curious about the provenance of that statement. Why would they buy all copies? What would be the purpose of scanning multiple copies? | | |
| ▲ | dataflow 9 hours ago | parent [-] | | I could see buying multiple copies being useful to mitigate problems, like damage. But buying all the copies is categorically different and I cannot imagine why they would attempt that, except perhaps to prevent their competition from getting a hold of the same text? |
| |
| ▲ | p-e-w 10 hours ago | parent | prev [-] | | It’s just another lie of the type these threads tend to be filled with nowadays. Of course they aren’t buying “all copies”, and that wouldn’t even be possible in most cases since such books are usually flea market/attic material and most copies aren’t for sale (or even catalogued) to begin with. I’d be interested to learn who comes up with such lies though. Is it really just random people venting their frustration, or some kind of organized astroturfing operation? | | |
| ▲ | diseasedyak 9 hours ago | parent | next [-] | | It really does seem like an organized operation, given that it's so prevalent and they all seem to be in lockstep with their specious claims. | | |
| ▲ | wongarsu 9 hours ago | parent [-] | | I wouldn't be surprised to find out that this outrage is fueled by the same actors as the AI data center water outrage. Whoever they are |
| |
| ▲ | vavos 6 hours ago | parent | prev [-] | | I think these type of lies usually come about as a result of a game of broken telephone and things get exaggerated |
|
| |
| ▲ | sajithdilshan 10 hours ago | parent | prev [-] | | what is your source? |
|