| ▲ | warkdarrior 11 hours ago |
| > Our ideal is to scan and upload all the world’s publications before publishers completely block knowledge, and before AI companies scan and destroy all the world’s books and papers. Isn't this just doing the work for the AI companies??? Then the AI companies can simply download a copy of Anna's Archive. |
|
| ▲ | throwatdem12311 11 hours ago | parent | next [-] |
| At least the knowledge will be available to everyone instead of mashed together and regurgitated poorly through proprietary LLMs. |
|
| ▲ | torh 11 hours ago | parent | prev | next [-] |
| At least there will be a copy left for us. The AI companies won't share these books in their original form. |
| |
| ▲ | brainwad 11 hours ago | parent [-] | | Because it's illegal. That's the whole reason they are shredding books in the first place, because copyright law forces them to do stupid things. Google wanted to share the whole of Google Books 15 years ago, too, but they were sued to hell, so now you get a watered down search functionality. | | |
| ▲ | Paratoner 11 hours ago | parent | next [-] | | The poow AI execs being forced to commit acts of intewwectual tewwowist when all they wanted was to cynicawwy make the wowld a wowse place | |
| ▲ | JohnFen 11 hours ago | parent | prev | next [-] | | > because copyright law forces them to do stupid things. Baloney. They aren't forced to destroy books. They could leave well enough alone and not scan them to begin with. They're voluntarily choosing to do this. | | |
| ▲ | brainwad 9 hours ago | parent [-] | | Let us not forget they own the books. They bought them legally on the open market. Also, no modern book truly is destroyed forever, there are always libraries of record with a copy. | | |
| ▲ | JohnFen 8 hours ago | parent [-] | | > Let us not forget they own the books. That's rather beside the point. I don't think anyone is arguing that what they're doing is in some way illegal. > no modern book I'm less concerned about modern books. | | |
| ▲ | brainwad 7 hours ago | parent [-] | | All the books being destroyed are modern enough to have deposited copies, or they wouldn't be still under copyright. The destruction is happening to satisfy judges that no illegal copying is happening, only a transformation of medium. |
|
|
| |
| ▲ | xandrius 11 hours ago | parent | prev | next [-] | | Google probably wanted to sell the whole of Google Books. | |
| ▲ | subscribed 11 hours ago | parent | prev | next [-] | | Copyright law didn't force them to torrent terabytes of books what they did and got caught doing so. I'm not convinced they destroy the books to obey the law, lol. | |
| ▲ | naasking 11 hours ago | parent | prev | next [-] | | Yes, government regulations are almost always behind commercial entities making seemingly irrational choices. | |
| ▲ | embedding-shape 11 hours ago | parent | prev [-] | | > because copyright law forces them to do stupid things This is such dangerous train of thought, to give them the benefit of being forced to destroy books. Why is that exactly, and who is forcing them? You can also, you know, find another way? Like the data centers who currently use very dirty energy acquisition methods (not all of them), are they also "forced" to do this, because they too need to make as much money as the other ones? How long would you continue this idea of others "forcing" for-profit companies to try to make more money, regardless of consequences? Destroying books used to be an obvious dumb, stupid and shit idea, not sure how somehow a for-profit company making of a digital copy for themselves of the book before destroying it, suddenly makes it not a shit idea for the rest of humanity. | | |
| ▲ | brainwad 11 hours ago | parent [-] | | It is a shit idea. But that's copyright law for you - if you want to digitise the work for yourself, you according to latest precedents have to destroy the copy you digitised ¯ \ _ ( ツ ) _ / ¯ I don't blame the companies for either wanting digitised works, nor following the law. I blame the absurd court ruling, and I blame the publishers for not having digitised the old works themselves, in which case they could just sell e-books to the labs... They are after all the only ones who can legally do it non-destructively. | | |
| ▲ | embedding-shape 10 hours ago | parent [-] | | But why do they have to do this at all? If it's a shit idea, and you cannot do something without negative side-effects of it, can't you just not do it? Why these companies absolutely have to do this? | | |
| ▲ | brainwad 9 hours ago | parent | next [-] | | Well, they want to use the book they own in digital format. They own it, they get to decide what to do with it. | |
| ▲ | famouswaffles 10 hours ago | parent | prev [-] | | What's shit about it ? They're buying books that would be headed for the pump or trash heap anyway. Millions of books are trashed or pulped every day. | | |
|
|
|
|
|
|
| ▲ | TaLiTr 6 hours ago | parent | prev | next [-] |
| > Then the AI companies can simply download a copy of Anna's Archive. And so can everyone else.
That's the ideal outcome. Only Anti-AI types are against this, they probably don't even care about the books, it's just a proxy for trying to stop the "evil AI companies." |
|
| ▲ | 11 hours ago | parent | prev | next [-] |
| [deleted] |
|
| ▲ | voidhorse 11 hours ago | parent | prev [-] |
| Yeah, which is completely fine. There's a major difference between: A. Company destroys a book forever for training. Its scan is locked away forever in company records. In this case: - This information is locked away in the improvisations of an LLM. It is no longer possible to directly access the information as written by the human being that authored it. This constitutes the loss of literary history, or at least loss of access to that history to the general public. - The price of the book is no longer distinct from the general price of "inference". It becomes increasingly impossible to pay for specific information, instead you are charged by the meter for general machine inference, which doesn't even give you access to a specific text. - The provenance of information is totally destroyed. This causes potentially unresolvable problems of authority and citation. If the original source is lost, how are we to know if a random LLM claim about an obscure topic or specific niche text is even true or just hallucinated? B: Company uses freely available scanned copy of the text: None of the issues above obtain, since anyone can still access the actual book. Most importantly, this reduces the power companies have to force everyone to continually pay for a derivative form of the book's information in perpetuity in the form of token costs. I much prefer B. |
| |
| ▲ | Filligree 11 hours ago | parent [-] | | B is illegal, and Anthropic ate a billion dollar fine for trying it, so you can’t even claim they don’t want to. | | |
| ▲ | voidhorse 11 hours ago | parent [-] | | Then give up training on antiquated books. Why does an LLM aimed at providing utility for people living in 2026 need to be trained on rare (thus probably obscure) texts of yore in the first place? Because these companies have no real strategy beyond trying to capture any and all information they possibly can to try and lock it away and charge the public for it in perpetuity. |
|
|