| ▲ | est31 3 hours ago |
| > You can reprint a bestseller. You can't replace the last three copies of an 18th-century botanical text once someone shreds them for training data. And the judge said it's legal. So it's going to accelerate. Aren't they shredding only the books still under copyright protection? How is an 18th century botanical text still under copyright? IDK about the shredding, it's not nice, but it's more a problem with copyright law than AI companies. Scanning books you own should be legal from a copyright point of view, and not require shredding. Second, one should think about abandoned property provisions for copyright works published more than 50 years ago and in danger of being forgotten: once challenged, either you as the owner have to prove that the work is preserved for future generations (e.g. in various libraries around the world), or you have to authorize further copies, or you give up copyright on the work. |
|
| ▲ | ACCount37 3 hours ago | parent | next [-] |
| Scanning books by taking them apart into singular pages and scanning those pages is faster and cheaper. AI training is a numbers game, so they want faster and cheaper. What happens to the pages after? No one needs them anymore, so they get mulched and recycled. That would be the dominant scanning method even if copyright wasn't a thing. But then again - if copyright wasn't a thing, there would be much less need to scan any physical media. The reason why OpenAI can't just go on Amazon, buy a "digital edition" of a 2018 book and use that is that it would violate the license in ten ways, and then the DMCA laws that forbid breaking DRM on top of it. |
| |
| ▲ | voakbasda 2 hours ago | parent | next [-] | | Think about that last point for a moment. Our “rights to read” are diminished significantly with digital works as compared to printed works. Right of resale. Right to lend. In the end, digital publishing just isn’t right and will lead to massive gap in our historical records. They require active curation and cannot be preserved simply by resting on a dusty shelf. | | |
| ▲ | butlike 2 hours ago | parent [-] | | Every innovation since the microprocessor isn't worth saving in the grand scheme of things. When today's algae evolve enough into tomorrow's sentient creatures, they're really only going to need up to the industrial revolution and should probably stop right before that. | | |
| ▲ | compass_copium 29 minutes ago | parent [-] | | I'm personally a fan of more than 50% of children surviving past the age of 6, something that didn't happen until the 20th century. |
|
| |
| ▲ | TeMPOraL 12 minutes ago | parent | prev | next [-] | | Machines for non-destructively scanning books were developed and perfected long ago. The destructive scanning is neither technological limitation nor an issue of expedience. It's an issue of copyright law and fair use. | |
| ▲ | edoloughlin an hour ago | parent | prev | next [-] | | > What happens to the pages after? No one needs them anymore, so they get mulched and recycled. Strictly speaking, no one needs the Sistine Chapel or the Pietà etc. It would be a shame if they were mulched and recycled, though. | | |
| ▲ | doublerabbit 7 minutes ago | parent [-] | | Same with the magna carta and the American constitution. ChatGPT know them, i'd count that as digitalised why keep the originals? |
| |
| ▲ | dragonwriter 2 hours ago | parent | prev | next [-] | | > if copyright wasn't a thing, there would be much less need to scan any physical media. Because there’d be much less content created in any media to capture in the first place. | | |
| ▲ | eru 2 hours ago | parent [-] | | Empirically, probably not. We had lots and lots of content before copyright, and people seem to produce lots of content even in jurisdictions with weaker copyright. |
| |
| ▲ | classified 2 hours ago | parent | prev [-] | | It's the law, logic doesn't enter into it. |
|
|
| ▲ | graemep 2 hours ago | parent | prev | next [-] |
| > You can't replace the last three copies of an 18th-century botanical text once someone shreds them for training data. And the judge said it's legal. An 18th century book would be out of copyright so why would it be illegal to keep the original and scan it? |
| |
|
| ▲ | sethops1 3 hours ago | parent | prev | next [-] |
| > Aren't they shredding only the books still under copyright protection? How is an 18th century botanical text still under copyright? It's cheaper to scan the books if you do it destructively. Cost. That's why they're shredding irreplaceable texts. Nothing to do with copyright. https://www.404media.co/ai-companies-are-buying-tons-of-old-... |
| |
| ▲ | cestith an hour ago | parent [-] | | One doesn’t need to pulp the pages after scanning though. After scanning, they could be rebound and put into a library. |
|
|
| ▲ | mc32 3 hours ago | parent | prev | next [-] |
| Old rare books where there are single digit copies should enjoy some sort of patrimonial protection just like museum pieces. You can own them but have the state have the option to buy it if you’re about to significantly deface it or destroy it. |
| |
| ▲ | soco 3 hours ago | parent [-] | | The only issue I see is, how could you tell which are those books? | | |
| ▲ | 9dev 2 hours ago | parent [-] | | We manage to do this for endangered wildlife too without anyone counting every single specimen; why shouldn’t we be able to estimate how rare a book is? | | |
| ▲ | phoghed 2 hours ago | parent | next [-] | | And if you instituted this, the commenters of this very website would surely decry it as a prime example of government overstep and waste. | | |
| ▲ | shimman an hour ago | parent [-] | | Commentators on this web site work for some of the most evil organizations on the planet and have beliefs that 95% of the population rejects. You can safely ignore the YC cohort of devs and be fine. |
| |
| ▲ | eru 2 hours ago | parent | prev [-] | | It's pretty expensive for the wildlife. Most rare books are rare because no one cared enough about them. Ie most rare books are rubbish. |
|
|
|
|
| ▲ | croes 3 hours ago | parent | prev [-] |
| Books that are shredded can’t be scanned by competitors. |
| |
| ▲ | jfyi 23 minutes ago | parent [-] | | Yeah, this is the point. I don't understand the bulk of this conversation. Copyright doesn't matter, the books themselves don't matter. All that matters is that their corpus of training data grows faster than their competitors. |
|