| ▲ | asaddhamani 18 hours ago |
| But that scan is never made available to us in its original form. So it getting scanned by the AI company does nothing to preserve the book. |
|
| ▲ | scarmig 18 hours ago | parent | next [-] |
| Dumpsters also don't typically come equipped with a robot scanner and network uplink built in. Like, I really don't know what people objecting to this imagine typically happens to old, unwanted books. They don't get sent to some magical library in the countryside if unpurchased where they are carefully maintained forever (next to where Rover spends the rest of his days). They are very literally thrown into the trash. That said, I'd be thrilled if the US government required AI companies to make them available to the public. I'd even settle for the US government making it legal for them to. |
| |
| ▲ | fmajid 16 hours ago | parent | next [-] | | The Internet Archive tries to be that magical library, but they can only scan and physically archive what is sent to them. | |
| ▲ | ralferoo 14 hours ago | parent | prev | next [-] | | > I really don't know what people objecting to this imagine typically happens to old, unwanted books. In the UK at least, people usually take them to a second hand / charity shop, who sort through them and send the valuable ones to auction (typically early editions, 100+ years old) and then either sell them themselves (for recent books that are easy to get rid of) or sell them to specialised second-hand bookshops. Most of the specialised second-hand bookshops rarely throw books away, usually if nobody buys them after a couple of years they end up in the extreme discount piles (20p, 50p etc) and probably only trashed if they still don't sell from there. | | |
| ▲ | theshrike79 14 hours ago | parent [-] | | So trashed, but with a bunch of extra steps then? | | |
| ▲ | tentacleuno 3 hours ago | parent [-] | | I would presume that the extra steps incrementally diminish the possibility of the book remaining unsold, and thus destroyed or sent elsewhere. | | |
| ▲ | theshrike79 2 hours ago | parent [-] | | And it also adds costs in every step. Someone needs to move thousands of unwanted books from high end stores to lower and lower end stores. Someone needs to store them in the proper environment etc. I do get the _idea_ of preserving books, but... people don't care. I just threw out well over a thousand books from my grandparents house this spring. There were ~6-10 "valuable" books there. Two because I personally knew someone who wanted old war-time books and a few 100+ year old bibles. And maybe two dozen books worth saving, mostly because they were from big-name authors or had stuff that nobody would print anymore (I have detailed instructions how to make laughing gas and how to build an underground chemical lab - hobby books in the 50s were ... interesting :D ) I literally couldn't give away the rest. And I tried. It was all just "interesting, but..." - no way to justify using the shelf space for books that, realistically, nobody will actually ever read again. |
|
|
| |
| ▲ | soperj 18 hours ago | parent | prev [-] | | They're buying the books from resellers, not rescuing these books from dumpsters. Stop being an apologist. | | |
| ▲ | skeledrew 17 hours ago | parent | next [-] | | Dumpster is where they go when the resellers fail to complete sales. | |
| ▲ | scarmig 18 hours ago | parent | prev [-] | | The magical library in the countryside, to be painfully explicit, does not exist. | | |
| ▲ | Ekaros 17 hours ago | parent | next [-] | | And they should not even be needed. In many places the issue is solved at start. Copy or copies of each commercially produced book is send to national library. Which with tax payer money keeps an archive. Meaning that at least one copy exist for research purposes if needed. | | |
| ▲ | scarmig 17 hours ago | parent [-] | | Unfortunately, that's not the case in the United States. The LOC only selects around half of published books to be permanently held. The rest are disposed of (usually returning them to the publisher, donating them to a library, or destroying them). | | |
| |
| ▲ | soperj 17 hours ago | parent | prev | next [-] | | How many times can you post the same thing in a thread? | |
| ▲ | exe34 16 hours ago | parent | prev [-] | | The internet archive |
|
|
|
|
| ▲ | dukeyukey 17 hours ago | parent | prev | next [-] |
| If it were legal they may well do that as a public branding exercise. Google already tried and got punished for it! |
|
| ▲ | skeledrew 17 hours ago | parent | prev | next [-] |
| It was never available to you/us in the original form either. |
|
| ▲ | red75prime 17 hours ago | parent | prev [-] |
| ...because it is illegal to copy copyrighted material. 70 years later they might do it. |
| |
| ▲ | rhdunn 17 hours ago | parent | next [-] | | 95 years after publication. Many other countries also have an X years after the author's death clause where X varies between countries but is at least 70. There are also other weird issues such as the UK having a clause protecting Peter Pan (so a children's hospital gets royalties) and the King James translation of the bible (under Crown copyright) that extend the copyright even further. In short, it's a mess. | |
| ▲ | mrweasel 16 hours ago | parent | prev [-] | | The thing I find most hypocritical though is that they are probably never share their libraries with anyone. After scanning, downloading, stealing, overloading websites and everything in between, to acquire enough data for their stupid machine, they're not going to share their data? I get that most of it can't be shared, but a lot can. There's no reason why you need to destroy multiple copies of a book from 1880, when it's free to share. At the same time I can understand keeping track of when each books enters public domain might also be an absolute nightmare, and I wouldn't blame the AI companies for not wanting to deal with that. For the stuff they absolutely know is clear, they should provide dumps for everyone to download. | | |
| ▲ | novok 16 hours ago | parent | next [-] | | This is solved by law, which is solved by 'we the people' and I bet many AI companies would be fine with something like the equivalent to patent law with bankruptcy escrow to the library of congress, where they must release the scans in 10 years for books that the vast majority will not give a flying shit about. By then the advantage is long gone in data moat. | |
| ▲ | fmajid 16 hours ago | parent | prev [-] | | Since they seem to leapfrog each others’ models every few months, the training data is one of the few ways they can build competitive advantage, and that explains why they don’t share, even if we don’t have to like this. | | |
|
|