Remix.run Logo
awakeasleep 2 hours ago

Anyone who is into books eventually finds out that we're permanently losing them all the time. Like they're thrown out, and lost forever. University libraries throwing out huge collections of out-of-print material to make room for new books, study spaces, even cafés. Municipal libraries turning over their collections. Books that never made it to libraries going out of print, tossed in the trash after yard sales.

The AI companies digesting this stuff is a net win for humanity. And I'm not a fanboy! Ideally they'd upload them to Anna's archive too, but even if they keep it private forever, at least these books live on in some way in the model weights. Thats better than a landfill.

ygjb an hour ago | parent | next [-]

It would be really great if the companies who are doing this would commit to placing the scanned files into a public trust that would coordinate with organizations like the Gutenberg project to ensure that the scanned materials enter the public domain on schedule. Publishing encrypted archives with the keys in escrow would be a good first step.

IMO that would go a long way to resolve any concerns about losing books. I still don't like the idea of extremely hard to find or last prints being actually destroyed for this, but it certainly makes it more palatable.

Aurornis an hour ago | parent | next [-]

They cannot do those things. The reason they’re scanning the books is because it’s not possible to legally obtain or transfer their digital copies. They have to do their own scans and keep them in house.

ygjb an hour ago | parent | next [-]

First, to be clear, I work for Amazon, and I don't work on this or related efforts, I am a security engineer. I also won't talk about or answer questions at work, and my comments are more generally about the practice (all of the major organizations building AI are doing destructive book scanning). These are my opinions, and do not reflect my employers (past or present).

You might be right that they can't do that now. The simple path forward is to have these companies simply make a public commitment to publish the data when the copyright expires. I also think there is a space to be carved out, probably through regulation, to ensure that there is a clear path for these scans to enter the public domain, at the very least.

This is not just important for these specific books, I have written on other platforms and in other spaces about the importance of media companies and those who benefit from strong copyright laws to protect and generate profits and revenues to repay the public for the cost of that enforcement over time by ensuring that at the appropriate time, those works fully enter the public domain. That could mean a restructuring of the Library of Congress in the United States to become a modern Library of Alexandria to host data, and shifting to a registered copyright model where to gain the protections of the court, you need to upload/submit your copyrighted works for storage and eventual release. I doubt it would ever happen there because of the amount of money invested in tying up IP in the United States, but perhaps a more amenable location like the EU could help with that.

However it might work, part of the promise of the Internet was that information would be liberated, but we see every day how much information gets sent down the memory hole when businesses, sites or services shut down, or how regularly companies abuse IP related regulations to attempt to strangle competition. It would be expensive now, but it would create an incredibly valuable legacy of information for the future, and it will only get more expensive to build such a thing as time goes on.

axus an hour ago | parent | prev [-]

Copyright does expire, for older books.

It would make more sense for government to accept digital copies for any book, and share the out-of-copyright ones. In the United States, we have a Library of Congress who could do it if a law were passed with funding.

cucumber3732842 an hour ago | parent | prev [-]

I wonder if there's any way this can be construed to get favorable tax treatment. That'd actually get them doing it.

Melatonic 42 minutes ago | parent [-]

The real solution right here.

bitmasher9 an hour ago | parent | prev | next [-]

I would feel much better about this process if they were uploaded and if it were framed as a knowledge preservation project. This would only slightly increase the cost of the project, but have a huge impact on its perception and its net positive impact.

Of course, actually benefiting humanity is only a minor, indirect concern for investors.

threetonesun an hour ago | parent | next [-]

I thought this was the original purpose of Google Books. It's actually a little surprising to me that there's anything left out there worth buying in physical form and scanning that hasn't already been digitized in some way.

sagarm an hour ago | parent | prev [-]

Google faced a decade of litigation for making books searchable. It's all downside and little upside for a business.

rtkwe 30 minutes ago | parent [-]

They could scan and store the data until the copyright expires without issue probably. It's a long time line to work with but you could probably get away with it that way.

Aurornis an hour ago | parent | prev | next [-]

> Ideally they'd upload them to Anna's archive too

Not only can they not do that, they must scan physical copies because they are forbidden from using digital pirated copies from sources like this.

Anthropic had a big settlement because they were caught using downloaded digital copies. As a response they’ve ramped up their book scanning and others have followed.

rcxdude an hour ago | parent [-]

Yeah, the very disjointed way that copyright is applied (I would argue, rights that should also transfer across to digital media but generally have not) is contributing to this situation in the first place.

alightsoul an hour ago | parent | prev | next [-]

If or when they go bankrupt or reach agi, they will just delete them. I hope Anna's archive already has them anyways. Apparently openai's newest unreleased model, gpt 6, is capable of continuous training at Inference time, like a person is. That might be enough to delete them

infecto an hour ago | parent [-]

And that’s why imo the outrage should be about copyright law not the businesses. I think it’s a hard problem to solve but right now with copyright on new books being the life of the author + 70 years it’s a bit silly.

alightsoul an hour ago | parent [-]

The only way to gatekeep a book behind payment should be, if it is in print.

infecto an hour ago | parent [-]

I like this and think it could make a lot of sense but you would need to thoughtfully tie it to volume or something similar. I could see publishers gaming the system. I would also add that once it drops out of print that it should belike generic drugs. Anyone can use it. IMO making it quicker free use stops most of what this article is describing.

red_green_yell an hour ago | parent | prev | next [-]

This is the right answer but the reason they can’t upload the scans is copyright law as demonstrated by Google having to settle with the publishers and allow them to remove their books and limit free access to 20% of text. The AI companies are essentially compressing the information in a huge swath of books that would otherwise be headed to landfill and making them 1000x more accessible. This is unquestionably one of those instances where capitalism is taking money from rich investors and benefiting the 99%.

Finnucane an hour ago | parent | prev [-]

That is the problem: they are digesting it. They are not creating a new kind of library, where you could say, show me the text of "How to Fix Your Ice Cream Problems". (An actual book I own) It is not being done for our future reference.

red_green_yell an hour ago | parent [-]

They are digesting it an incorporating into the weights. You won’t be able to get the exact page but you will (if the model is good) be able to get the knowledge out of it in a likely far more concise, relevant, and certainly more widely accessible than the book sitting on your shelf.