Remix.run Logo
RobotToaster 4 hours ago

Do they offer bulk torrent downloads as an alternative?

QuantumNomad_ 3 hours ago | parent | next [-]

Once upon a time some people explored backing up the Internet Archive.

However, that experiment ended. They mention there were some learnings and they then say:

> The Internet Archive continues to explore methods and code to decentralize the collection, to have a mirror running in various ways - these include IPFS, FileCoin, and others. The INTERNETARCHIVE.BAK project also added general mirroring and tracking code to a number of projects that are still in use.

https://wiki.archiveteam.org/index.php/INTERNETARCHIVE.BAK

I would really like to know if any sort of thing like that is still ongoing and if it’s accessible to people in general. Would be nice to mirror some data from IA to my local drives, for example via BitTorrent or IPFS, to have it for offline exploration and personal archive.

I know that individual items have torrents. And I’ve downloaded a few that way but always it ends up only using the “web seed” (i.e. the BitTorrent client is retrieving the files from IA via HTTP) because there are no one seeding some random single item I found. Plus, those torrents are unreliable sometimes because they include meta data files that were since updated but the torrent was not updated and so the web seed is giving the updated files that don’t match what the torrent says their hashes should be. So then you have to jump through some extra hoops to fix that and then resume the download, and all the while the HTTP connections to IA servers time out because their servers are overloaded. So when I say I wonder about possibilities of using BitTorrent I mean to retrieve whole collections of many items instead of individual ones, and with actual other peers instead of just having it put load on IA HTTP servers.

alightsoul 41 minutes ago | parent | next [-]

The internet archive's decentralization project is paused as far as I can tell. They have too many things to do and too little funding to do it all. Their current strategy seems to be establishing new legal entities outside the us like in Canada and Switzerland, but they don't accept web traffic even though they hold full copies of the internet archive. There used to be a full copy in Egypt at the library of Alexandria and another in the Netherlands. Not sure if they're still in use, but they did accept web traffic. They hold a decentralized web camp every year in the middle of a forest

giantrobot an hour ago | parent | prev [-]

The Internet Archive's torrents are a sick joke. I've yet to find one that actually manages to complete. They always get stuck at 90-something percent but that final blocks always fail verification and get retried, fail, and the process repeats forever. Because they're web seeds they're hitting IA infrastructure and not offloading to a real swarm. So their broken torrents are just screwing themselves.

echelon 4 hours ago | parent | prev [-]

I would love to be able to download every page of a given domain as an archive, and I'd pay to do this.

msephton 4 hours ago | parent | next [-]

They provide a free cli tool to do this.

petcat 4 hours ago | parent | prev | next [-]

isn't that what wget -m does? what is there to pay for?

carlosjobim 3 hours ago | parent | prev [-]

You'd pay the domain owner for it? How much?