Remix.run Logo
bsimpson 5 hours ago

It's an open secret that you can often circumvent paywalls by searching Wayback.

zymhan 5 minutes ago | parent | next [-]

Only some of them, it is not universal.

gambiting 5 hours ago | parent | prev [-]

Every single paid article linked on HN has the way back machine link as the very first comment.

ValentineC 5 hours ago | parent [-]

The links are usually to Archive.today (aka archive.ph and a bunch of other domains), not Wayback Machine (which is run by Internet Archive).

eek2121 2 hours ago | parent | next [-]

Correct:Also, archive.* has actively edited archived sites to promote their agenda. Why folks continue to use them confuses me. One would think the big wikipedia purge would curb such behavior.

normie3000 25 minutes ago | parent | next [-]

> Why folks continue to use them confuses me.

I use them. I haven't ever heard mention that the content is edited. Do you have a source?

DaSHacka an hour ago | parent | prev [-]

Ironically, your framing of the situation is infinitely more disingenuous to push a personal agenda versus anything the archive.today guy did.

petcat 4 hours ago | parent | prev [-]

ehh it's a distinction without a difference. The point is that alternative links are available to circumvent paid access for anyone that wants them.

sandcat_ 4 hours ago | parent | next [-]

That isn’t the point being discussed. The point being discussed is that it’s bad form to abuse a service (archive.org) that is provided for free, for the public good in order to run commercial scraping operations.

petcat 3 hours ago | parent [-]

It's bad form to scrape the scrapers?

sandcat_ 3 hours ago | parent [-]

Yes, arguably, and for reasons I already gave. I’d genuinely spend a bit more time reading and thinking rather than replying. Your replies are pithy but you’re missing details and frankly making cognitive mistakes. (Apologies if this seems harsh, I don’t mean it as an insult, but this thread has blown up entirely unnecessarily- and yes, I know I’m not helping either!)

petcat 3 hours ago | parent [-]

You seem to think that scraping websites "for the public good" is somehow different than scraping websites for any other reason.

The end result is exactly the same.

sippingabonedry 2 minutes ago | parent | next [-]

You missed the XCancel flamewar yesterday. The consensus is we should be allowed to scrape data and bypass login walls if we dislike the site owner, or feel we are owed free access by arbitrary criteria, it's sort of an unwritten rule here.

DaSHacka an hour ago | parent | prev [-]

The minuscule traffic generated by the wayback machine, which serves to preserve the content for years to come, is completely incomparable to the scrapers that hammer every single href linked on a website.

organsnyder 4 hours ago | parent | prev [-]

They're different sites, with different goals, run by different people.

petcat 4 hours ago | parent [-]

That provide the same functional service....

Hence, distinction without a difference.

celsoazevedo 3 hours ago | parent | next [-]

They are 2 different services, run by different people, one goes out of their way to bypass paywalls while the other doesn't, one is banned by Wikipedia and the other isn't, etc.

I think it's a distinction worth making.

Not to mention that the Wayback Machine itself isn't exactly a good tool to bypass paywalls as most paid sites don't let them archive paywalled content anyway.

fluffybucktsnek 4 hours ago | parent | prev | next [-]

Given that the root of the discussion is about Internet Archive being hit with huge traffic and not the functionalities provided by Wayback Machine, it very much is a distinction with a difference.

petcat 3 hours ago | parent [-]

Bot traffic or human traffic doesn't matter. The goal is to read websites without having your own access.

So Internet Archive, Archive.today, Archive.ph, etc. are all just means to the same end.

HDBaseT a few seconds ago | parent [-]

I am not sure if you have the wrong impression of the Internet Archive.

The internet archive is not designed to circumvent anything. It is not designed to "grant access without having your own access".

rpdillon 3 hours ago | parent | prev [-]

Yeah, you're mistaken. One archives web pages, the other maintains a list of paid-access accounts and fetches information from behind paywalls as a service.

DaSHacka an hour ago | parent [-]

Exactly this

archive.org is the more straight-laced archive that doesn't circumvent sites that try to block it, and removes content they deem 'problematic' even if not illegal or requested by the site owner.

Meanwhile archive.today/ph/is/etc is the guerrilla alternative run by a die-hard datahoarder that seeks to archive the information itself, bypassing whatever blockers/login pages/whathaveyou to achieve the result.

It's nice to have both options. When I archive a site, I usually use both for added resiliency.