Remix.run Logo
▲ slow_typist 3 hours ago

Of course they know. I have seen local newspapers who managed to block archive.ph's mechanism. If a bigger outlet does not use technical measures against services like wallhop, it is safe to assume that’s deliberate.

It makes sense economically since the part of the population using such services is so small and may have specific demographics that justify specific treatment just from a marketing POV.

Then you want search engines to get the whole content (okay at least before they started to summarise everything but the kitchen sink) - that will always leave a loophole.

▲dns_snek 2 hours ago | parent [-]

I think you have that the wrong way around. Archive services have to explicitly bypass paywalls so most local newspapers are simply unsupported.

▲walrus01 an hour ago | parent [-]

Based on some of my recent experience with building a scraper for purposes that aren't a news website, what has changed in the last 12 months or so is the availability of very low cost (or free, if you can run it on a 256 or 512GB system in your own office) LLM that are good enough to give it a target of something and have it build a custom scraper/paywall bypass profile on a per site basis. I'm referring specifically to things that can be run in harnesses and score well on terminalbench 4.0 and SWE LLM benchmarks.

When previously nobody would have gone through the effort to build and maintain a custom scraper against a moving target, for some small to medium size city's newspaper, now it's just one of a myriad of 'scraper profiles' you can have a near fully automated tool create.