| ▲ | nneonneo 7 hours ago | ||||||||||||||||
I disagree. The kernel finds it effective - 66% of scrapers are turned away directly. The scraper problem now is fleets of residential proxy devices - often things like smart TVs, phones, and browsers with some “proxy SDK” installed as part of an app’s monetization scheme. They make a couple of requests to a site - just enough to fly under the radar - and move on to a different site. If each new site they hit forces them to solve a proof-of-work, that’s a meaningful dent in their scraping performance. Many of these boxes may not even have the spare CPU power to efficiently solve so many proofs of work - and anything that makes an owner notice their device is running slow is something that could meaningfully impede adoption of these SDKs, or force the operators to choose between minimizing performance impact or scraping more sites. | |||||||||||||||||
| ▲ | tptacek 7 hours ago | parent | next [-] | ||||||||||||||||
It's weird to believe data center based, Internet-scale scraping operations will be less able to allocate compute to proof-of-work challenges than individual users. This is design problem with things like Anubis: proof-of-work depends on a cost asymmetry between attacker and defender. But in scraping, both legitimate users and scrapers get the same value out of a transaction. | |||||||||||||||||
| |||||||||||||||||
| ▲ | Y_Y 7 hours ago | parent | prev | next [-] | ||||||||||||||||
The implication here is that the proxy fridge forwards the Anubis challenge to a dedicated rig controlled by the scraper who efficiently solves it and returns the answer. | |||||||||||||||||
| |||||||||||||||||
| ▲ | oasisbob 5 hours ago | parent | prev | next [-] | ||||||||||||||||
> Many of these boxes may not even have the spare CPU power ... I don't think that's generally how these networks use residential exit proxies. There are at least a dozen well-developed frameworks out there for decoupling the crawler from the network exit point. Most res proxy exits are just slinging bytes for clients using SOCKS, or another tunneling protocol. If nothing else, a modern scraper will want better control over their TLS fingerprints, and you can't get that if you're depending on the on-device TLS libraries alone. | |||||||||||||||||
| ▲ | lxgr 6 hours ago | parent | prev | next [-] | ||||||||||||||||
Why would they even run a browser engine on the devices they're hosted on? All they need to do is forward traffic and launder its IP origin. They don't even need to be able to (and would actually be well advised not to) decrypt TLS streams. | |||||||||||||||||
| ▲ | lxgr 6 hours ago | parent | prev | next [-] | ||||||||||||||||
> 66% of scrapers are turned away directly. Until they discover this neat trick [1] and solve challenges orders of magnitudes more efficiently than legitimate users. The game theory of Anubis is not sound. It makes fundamentally less sense than Captchas, and even those have been on the way out for a while. | |||||||||||||||||
| ▲ | semiquaver 5 hours ago | parent | prev | next [-] | ||||||||||||||||
Until you actually do the math and realize that it is not meaningful at all. It’s equivalent to the blogs that have a custom “bot protector” that asks you “what’s 2+2” every time you submit a comment. It might work temporarily as an inconvenience, but nothing more. | |||||||||||||||||
| ▲ | graemep 6 hours ago | parent | prev [-] | ||||||||||||||||
How s that measured? How do you count human users who have been turned away? | |||||||||||||||||