Remix.run Logo
wongarsu 5 hours ago

Isn't the issue here the three order of magnitude difference between the (presumably optimized) JS implementation and the optimized C kernel on your browser? If the two stay within an order of magnitude of each other Anubis is a perfectly viable and scalable solution. Bots want to open many orders of magnitude more pages than a normal user, so the resource investment for each single page matters a lot more to them

For reference, the challenge on lists.ffmpeg.org takes 8 seconds on Firefox on my three year old laptop CPU that has worse benchmark scores than the iPhone 17 (tbf, the laptop also cost less than an iPhone 17). 8 seconds doesn't run against thermal limitations, so I really don't see why Safari on a modern iPhone should be so slow at this

semiquaver 5 hours ago | parent | next [-]

  >  I really don't see why Safari on a modern iPhone should be so slow at this
me neither, but I don't think it changes the argument. There's always going to be someone on a low-end device. Your adversaries already have superhuman coding ability and infinite patience. Why would you expect the long-term advantage to be with the defenders?
paytonjjones 4 hours ago | parent [-]

In this case, because there's a vastly more efficient economic path for the adversaries (cloning).

They're not trying to engage in an arms race, they're trying to channel a racing river into its natural course.

pbronez 4 hours ago | parent [-]

It’s more economical at a compute level, but not at the developer level. The moment you start customizing your crawler to use protocol X for site Y your scale story collapses.

paytonjjones 3 hours ago | parent [-]

It's a good point, but in practice it depends on how easy those customizations are to implement / maintain, and how much money and effort you save. At some point the compute cost can disrupt even the nicest scale story.

I think the path forward is that websites offer one path for humans, and another for scrapers. But the huge catch is the path for scrapers must be _genuinely_ and _reliably_ the more economical and scalable path (either through something like PoW arms races, or through fear of litigation). Otherwise they will continue to ignore instructions and intrude on the human path.

inigyou 2 hours ago | parent [-]

Why aren't we litigating against scrapers, anyway? DDoS is a felony.

ekidd 2 hours ago | parent [-]

Largely because they're residential botnets in places like Brazil (a real example from one of my sites that was crawled to near-destruction). Someone could probably do something about this, but it's out of reach for individual site owners.

inigyou 2 hours ago | parent [-]

If you block Brazil, they'll find an alternative, maybe then you can sue them.

embedding-shape 29 minutes ago | parent | prev | next [-]

> which takes ~180sec for my iPhone 17 to solve at ~100KH/s

> so I really don't see why Safari on a modern iPhone should be so slow at this

FWIW, my iPhone 12 Mini also does ~110KH/s with Anubis on lists.ffmpeg.org, so seems fairly likely that Safari somehow here isn't working as expected.

Aurornis 4 hours ago | parent | prev [-]

> Bots want to open many orders of magnitude more pages than a normal user, so the resource investment for each single page matters a lot more to them

Depending on the configuration, Anubis will supply a token after the challenge that bypasses the challenge for a time.

So any scraper that retains basic cookies will be able to bypass the challenge for a number of page views.

A user who needs to load a single page and a bot that wants to scrape a number of pages may pay the same cost.

The amortized per-view cost is highest for the real user.

rplnt 2 hours ago | parent [-]

Now you have a session of sorts and can limit the requests for that client, right? They can be fast, just limited in volume - regular user isn't punished.

JsonCameron an hour ago | parent [-]

yep, that's the exact play. Or better fingerprinted & blocked in other ways