| ▲ | delichon 6 hours ago | ||||||||||||||||||||||||||||||||||||||||
> Why is git.kernel.org “interesting” to crawlers Interesting to crawlers is not a narrow scope. We have the same problem on a B2B car wash site. | |||||||||||||||||||||||||||||||||||||||||
| ▲ | wiredfool 5 hours ago | parent | next [-] | ||||||||||||||||||||||||||||||||||||||||
Seeing the exact same thing on (somewhat high profile) open data sites I run. The crawlers get stuck in a loop requesting the dataset listing page with every. single. combination. of. facets. At essentially as fast as it can be pumped out or blocked. Had one bot super interested in a single organization to the tune of 1000 r/min for 24+ hours. Realistic rotating user agents, realistic sec-* headers, no ip address seen more than a couple times in 10 minutes. The only commonality was the route. | |||||||||||||||||||||||||||||||||||||||||
| |||||||||||||||||||||||||||||||||||||||||
| ▲ | iririririr 5 hours ago | parent | prev [-] | ||||||||||||||||||||||||||||||||||||||||
so true. the article authors wishing crawlers will use git instead is so funny because the crawlers don't care at all. they are scrapping everything with brute force. they don't care about your content or effective alternatives, and one more site driving their real users crazy with Anubis is nothing more than a new blip in their dashboard. the crawler operators will not even look at the url. | |||||||||||||||||||||||||||||||||||||||||