| ▲ | djoldman 3 hours ago | |||||||
Just FYI, the bigger companies all allow you to block crawlers via robots.txt: | ||||||||
| ▲ | pixelesque 7 minutes ago | parent | next [-] | |||||||
Claude Bot still (at least last month, and it's been doing it for over a year now) seems to have a bug when traversing (at least my sites), wherein it drops the trailing slash of a directory (which is present in the a href tag), then makes the request to the subdirectory without the slash I put in the link, then Caddy automatically responds to that (via the built-in File handler) with a redirect telling it to add the trailing slash, and Claude Bot then makes the request again with the trailing slash that should have been there in the first place. So most subdirectory URLs get two requests from Claude bot, the first one needless because that wasn't the URL in the tag. | ||||||||
| ▲ | frereubu 3 hours ago | parent | prev | next [-] | |||||||
This is good to know, but a bit of all-or-nothing. It's a shame that, for example, Google doesn't support the crawl-delay field so you can tailor their crawling to your setup: https://developers.google.com/crawling/docs/robots-txt/robot... I presume it would also cut you off even more from referral traffic. | ||||||||
| ▲ | spiderfarmer 2 hours ago | parent | prev [-] | |||||||
90% of bot traffic on my network of websites is through headless Chrome, via residential bots nowadays. Impossible to block. Not even for Google, as they inflate my Adsense numbers as well. | ||||||||
| ||||||||