| ▲ | doginasuit 5 hours ago | |
Let's not conflate crawlers with the traffic that bot protection services block. A crawler that respects robots.txt is a good internet citizen and can provide a vital service. | ||
| ▲ | senko an hour ago | parent | next [-] | |
A crawler that respects robots.txt is useless in practice since many sites only allow Googlebot and maaaybe Bing - by name. | ||
| ▲ | madibo3156 4 hours ago | parent | prev | next [-] | |
However, so-called AI crawlers are not the same as crawlers of yore. They hit live pages every time a user prompt triggers a web search. This Web Search API, unlike an AI crawler, only fetches periodically. It feels like a step in the right direction for managing resource strain across the internet. If only the LLM giants could do something similar. | ||
| ▲ | sreekanth850 4 hours ago | parent | prev [-] | |
And you think all this web AI crawlers will respect robots.txt. That era is gone. | ||