| ▲ | xena 5 hours ago | |
AI bros think they should be exempt from robots.txt. Administrators of big services beg to differ. No solid consensus has arisen. I bet it's gonna take a lawsuit or two to see how it shakes out. | ||
| ▲ | Dylan16807 2 hours ago | parent | next [-] | |
wget ignores robots.txt outside of recursive mode. I think it's correct to do so, and I think an AI loading a handful of pages in response to a command should be about the same. | ||
| ▲ | ghaff 4 hours ago | parent | prev | next [-] | |
From the start, robots.txt has always been an indicator of a site's preference with no actual legal significance. | ||
| ▲ | recursive 4 hours ago | parent | prev [-] | |
If a new directive was introduced that allows for an explicit setting in robots.txt, do you think the bros would follow it anyway? Something like `ALLOW AGENTS` or `DISALLOW AGENTS` | ||