Remix.run Logo
bradly 4 hours ago

Do you think there should be a way for a site to tell an agent it isn't allowed access? I'm not sure where I land on this exactly tbh, so no judgment cast.

aaron_m04 4 hours ago | parent | next [-]

robots.txt?

bradly 4 hours ago | parent [-]

Has it been settled whether robots.txt applies to user-driven chat sessions and if things like the crawl delay should be applied to say an end-user, an ip address, a harness provider, etc? My understanding is robots.txt is more for training exclusions, but less so for agent work.

bityard an hour ago | parent | next [-]

robots.txt applies (or should, in my opinion) to anything that automatically follows a link. Basically any software that is not a human-controlled web browser or single-shot curl command. Everything else: robot.

xena 3 hours ago | parent | prev [-]

AI bros think they should be exempt from robots.txt. Administrators of big services beg to differ. No solid consensus has arisen. I bet it's gonna take a lawsuit or two to see how it shakes out.

Dylan16807 32 minutes ago | parent | next [-]

wget ignores robots.txt outside of recursive mode. I think it's correct to do so, and I think an AI loading a handful of pages in response to a command should be about the same.

ghaff 2 hours ago | parent | prev | next [-]

From the start, robots.txt has always been an indicator of a site's preference with no actual legal significance.

recursive 2 hours ago | parent | prev [-]

If a new directive was introduced that allows for an explicit setting in robots.txt, do you think the bros would follow it anyway? Something like `ALLOW AGENTS` or `DISALLOW AGENTS`

Analemma_ 3 hours ago | parent | prev [-]

I want agents to be able to act on my behalf, that’s the entire point. An agent should be able to do anything I can do sitting at my browser.

compiler-guy 3 hours ago | parent | next [-]

I suspect most people would be ok with this if they could only do it at the rate and frequency you yourself can do it. The problem is largely one of scale.

daveoc64 2 hours ago | parent | next [-]

Is scale what we're discussing though?

e.g. a prompt of "fetch <article URL> and summarise it for me" is very close to what a human would be doing with a web browser, and doesn't seem to involve any kind of scaling issue.

compiler-guy 14 minutes ago | parent [-]

It’s easy to write instructions that have the agent check once every fifteen minutes, or even once an hour, in perpetuity, which never sleeps. And people do write such instructions. A human can’t do that by hand for very long.

The problem is that it is hard to distinguish your one off (which seems perfectly fine) from the tidal wave of bad actors.

cruffle_duffle 3 hours ago | parent | prev [-]

Then make agent friendly content. Take the text and make a markdown version.

fineIllregister 2 hours ago | parent [-]

People doing this say it makes things worse because then the bots download both.

compiler-guy 2 hours ago | parent [-]

Not to mention that it solves none of the rate issues. If the scrapers are hitting your site 10,000 times a day, adding markdown isn’t going to change that at all.

ryandrake an hour ago | parent | prev | next [-]

Technically, even your browser is an agent. It says it in the HTTP: User-Agent. So is cURL. Every application the user runs is acting on the user's behalf.

cruffle_duffle 3 hours ago | parent | prev [-]

Dunno why the downvotes. I feel that is reasonable as well. Owners that block that stuff are doing so only to their detriment.