Remix.run Logo
TeMPOraL a day ago

Where on Earth do people get the belief that:

- It's the SOTA companies doing it?

- Scrapers are doing it for training data?

Those are two assumptions I see in posts and threads around Anubis, that are taken at faith, and never once substantiated.

shakna a day ago | parent | next [-]

Because Anthropic already admitted it? [0]

[0] https://www.ft.com/content/07611b74-3d69-4579-9089-f2fc2af61...

inigyou 17 hours ago | parent [-]

It says "accused of"

shakna 12 hours ago | parent [-]

... It also has a quote from Anthropic, admitting they backed off their scraper after a robots.txt update.

inigyou 6 hours ago | parent [-]

Nobody who gives a shit about their software actually working obeys robots.txt. most robots.txt block everything except googlebot, which is an insane policy and you don't have to follow it.

shakna 3 hours ago | parent [-]

Which isn't all that relevant a discussion here. Anthropic saying they did so in this case, means that they admit they were the scraper flooding the domain.

mschuster91 a day ago | parent | prev | next [-]

> It's the SOTA companies doing it?

There are more than just the American top dogs (OAI, Anthropic, SpaceX, Facebook)... especially the Chinese government with all its infinite cash resources and next to zero ethical constraints.

I don't trust the US top dogs at all, but I think the fear of discovery alone would lead them to not use "residential proxy" services. Non-US/EU entities however... who cares?

figglestar a day ago | parent [-]

Why would they directly use a proxy service? I'd just launder the data scraping through some third party company that I could slough off if it ever turned into a news story. Not that anything would happen to them if they directly used these services anyway.

inigyou a day ago | parent | prev [-]

What's your alternative hypothesis?