Remix.run Logo
qbane 5 hours ago

> And yes, my site gets its data by scraping those public documents. So I'm a scraper writing a blog post complaining about scrapers. I'm aware of how that sounds.

ihuman 4 hours ago | parent | next [-]

There's a difference between someone running a scraping tool occasionally and bots constantly and rapidly re-scraping the same site over and over again

Lalabadie 4 hours ago | parent | next [-]

The author's website is responsible for storing its own data. AI services currently treat the entire web as their storage and cache layer.

qbane 4 hours ago | parent | prev [-]

This is an important context: the site is more likely to be targeted by scrapers because it is a curated collection of scraped information.

0cf8612b2e1e 4 hours ago | parent | next [-]

Do the bots care? Seemingly very little intelligence in many of them. Could be as simple as the site has more pages, so more traffic.

Loot first, ask questions later.

everybodyknows 4 hours ago | parent | prev [-]

A curated collection of people who give away money. Was ever sweeter honey ever found in a pot?

aeturnum 4 hours ago | parent | prev | next [-]

Similar to the dose making the poison - the thing that jumped out at me in this blog was the ratio of scraping to visits. Unless OP is scraping thousands of times a day I don't really think they're in the same class as the bots they are blocking.

nickgray 4 hours ago | parent [-]

OP here: I'm not scraping thousands of times per day! Usually just a few times per year.

jader201 2 hours ago | parent [-]

But you could just be one of thousands of bots targeting the same sources you’re scraping.

ethersteeds 5 hours ago | parent | prev | next [-]

Who scrapes the scrapemen?

alansaber 4 hours ago | parent | prev [-]

Live by the scraper, die by the scraper