| ▲ | qbane 5 hours ago |
| > And yes, my site gets its data by scraping those public documents. So I'm a scraper writing a blog post complaining about scrapers. I'm aware of how that sounds. |
|
| ▲ | ihuman 4 hours ago | parent | next [-] |
| There's a difference between someone running a scraping tool occasionally and bots constantly and rapidly re-scraping the same site over and over again |
| |
| ▲ | Lalabadie 4 hours ago | parent | next [-] | | The author's website is responsible for storing its own data. AI services currently treat the entire web as their storage and cache layer. | |
| ▲ | qbane 4 hours ago | parent | prev [-] | | This is an important context: the site is more likely to be targeted by scrapers because it is a curated collection of scraped information. | | |
| ▲ | 0cf8612b2e1e 4 hours ago | parent | next [-] | | Do the bots care? Seemingly very little intelligence in many of them. Could be as simple as the site has more pages, so more traffic. Loot first, ask questions later. | |
| ▲ | everybodyknows 4 hours ago | parent | prev [-] | | A curated collection of people who give away money. Was ever sweeter honey ever found in a pot? |
|
|
|
| ▲ | aeturnum 4 hours ago | parent | prev | next [-] |
| Similar to the dose making the poison - the thing that jumped out at me in this blog was the ratio of scraping to visits. Unless OP is scraping thousands of times a day I don't really think they're in the same class as the bots they are blocking. |
| |
| ▲ | nickgray 4 hours ago | parent [-] | | OP here: I'm not scraping thousands of times per day! Usually just a few times per year. | | |
| ▲ | jader201 2 hours ago | parent [-] | | But you could just be one of thousands of bots targeting the same sources you’re scraping. |
|
|
|
| ▲ | ethersteeds 5 hours ago | parent | prev | next [-] |
| Who scrapes the scrapemen? |
|
| ▲ | alansaber 4 hours ago | parent | prev [-] |
| Live by the scraper, die by the scraper |