| ▲ | hyperhello 6 hours ago |
| If the scrapers are going to get it anyway, put the data up as a zip somewhere. |
|
| ▲ | jjgreen 5 hours ago | parent | next [-] |
| It would not make any difference. |
| |
| ▲ | codemonkey-zeta 5 hours ago | parent [-] | | Indeed, the article mentions Wikipedia experiencing similar scraping pains, even though they already DO have bulk data available. | | |
| ▲ | HeatrayEnjoyer 4 hours ago | parent [-] | | Who are running these bots? I presume developers at all of the frontier labs know (or at least would know to look for) Wikipedia has bulk APIs for automated access. Unnecessary scraping increases their workload/costs too, so why in 2026 is this still a problem? | | |
| ▲ | esseph 4 hours ago | parent [-] | | Black market and gray market data. All the firms want data. All the other firms want data. The banks want data. The other criminals also want data for their crimes and schemes. Oh insurance companies, and the ATS systems. Everybody wants as much data as they can get and they don't care how they get it. | | |
| ▲ | esseph 2 hours ago | parent [-] | | This data selling also happens with leaks of all kinds like medical data, often to current or future employers, health insurance companies, etc. |
|
|
|
|
|
| ▲ | antisthenes 4 hours ago | parent | prev [-] |
| > Read the Docs, a non-profit that hosts documentation for open-source software, who watched a single crawler download 73 terabytes of zipped HTML in one month, costing it over $5,000 in bandwidth From the article. Not the same site, but an example of the same issue. |