| ▲ | simonw 5 hours ago |
| The big news here is that Googlebot will be blocked from September 15th onwards by one the "block training" policies, because Google use the same crawler infrastructure for their search index AND for training Gemini: > Another change that will apply on September 15 is that multi-purpose crawlers (specifically those that combine Search with Training) will be allowed/blocked according to all of their behaviors, in line with our call for transparency for website owners. Since the defaults will be enforced by the most restrictive applicable rules, multi-purpose crawlers such as Googlebot, Applebot, and BingBot will be blocked by customers who have selected to block Training (either through the new options to manage AI traffic, or through the legacy Block AI bots service). |
|
| ▲ | dannyw 2 hours ago | parent | next [-] |
| Good. Google's approach here is manifestly predator, unfair, and IMO illegal. They deserve to be in court for this behaviour, and mandating owners give consent for AI training or drop out of Google; which is just a non-starter because they're a search monopoly. That's exactly what antitrust laws are supposed to do, and I hope at least EU regulators take action. Every single Googlebot crawl in your access logs is a trace for damages. |
| |
| ▲ | troyvit an hour ago | parent [-] | | I think it's bad, because everybody is desperate to hold onto every last bit of google search traffic they can, so they're going to allow training to do so. Google's predatory, unfair and illegal actions will continue as they have with a few $100 million slaps in the wrist from the EU and a few more white house dinners for their CEO. |
|
|
| ▲ | miohtama 20 minutes ago | parent | prev | next [-] |
| People will use something for search and something needs to index pages, either for LLM or old school search engine. |
|
| ▲ | jofzar 5 hours ago | parent | prev | next [-] |
| We had googlebot blast a random customer system and almost cause an outage, this is when I first learnt that google will use it for AI training also. It's honestly kind of frustrating also because you then search on it and theres (was) nothing on how you are meant to "correctly" tell google to fuck off, and not use it like that. |
| |
| ▲ | bobbiechen 22 minutes ago | parent | next [-] | | The vast majority of Googlebot user agents are lying. Real Googlebot is pretty well behaved in my experience. You should use reverse DNS or IP lists to check: https://developers.google.com/crawling/docs/crawlers-fetcher... | |
| ▲ | 20k 4 hours ago | parent | prev [-] | | Google's web scraping functionality has been acting as a ddos for more than two decades. I've seen literally hundreds of reports of them attacking websites and taking them down, where there's nothing you can do but accept the traffic, or get delisted This is unfortunately nothing new. There's no correct way to tell them to fuck off, they do not care, and they never will do. People have even taken them to court over this | | |
| ▲ | weird-eye-issue 28 minutes ago | parent [-] | | If a site cannot handle traffic from the real Googlebot that is a serious issue with the site itself since it's actually pretty conservative Also I should note there are lots of fake Googlebots... |
|
|
|
| ▲ | inigyou 5 hours ago | parent | prev [-] |
| [flagged] |
| |
| ▲ | Cider9986 4 hours ago | parent | next [-] | | Why should I use something other than Cloudflare pages for a simple app landing page? | | | |
| ▲ | ajmurmann 4 hours ago | parent | prev [-] | | Why is this? | | |
| ▲ | ceejayoz 4 hours ago | parent [-] | | It's a planet-scale MITM? | | |
| ▲ | TurdF3rguson 3 hours ago | parent | next [-] | | It's a cache. My tiny websites couldn't survive getting hammered by AI bots without them. | |
| ▲ | dbbk 3 hours ago | parent | prev [-] | | So you're against all CDNs? | | |
| ▲ | 3 hours ago | parent | next [-] | | [deleted] | |
| ▲ | fc417fc802 3 hours ago | parent | prev [-] | | A CDN doesn't necessarily have to perform a MitM. We really need more nuanced terminology to distinguish the various approaches. | | |
| ▲ | gruez 3 hours ago | parent | next [-] | | Right, but practically speaking all CDNs are MITMs. If you're against cloudflare you should be against cloudfront, akamai, etc. as well. | |
| ▲ | edaemon 3 hours ago | parent | prev | next [-] | | How would they cache and serve responses without decrypting the traffic? | |
| ▲ | sandeepkd 3 hours ago | parent | prev [-] | | Ideally yes, the TLS termination does not need to happen for caching purposes. Challenge is that in practice every business wants to be sticky and try to provide more functionalities which do require TLS termination. Most people either trust CDN's or they do not understand MitM so it does not concerns them. Plus they are getting certificate management and DDOS prevention capabilities. |
|
|
|
|
|