Bunny Honey ClubBunny Honey/blog
Work with us
← back to indexblog / seo / should-you-block-ai-crawlers-from-your-website
SEO

Should You Block AI Crawlers From Your Website?

Cloudflare defaults to blocking AI crawlers on Sept 15. The same setting can quietly take Googlebot down with it. Here's what to check.

AH
Arthur HofFounder, Bunny Honey Club AI
publishedAug 27, 2026
read6 min
Should You Block AI Crawlers From Your Website?

On September 15, 2026, Cloudflare starts blocking AI crawlers by default on every new domain. That alone is a mid-tier CDN policy update. The part worth five minutes of your week is that the same setting can quietly take Googlebot down with

On September 15, 2026, Cloudflare starts blocking AI crawlers by default on every new domain. That alone is a mid-tier CDN policy update. The part worth five minutes of your week is that the same setting can quietly take Googlebot down with it.

Search Engine Journal flagged it directly: block Cloudflare's "Training" crawler category and you also block Googlebot, Applebot, and Bingbot, because all three run mixed-use crawlers that do search indexing and AI training in a single pass (Search Engine Journal). If a site already flipped Cloudflare's older, blunter "Block AI Bots" toggle at some point, that toggle now falls under the same rule.

Roughly a quarter of the web sits behind Cloudflare, 24.9% of all websites as of this month according to W3Techs. That's not a niche edge case. If your site runs through Cloudflare and you've ever touched a bot-blocking setting to keep AI scrapers off your content, this is worth checking before the deadline, not after your organic traffic drops and you're troubleshooting blind.

Cloudflare now sorts every crawler into three buckets

Cloudflare used to give site owners one blunt lever: allow AI bots, or block AI bots. As of this year, that's gone. Every crawler hitting a Cloudflare-fronted site now gets sorted into one of three categories, described on Cloudflare's own blog (Cloudflare):

  • Search: indexes your content to answer questions about it later. This is classic Google-style crawling.
  • Agent: a real-time fetch acting on someone's behalf right now, the kind of request ChatGPT or Perplexity makes when a user asks it a question and it goes to pull your page to answer.
  • Training: a crawler pulling your content to train or fine-tune a model, with no real-time user request behind it.

This distinction is the whole story. A site owner who wants to stay visible in Google, stay eligible to get cited by AI answer engines, and still keep their content out of somebody's next foundation-model training run needs all three settings pointed differently. One toggle can't do that. Three can.

The old block-everything switch takes Googlebot with it

Here's the mistake that actually costs money. A lot of site owners, or the agency that set up their site, flipped Cloudflare's "Block AI Bots" toggle sometime in the last year or two as a reasonable-sounding move to stop AI companies scraping their content for free. That instinct wasn't wrong. The execution now has a side effect nobody warned them about.

Google, Apple, and Microsoft all run crawlers that do double duty: they index your pages for search results, and the same crawl also feeds each company's own AI training pipeline. Cloudflare's system doesn't split that behavior apart per crawler. It applies the strictest rule that matches. Block Training, and any crawler that does Training as part of its job gets blocked entirely, search function included.

If someone toggled "block AI bots" eight months ago to keep your content out of an AI training set, and hasn't looked at that setting since, you don't actually know whether Google is still fully crawling your site. That's not a hypothetical risk. It's a five-minute dashboard check you're overdue for.

the setting nobody remembers changing

Losing Googlebot access isn't a slow bleed you'll notice next quarter. It's a network-level block, not a robots.txt suggestion Google can choose to ignore. Pages stop getting recrawled, updates stop propagating to the index, and rankings erode with no error message pointing at the cause.

September 15 sets the default for every new domain

The specific deadline matters less than what it signals. Starting September 15, 2026, any domain newly onboarding to Cloudflare gets Training and Agent crawlers blocked by default on pages that carry advertising, while Search crawlers stay allowed by default across the board (Cloudflare). Existing domains aren't touched by the new default. They keep whatever settings they already have, which is exactly why "I didn't change anything" isn't the same as "nothing changed."

24.9%of all websites run behind Cloudflare (W3Techs, Aug 2026)
Sept 15, 2026new domains default to blocking Training + Agent bots on ad-carrying pages
118:1–50,000:1range of AI crawl-to-referral ratios in Cloudflare's own dashboard data
3crawler categories every bot now gets sorted into: Search, Agent, Training

There's a second layer under the dashboard toggle worth knowing about: Cloudflare's Content Signals Policy, an extension to robots.txt with three explicit signals, search, ai-input, and ai-train. When you set bot categories in the dashboard, Cloudflare writes matching directives into your robots.txt automatically. If someone on your team or a past developer hand-edited robots.txt separately, those two layers can now disagree with each other, and it's worth a five-minute look to confirm they don't.

Blocking training doesn't have to mean losing citations

This is where most coverage of the Cloudflare change stops short, and it's the part that actually matters for a business trying to get found by AI answer engines. Blocking Training doesn't block Agent. Those are separate settings for a reason.

A business that wants to show up when someone asks ChatGPT or Perplexity a question in their category needs Agent crawlers allowed, because that's the live, in-the-moment fetch those tools make to pull a page and cite it in an answer. Training is a different thing entirely: a slower, background harvest with no specific user question behind it, feeding whatever model gets built next. You can allow one and block the other. Most site owners don't realize that's even an option, because the old single toggle never let them.

Our own take, and the setting we ship by default on client sites: allow Search, allow Agent, block Training. That gets you full visibility in Google, eligibility to get cited the moment someone asks an AI assistant a relevant question, and a real opt-out of your content becoming free training data for a model you'll never see a cent from. It's a close cousin of the crawl-directive work behind getting a site discovered by Brave's index, which runs on the same logic: being crawlable by the right systems is what makes you eligible to be quoted by an assistant later, whether or not anyone ever visits your site through the search box directly.

What to actually check in your dashboard this week

Five checks, in order of how much they can cost you if skipped:

  1. Open Cloudflare, go to Security, then Bots or AI Crawl Control. Confirm you're looking at the current three-category system, not assuming settings from a year ago still apply the same way.
  2. If the legacy "Block AI Bots" toggle is on, don't leave it there on autopilot. It's now covered by the new mixed-crawler rule. Decide deliberately whether that's still what you want, since it may be quietly limiting Google along with everything else.
  3. Set Search to allowed, full stop. There's rarely a good reason to block the crawler that puts you in Google results.
  4. Decide Agent and Training separately, not as one bundle. Agent on, Training off is the default we'd recommend for most content-driven small business sites chasing both search rankings and AI citations.
  5. If your site displays ads, know the September 15 default doesn't override settings you've already made. It only applies where nobody has touched these controls yet.

None of this requires touching code. It's a dashboard, three category toggles, and a decision that was previously being made for you by whoever set up your Cloudflare account, possibly years before ChatGPT existed. Our AEO playbook for ranking in ChatGPT, Perplexity, and Gemini covers the content side of showing up in AI answers. This is the infrastructure side, the part that decides whether the content ever gets read by the systems doing the answering in the first place, alongside the technical groundwork in building a blog that ranks and gets cited by LLMs.

We check this on every site we launch, not as an afterthought bolted on after a client complains about a traffic dip. Bot categories, robots.txt content signals, and canonical structure all get set correctly before the site goes live, because a five-day build that quietly blocks Google in month three isn't actually fast. If your current site was handed off by someone who hasn't looked at it since launch, that's the audit we run before touching anything else.

The bigger pattern here matters more than Cloudflare's specific deadline. Every major CDN and bot-management vendor is going to end up drawing the same three-way line between search, live AI fetches, and model training, because "AI bot" stopped being one thing the moment agents started browsing the web on a person's behalf in real time. Cloudflare just moved first and loudest. Treat September 15 as the reminder to check the setting, not as the only day it'll ever matter.

— filed underSEOAIStrategy
— share
— keep reading

Three more from the log.

How to get discovered by Brave Search in 2026
002 · SEO

How to get discovered by Brave Search in 2026

Brave runs its own index, and getting into it is now an AI-visibility play. Here's how Brave Search actually discovers pages, and how to get yours in.

Jul 09, 2026 · 7 min