On September 15, 2026, Cloudflare flipped its default AI traffic controls for every new domain and for its entire free tier. More than 20% of the world's web domains sit behind its network — if your SaaS or e-commerce site is one of them, a setting may have changed under your feet without a clear heads-up. This isn't accidental blocking like the classic Bot Fight Mode issue that blocks your AI crawlers without you knowing: it's a new, deliberate classification layer built specifically to separate AI citation from model training.
Search, Agent, Training: three categories, three different GEO stakes
Cloudflare now sorts all AI bot traffic into three distinct families, each with its own control logic in the dashboard:
- Search — traffic that indexes your content to answer questions later, with the expectation of referral traffic in return. This is the category that directly feeds your citations in Google AI Overviews, ChatGPT answers, or Perplexity.
OAI-SearchBotandPerplexityBotfall into it. - Agent — automated behavior acting in real time on a person's behalf, "to get something done right now":
ChatGPT-Userfetching a page mid-conversation, or a browsing agent like Perplexity Comet. This is the layer that makes your site actionable for agentic AI. - Training — crawlers that absorb your content to train or fine-tune models, permanently.
GPTBotandClaudeBotare the best-known examples.
These three families don't necessarily line up with "good" or "bad" bot in the anti-spoofing sense: a perfectly authentic GPTBot is still a Training bot, whether or not blocking it makes sense for your editorial strategy.
What flips by default since September 15, 2026
Before that date, new Cloudflare domains started with all three categories broadly open. Since then, the default behavior changes specifically on pages that display ads:
| Category | Default before 09/15 | Default after 09/15 (ad pages) |
|---|---|---|
| Search | Allowed | Allowed |
| Agent | Allowed | Blocked |
| Training | Allowed | Blocked |
Cloudflare's stated reasoning: an ad is a signal that the site owner expected a human visitor on that specific page, not an agent or a training crawler. The change applies automatically to new domains and to the entire free tier; existing paying customers could opt out of this new default before September 15 through their security settings — after which the old behavior stays active until they change it themselves.
To check your setup: in the Cloudflare dashboard, under Security, each category (Search / Agent / Training) is configured independently across three levels — allow, block everywhere, or block only on ad-monetized pages. "Ad page" detection relies on Cloudflare's own internal signals (detected ad scripts, known ad-tech networks) rather than a manual declaration from you, which means the setting can apply to pages you didn't even realize were monetized through a third-party plugin or a legacy component. If your blog runs ad revenue and you're counting on ChatGPT or Perplexity citations for your GEO visibility, this setting deserves a manual check rather than blind trust in the default.
Content Signals Policy: the declarative layer that complements blocking
Alongside actual WAF-level blocking, Cloudflare has been pushing a robots.txt extension called the Content Signals Policy since September 2025. The syntax fits in one line:
User-agent: *
Content-Signal: search=yes, ai-input=yes, ai-train=no
Three signals, three distinct uses: search (classic indexing), ai-input (real-time retrieval for RAG or grounding — the mechanism behind an AI citation), and ai-train (training or fine-tuning a model). The semantics are simple: yes allows the corresponding use, no forbids it, and no signal at all neither grants nor restricts it explicitly. Cloudflare goes as far as framing these restrictions as reservations of rights under Article 4 of the EU Database Directive — a real legal anchor, but one that only matters if the crawler on the other end actually respects the declaration.
That's the structural limit: Content-Signal is declarative, not a technical enforcement mechanism. Nothing stops a crawler from simply ignoring it. Cloudflare says so itself: signals work best paired with real bot control (WAF, the Agent/Training rules above), not instead of it.
Content Signals also differs from RSL (Really Simple Licensing), which adds a monetization model (pay-per-crawl, subscription) behind the declaration rather than a plain allow/deny — useful if you're looking to get paid rather than just allow or forbid.
Who actually ships the directive today: a live check
Rather than repeating the press release, we checked public robots.txt files directly. cloudflare.com and blog.cloudflare.com both show the same line: Content-Signal: ai-train=yes, search=yes, ai-input=yes — Cloudflare opens everything on its own marketing domains, consistent with a company that benefits from being cited and trained on everywhere. stackoverflow.com, by contrast, shows Content-signal: search=no, ai-train=no — a notably more defensive stance that closes even the classic indexing signal, not just AI.
Across another dozen or so major domains we checked (NYTimes, Reuters, Wired, Forbes, BBC, CNN, Medium, GitHub, The Verge, The Economist, Ars Technica, Search Engine Journal), none show the directive yet. More than a year after launch, adoption remains niche — a sign that most SEO teams simply haven't put this on their list, despite an implementation cost close to zero.
What this means for your GEO audit
Three decisions to make separately, not a single checkbox:
- Search — leave it open if you want to exist in Google AI Overviews, ChatGPT, or Perplexity. Closing it means voluntarily opting out of the GEO race.
- Agent — depends on your business model. A SaaS that wants a Comet-style or ChatGPT Atlas agent to fill out a form or compare an offer on a user's behalf should keep this category open on product and pricing pages, while blocking it on purely ad-monetized pages where it adds nothing.
- Training — the one category where blocking costs almost nothing in immediate GEO visibility: your content can still be cited without being absorbed into a future model's training run. It's the most consensus-friendly setting to close if you want to protect your IP without sacrificing citations.
To go deeper on the overall logic behind these trade-offs, the complete guide to GEO covers the citation mechanics behind answer engines beyond this one Cloudflare case.
A 5-point checklist for your audit:
- Check your Cloudflare dashboard (Security > Bots or AI Crawl Control) for the current state of the three categories, page by page if you run mixed ad/product zones.
- If you're a new domain or on the free tier, confirm the September 15 default actually matches your intent, not just whatever Cloudflare chose for you.
- Add the
Content-Signalline to yourrobots.txtas a complement — it costs nothing and documents your position for crawlers that respect it. - Don't confuse this setting with classic accidental Bot Fight Mode blocking: these are two different mechanisms, worth auditing separately.
- Cross-check your allowed bot list against an authenticity check (reverse DNS or official IP ranges) — a spoofed
GPTBotimpersonator won't respect any of these signals anyway.
A concrete example: a B2B SaaS with a blog monetized through a display ad banner, but not its /pricing or /product pages. On a recently created Cloudflare domain, agent and training crawlers are now blocked by default on the blog posts (where the ads run), but stay allowed on the product pages — exactly the trade-off recommended above. The owner didn't have to configure anything to get this result, but would be wrong not to verify that's actually what happened on their account.
Run a free audit on SeAudit to check where your site stands on these settings, or go straight to the full report to see the level of technical detail to expect.
Key takeaways
- Since September 15, 2026, Cloudflare blocks the Agent and Training categories by default on ad-monetized pages, for new domains and the free tier — Search stays open.
- The Content Signals Policy (
Content-Signal: search=, ai-input=, ai-train=) is a distinct declarative layer, complementary to WAF blocking, not a substitute for it. - Real adoption of the directive remains rare more than a year after launch: across roughly fifteen major sites checked, only two show it.
- For your GEO, the simplest call is to close Training (protects your IP at no visibility cost) and keep Search open (without it, no AI citation is possible).
FAQ
Does the September 15, 2026 Cloudflare block apply to my existing site?
Only if you're a new domain added after that date, or on the free tier. Existing paying customers keep their previous behavior until they change the setting themselves.
Are Content Signals Policy and Bot Fight Mode the same thing?
No. Bot Fight Mode and the Agent/Training rules actually block traffic at the WAF level. Content Signals Policy is a declarative line in robots.txt expressing a preference, with no technical blocking power of its own.
Should I block all AI crawlers to protect my content?
Not if you want to exist in AI answers. Blocking Search takes you out of the citation race entirely. Blocking Training alone protects your content from model training without hurting your GEO visibility.
How do I know if a crawler claiming to be GPTBot or ClaudeBot is genuine?
Through reverse DNS verification or the official IP lists each provider publishes — the full method is covered in our dedicated guide to AI bot anti-spoofing verification.
Check in 2 minutes whether your Cloudflare settings and robots.txt line up with your GEO strategy: run your free audit, or go straight to the full report for a detailed technical diagnosis.
