Cloudflare Disallow AI Training: A Crawler Migration Map for Search Teams
Cloudflare Disallow AI Training separates training preferences from search access, but Training Block can stop mixed-use search crawlers. Map provider support, Bing's early 2027 gap and rollback checks without confusing crawl access with training compliance.
What changed: Disallow AI Training is not the same as Block
Cloudflare Disallow AI Training publishes training preferences while keeping Accountable mixed-use search crawlers allowed. Training Block can stop those same crawlers. That is the practical distinction in Cloudflare's September 15, 2026 release. Bing's robots.txt no-training support, meanwhile, remains targeted for early 2027.
Start with the outcome you want. Keeping pages available to search does not settle whether their content can be used for training, and neither decision settles AI summary inclusion. Each needs a supported control. A configured dashboard is a poor definition of done. Check whether the intended crawler can still reach the pages that matter.
The release describes how the controls work; it does not establish that sites lost rankings when the release shipped.
Three outcomes that need separate controls
With Disallow AI Training, Bot Preference Sync publishes applicable training preferences through robots.txt. Accountable mixed-use crawlers remain allowed for search. Other training crawlers are blocked, including training-only routes from Amazon, Anthropic, Meta and OpenAI. Their separate search crawlers are different targets.
Google-Extended and Applebot-Extended are usage-control tokens, not substitute names for Googlebot and Applebot. Declining training use does not automatically decline every generated answer. The SEO, AEO and GEO guide covers the wider context; here the focus is migration decisions and the evidence needed to defend them.
Map existing Cloudflare settings before changing search access
Settings and effects: Allow, Disallow and Block
Record the zone's current values before changing anything. An Allow selection does not override a separate setting or an unrelated WAF rule. Training Block and Block on pages with ads now include mixed-use Applebot, Bingbot and Googlebot, so an assumption about what counts as AI traffic can lead to unintended search denial.
The following settings and migration outcomes are documented in Cloudflare's September release.
| Existing setting or proposed action | Documented outcome | Search implication | Remaining check |
|---|---|---|---|
| Legacy Block AI disabled | Search Allow, Training Allow, Agent Allow | No category block added by migration | Existing rules and robots.txt |
| Legacy Block AI Block or ad-only block | Search Allow, Training Disallow AI Training, Agent Block on pages with ads | Mixed-use search remains allowed by this training selection | Provider support |
| Previously granular Training Block or ad-only block | Training becomes Disallow; Search and Agent selections stay unchanged | Preserved Search restrictions still matter | Actual migrated values |
| Select Training Disallow AI Training | Preferences for Accountable mixed-use crawlers; other training crawlers blocked | Search can remain reachable | Bing gap and WAF |
| Select Training Block | All training crawlers blocked, including mixed-use crawlers | Can deny search access | Whether denial is intended |
| Select Training Block on pages with ads | Training crawlers blocked on detected ad pages | Those pages can lose search access | Ad detection and affected paths |
Automatic migration is not a new-domain preset
Cloudflare says most existing settings migrate automatically. The new meaning of Block does not mean every previously protected site suddenly blocks Googlebot. Check the migrated values before deciding whether a change is needed.
New-domain recommendations are a separate topic. For ad-supported sites, the preset offers Search Allow, Training Disallow AI Training and Agent Block on pages with ads. For sites that are not ad-supported, the preset offers Allow for all three. Bot Preference Sync is enabled in both cases, and owners can change those recommendations.
The controls are available on all plans at domain or zone scope. There is no ad-only Disallow AI Training option and no Agent Disallow option. Testing individual paths does not turn the controls into path-scoped settings.
The Block AI Bots documentation still showed July-era wording about September defaults when checked. Where the accounts conflict, this article attributes the newer migration behavior to the September release. Check the dashboard before following an assumed click path.
Check provider capabilities, especially Bing's early 2027 gap
Accountable status covers more than mechanisms that are already live. It requires training opt-out, summary opt-out, URL-level transparency and assurance that training opt-out does not affect traditional search, either available now or covered by commitments. Each mechanism still needs its own check.
| Provider | Search and training distinction | Summary mechanism | Commitment or gap |
|---|---|---|---|
| Googlebot handles search; Google-Extended controls specified training and grounding uses | Search has snippet and indexing controls | Cloudflare reports additional Extended URL transparency expected within weeks | |
| Apple | Applebot crawls; Applebot-Extended controls foundation-model training use | nosnippet covers specified generative answers | Cloudflare reports URL-level inspection work for 2027 |
| Bing | Bingbot remains allowed under Disallow; automatic robots.txt no-training delivery is unsupported | Historical Bing Chat controls couple some answer and training restrictions | Robots.txt no-training support targets early 2027 |
Google: Google-Extended is not Googlebot
Google's crawler documentation describes Google-Extended as a product token without a separate HTTP request user-agent string. It controls specified Gemini training and grounding uses. Google says it does not affect Search inclusion or ranking.
For Google AI Overviews and AI Mode, Search AI guidance points to Googlebot as the crawl-access control and lists nosnippet, data-nosnippet, max-snippet and noindex as ways to limit information shown. These controls have different consequences. In particular, noindex is not a training-only preference and should not be treated like one.
After changing preview controls, use Search Console URL Inspection to see the HTML Googlebot received, then allow time for recrawling and processing. Do not look for a separate Google-Extended crawler in access logs as proof that the preference worked, because the documentation does not describe one.
Cloudflare also describes a generative-search portal toggle. The Google documentation checked here does not independently establish that exact universal interface. This map therefore does not prescribe it or treat promised URL transparency as already shipped.
Apple: Applebot-Extended is not a search-crawler block
Apple says Applebot-Extended does not crawl pages. Disallowing it controls training use of content collected by Applebot; pages can still remain discoverable in search.
Apple separately documents nosnippet for specified generative answers, with effects on descriptions and web answers. Review those presentation effects before rollout, because they touch how content appears rather than only how it is used for training. Cloudflare attributes Apple's URL-inspection work to next year, meaning 2027, without giving a launch date.
Bing: a preference gap is not a removal recommendation
Until Bing's targeted early 2027 support arrives, Disallow AI Training does not automatically deliver a no-training preference through robots.txt. Record that state plainly as unresolved, not implemented.
Microsoft's September 2023 announcement says NOARCHIVE excludes content from Bing Chat answers and prospective foundation-model training while retaining search-result eligibility. NOCACHE takes precedence when both appear. Those are historical documented semantics, not independently verified coverage of every current Microsoft surface.
Cloudflare mentions NOARCHIVE and removal tools. Do not use removal as a training-only substitute, because indexing consequences need separate verification. This article does not prescribe removal procedures. Cloudflare's unified control over the amount of content in summaries is also a goal for early 2027, not a September feature.
Use the Crawler Intent Migration Map
Record baseline and desired outcomes
The Crawler Intent Migration Map is a working template introduced in this article, not a Cloudflare feature or a claimed client methodology. Use it to connect the recorded zone state with the outcomes you want. Alongside each proposed control, keep the provider's support status and the evidence needed to accept the change. Include the rollback path.
Export settings where possible or record exact Search, Training and Agent values, Bot Preference Sync state, the public robots.txt response and relevant WAF rules. Timestamp the record. Redact sensitive rule detail before it leaves the operational team.
Write down whether search should remain open, then specify the training preference. Give summary policy its own decision. An unsupported preference remains a gap even when the dashboard appears configured; the setting alone proves neither provider support nor compliance.
Choose representative indexable paths, including ad-serving and non-ad-serving pages where relevant. Capture available crawl and index observations before the change. Agree which unexpected denials will trigger rollback and name the person who can restore settings. Record any unrelated protections that must remain intact.
This illustrative record is not a Cloudflare API payload or a production observation. It proposes declining training while leaving summaries undecided.
{
"zone": "example.com",
"observedAt": null,
"desiredOutcomes": {"search": "allow", "training": "decline", "summaries": "undecided"},
"currentControls": "not_checked",
"proposedControls": {"Search": "Allow", "Training": "Disallow AI Training", "Agent": "retain_recorded_baseline"},
"providerSupport": {"bingRobotsTraining": "target_early_2027_not_shipped"},
"evidence": {"robotsTxt": "not_checked", "verifiedCrawlerAccess": "not_checked"},
"rollback": "restore_recorded_baseline_after_unintended_search_denial"
}Inspect robots.txt and crawler access
Bot Preference Sync prepends generated content to existing robots.txt material. Inspect the entire response, not only the newly generated lines. Check applicable user-agent groups and directive precedence against each crawler operator's rules. The order in which lines appear is not an acceptance test. Cloudflare also says this category-wide sync does not directly read individual custom rules with more complex logic. Do not assume a custom WAF exception has been translated into robots.txt.
| Migration test | Evidence | Failure condition | Response |
|---|---|---|---|
| Baseline captured | Controls, rules and robots.txt | Previous state cannot be reconstructed | Defer change |
| Preference expressed | Served robots.txt and provider token | Missing or conflicting directive | Restore prior configuration; investigate sync |
| Search paths reachable | Verified requests, paths and statuses | Newly denied intended search crawler | Restore changed controls; identify rule |
| Redirects work | Response chains and destinations | Unexpected or inaccessible destination | Reverse responsible change |
| Other protections preserved | WAF events and rule evidence | Unexplained broad bypass needed | Stop; avoid blanket exceptions |
| Provider gap recorded | Dated source and support status | Promised support marked implemented | Correct record; defer unsupported outcome |
A copied Googlebot user-agent can inspect a response, but it cannot prove access by the real Googlebot. Use documented identity verification where available. Cloudflare's crawler-management documentation describes user-agent identification on the free plan and more thorough detection on upgraded plans. Label confidence accordingly.
The same documentation warns that unsuccessful requests can come from other rules or response errors. Diagnose the responsible control before changing training policy. Otherwise, a team can weaken the wrong rule and still leave the crawler problem unsolved.
Measure outcomes and retain rollback
The AI search measurement stack gives broader measurement context. This migration needs a narrower record that ties each observation to the specific setting change.
| Measurement | Evidence | Interpretation limit |
|---|---|---|
| Preference delivery | Timestamped robots.txt and applicable directives | Publication is not compliance |
| Search reachability | Verified requests by provider, path and status | Untested paths remain unknown |
| Crawl and index behavior | Search Console and Bing Webmaster observations | Revisit timing and reporting lag matter |
| Search and summary outcomes | Search performance and separate visibility observations | Change does not establish causality |
| Commercial outcomes | Relevant referrals and conversions | Demand and other releases confound attribution |
Google includes AI Overviews and AI Mode in Search Console's overall Web reporting. Do not label an aggregate change as an isolated summary effect. Choose observation windows around actual crawler visits and processing, not an assumed immediate verdict.
If intended search access is newly denied, restore the recorded settings and recheck affected paths. Restoration cannot guarantee immediate recrawl or index recovery. The Cloudflare Radar evidence-trace test provides related evidence-handling context for a different product, not this control rollout.
Common mistakes when separating search from training
The central mistake is selecting Training Block while expecting a mixed-use crawler to continue searching. Inspect migrated values rather than relying on the old label's meaning. Search Allow does not cancel independent restrictions elsewhere in the configuration.
Agent settings cannot solve a missing training mechanism. An ad-only Disallow option cannot solve it either, because none is offered. Accountable status does not close Bing's gap, and a training directive is not a summary opt-out.
Be careful about what the evidence actually shows. A copied user-agent string tells you little about crawler identity, while robots.txt records a published preference without establishing that training stopped. Stable rankings can also be reassuring for the wrong reason: they do not tell you whether every important path remains reachable.
Avoid universal tag recipes. Microsoft's NOARCHIVE and NOCACHE interaction shows how combined directives can change the result. Verify product scope before deployment, especially before using removal tools.
Caveats: what this map cannot prove
Publishing a preference does not prove compliance. It also cannot remove content from existing models or establish how earlier collections were used. Evidence about crawl enforcement cannot, on its own, answer questions about identity or downstream data use.
Budget for implementation effort, changing documentation, cached responses, revisit timing and incomplete logs. Keep evidence retention and access limited where privacy requires it.
No setting guarantees rankings, traffic, citations or conversions. This article is based on public-source review, not a live Cloudflare configuration change, client deployment or controlled SEO experiment. Keep unsupported outcomes visible instead of calling the migration complete for every site.
Key Takeaways
- 1Disallow AI Training combines robots.txt preferences for Accountable mixed-use crawlers with blocking of other training crawlers.
- 2Selecting Training Block can deny mixed-use search crawlers, but existing settings generally migrate instead of automatically adopting the new blocking behavior.
- 3Accountable includes capabilities and commitments; assess each provider mechanism separately.
- 4Bing's robots.txt no-training support targets early 2027, so Disallow AI Training does not automatically convey that preference to Bing today.
- 5Treat summaries separately, verify crawler access and retain rollback evidence without claiming proof of training non-use.
Conclusion
Keep search access open deliberately, with a recorded baseline and clear rollback triggers. Training Block can deny mixed-use search crawlers; Bing's robots.txt no-training support remains targeted for early 2027. Neither a dashboard label nor a provider commitment closes that gap. Set summary policy separately and leave unsupported outcomes marked unresolved. If useful, Optijara can help scope the settings map and measurement requirements before implementation.
Frequently Asked Questions
What does Cloudflare Disallow AI Training do?
It publishes applicable robots.txt training preferences through Bot Preference Sync for Accountable mixed-use crawlers and blocks other training crawlers. Provider support gaps remain; publication is not proof that training stopped.
Can Training Block prevent Googlebot or Bingbot from reaching pages?
Yes. The new Training Block semantics include mixed-use Googlebot, Bingbot and Applebot. Ad-only blocking can affect detected ad pages. Existing settings generally migrate, so this does not mean sites automatically lost search access.
Does disallowing Google-Extended or Applebot-Extended block traditional search?
These usage-control tokens are distinct from Googlebot and Applebot. Google and Apple document training controls separately from traditional search. Neither token is a universal AI summary opt-out.
Does Disallow AI Training send Bing a no-training preference today?
Not automatically through robots.txt. Cloudflare's September 15, 2026 release targets Bing support for early 2027. Historical NOARCHIVE guidance requires current-scope review; removal tools are not training-only substitutes.
Does Accountable mean every control is already available?
No. Accountable includes current capabilities and time-bound commitments. Verify each mechanism. Cloudflare's unified control over summary amounts is a future goal, not a feature shipped in this release. Record provider gaps, check verified crawler access and define rollback triggers rather than treating the designation as proof of implementation.
Sources
- https://blog.cloudflare.com/accountable-mixed-use-ai-crawlers/
- https://blog.cloudflare.com/bot-preference-sync/
- https://support.apple.com/en-us/119829
- https://developers.google.com/crawling/docs/crawlers-fetchers/google-common-crawlers
- https://developers.google.com/search/docs/appearance/ai-features
- https://blogs.bing.com/webmaster/september-2023/Announcing-new-options-for-webmasters-to-control-usage-of-their-content-in-Bing-Chat
- https://developers.cloudflare.com/bots/additional-configurations/block-ai-bots/
- https://developers.cloudflare.com/ai-crawl-control/features/manage-ai-crawlers/
Written by
Hamza DiazHamza Diaz is the founder of Optijara, where he builds practical AI agents, automation systems, and Copilot workflows for service businesses. He writes about AI operations, agent strategy, and real-world implementation for teams that want usable systems instead of hype.
