← Back to Blog
Marketing & Growth

Cloudflare Disallow AI Training: A Crawler Migration Map for Search Teams

Cloudflare Disallow AI Training separates training preferences from search access, but Training Block can stop mixed-use search crawlers. Map provider support, Bing's early 2027 gap and rollback checks without confusing crawl access with training compliance.

Written by Hamza Diaz
September 16, 202610 min read8 views

What changed: Disallow AI Training is not the same as Block

Cloudflare Disallow AI Training publishes training preferences while keeping Accountable mixed-use search crawlers allowed. Training Block can stop those same crawlers. That is the practical distinction in Cloudflare's September 15, 2026 release. Bing's robots.txt no-training support, meanwhile, remains targeted for early 2027.

Start with the outcome you want. Keeping pages available to search does not settle whether their content can be used for training, and neither decision settles AI summary inclusion. Each needs a supported control. A configured dashboard is a poor definition of done. Check whether the intended crawler can still reach the pages that matter.

The release describes how the controls work; it does not establish that sites lost rankings when the release shipped.

Three outcomes that need separate controls

With Disallow AI Training, Bot Preference Sync publishes applicable training preferences through robots.txt. Accountable mixed-use crawlers remain allowed for search. Other training crawlers are blocked, including training-only routes from Amazon, Anthropic, Meta and OpenAI. Their separate search crawlers are different targets.

Google-Extended and Applebot-Extended are usage-control tokens, not substitute names for Googlebot and Applebot. Declining training use does not automatically decline every generated answer. The SEO, AEO and GEO guide covers the wider context; here the focus is migration decisions and the evidence needed to defend them.

Map existing Cloudflare settings before changing search access

Settings and effects: Allow, Disallow and Block

Record the zone's current values before changing anything. An Allow selection does not override a separate setting or an unrelated WAF rule. Training Block and Block on pages with ads now include mixed-use Applebot, Bingbot and Googlebot, so an assumption about what counts as AI traffic can lead to unintended search denial.

The following settings and migration outcomes are documented in Cloudflare's September release.

Existing setting or proposed actionDocumented outcomeSearch implicationRemaining check
Legacy Block AI disabledSearch Allow, Training Allow, Agent AllowNo category block added by migrationExisting rules and robots.txt
Legacy Block AI Block or ad-only blockSearch Allow, Training Disallow AI Training, Agent Block on pages with adsMixed-use search remains allowed by this training selectionProvider support
Previously granular Training Block or ad-only blockTraining becomes Disallow; Search and Agent selections stay unchangedPreserved Search restrictions still matterActual migrated values
Select Training Disallow AI TrainingPreferences for Accountable mixed-use crawlers; other training crawlers blockedSearch can remain reachableBing gap and WAF
Select Training BlockAll training crawlers blocked, including mixed-use crawlersCan deny search accessWhether denial is intended
Select Training Block on pages with adsTraining crawlers blocked on detected ad pagesThose pages can lose search accessAd detection and affected paths

Automatic migration is not a new-domain preset

Cloudflare says most existing settings migrate automatically. The new meaning of Block does not mean every previously protected site suddenly blocks Googlebot. Check the migrated values before deciding whether a change is needed.

New-domain recommendations are a separate topic. For ad-supported sites, the preset offers Search Allow, Training Disallow AI Training and Agent Block on pages with ads. For sites that are not ad-supported, the preset offers Allow for all three. Bot Preference Sync is enabled in both cases, and owners can change those recommendations.

The controls are available on all plans at domain or zone scope. There is no ad-only Disallow AI Training option and no Agent Disallow option. Testing individual paths does not turn the controls into path-scoped settings.

The Block AI Bots documentation still showed July-era wording about September defaults when checked. Where the accounts conflict, this article attributes the newer migration behavior to the September release. Check the dashboard before following an assumed click path.

Check provider capabilities, especially Bing's early 2027 gap

Accountable status covers more than mechanisms that are already live. It requires training opt-out, summary opt-out, URL-level transparency and assurance that training opt-out does not affect traditional search, either available now or covered by commitments. Each mechanism still needs its own check.

ProviderSearch and training distinctionSummary mechanismCommitment or gap
GoogleGooglebot handles search; Google-Extended controls specified training and grounding usesSearch has snippet and indexing controlsCloudflare reports additional Extended URL transparency expected within weeks
AppleApplebot crawls; Applebot-Extended controls foundation-model training usenosnippet covers specified generative answersCloudflare reports URL-level inspection work for 2027
BingBingbot remains allowed under Disallow; automatic robots.txt no-training delivery is unsupportedHistorical Bing Chat controls couple some answer and training restrictionsRobots.txt no-training support targets early 2027

Google: Google-Extended is not Googlebot

Google's crawler documentation describes Google-Extended as a product token without a separate HTTP request user-agent string. It controls specified Gemini training and grounding uses. Google says it does not affect Search inclusion or ranking.

For Google AI Overviews and AI Mode, Search AI guidance points to Googlebot as the crawl-access control and lists nosnippet, data-nosnippet, max-snippet and noindex as ways to limit information shown. These controls have different consequences. In particular, noindex is not a training-only preference and should not be treated like one.

After changing preview controls, use Search Console URL Inspection to see the HTML Googlebot received, then allow time for recrawling and processing. Do not look for a separate Google-Extended crawler in access logs as proof that the preference worked, because the documentation does not describe one.

Cloudflare also describes a generative-search portal toggle. The Google documentation checked here does not independently establish that exact universal interface. This map therefore does not prescribe it or treat promised URL transparency as already shipped.

Apple: Applebot-Extended is not a search-crawler block

Apple says Applebot-Extended does not crawl pages. Disallowing it controls training use of content collected by Applebot; pages can still remain discoverable in search.

Apple separately documents nosnippet for specified generative answers, with effects on descriptions and web answers. Review those presentation effects before rollout, because they touch how content appears rather than only how it is used for training. Cloudflare attributes Apple's URL-inspection work to next year, meaning 2027, without giving a launch date.

Bing: a preference gap is not a removal recommendation

Until Bing's targeted early 2027 support arrives, Disallow AI Training does not automatically deliver a no-training preference through robots.txt. Record that state plainly as unresolved, not implemented.

Microsoft's September 2023 announcement says NOARCHIVE excludes content from Bing Chat answers and prospective foundation-model training while retaining search-result eligibility. NOCACHE takes precedence when both appear. Those are historical documented semantics, not independently verified coverage of every current Microsoft surface.

Cloudflare mentions NOARCHIVE and removal tools. Do not use removal as a training-only substitute, because indexing consequences need separate verification. This article does not prescribe removal procedures. Cloudflare's unified control over the amount of content in summaries is also a goal for early 2027, not a September feature.

Use the Crawler Intent Migration Map

Record baseline and desired outcomes

The Crawler Intent Migration Map is a working template introduced in this article, not a Cloudflare feature or a claimed client methodology. Use it to connect the recorded zone state with the outcomes you want. Alongside each proposed control, keep the provider's support status and the evidence needed to accept the change. Include the rollback path.

Export settings where possible or record exact Search, Training and Agent values, Bot Preference Sync state, the public robots.txt response and relevant WAF rules. Timestamp the record. Redact sensitive rule detail before it leaves the operational team.

Write down whether search should remain open, then specify the training preference. Give summary policy its own decision. An unsupported preference remains a gap even when the dashboard appears configured; the setting alone proves neither provider support nor compliance.

Choose representative indexable paths, including ad-serving and non-ad-serving pages where relevant. Capture available crawl and index observations before the change. Agree which unexpected denials will trigger rollback and name the person who can restore settings. Record any unrelated protections that must remain intact.

flowchart TD A[Capture zone baseline] --> B[Specify search, training and summary outcomes] B --> C{Provider supports intended combination?} C -->|No or unknown| D[Record gap or defer change] C -->|Yes| E[Record proposed controls] E --> F[Inspect robots.txt and crawler access] F --> G{Intended search access preserved?} G -->|No| H[Restore baseline and investigate] G -->|Yes| I[Observe crawl and index signals] I --> J[Retain evidence and unresolved limits]

This illustrative record is not a Cloudflare API payload or a production observation. It proposes declining training while leaving summaries undecided.

{
  "zone": "example.com",
  "observedAt": null,
  "desiredOutcomes": {"search": "allow", "training": "decline", "summaries": "undecided"},
  "currentControls": "not_checked",
  "proposedControls": {"Search": "Allow", "Training": "Disallow AI Training", "Agent": "retain_recorded_baseline"},
  "providerSupport": {"bingRobotsTraining": "target_early_2027_not_shipped"},
  "evidence": {"robotsTxt": "not_checked", "verifiedCrawlerAccess": "not_checked"},
  "rollback": "restore_recorded_baseline_after_unintended_search_denial"
}

Inspect robots.txt and crawler access

Bot Preference Sync prepends generated content to existing robots.txt material. Inspect the entire response, not only the newly generated lines. Check applicable user-agent groups and directive precedence against each crawler operator's rules. The order in which lines appear is not an acceptance test. Cloudflare also says this category-wide sync does not directly read individual custom rules with more complex logic. Do not assume a custom WAF exception has been translated into robots.txt.

Migration testEvidenceFailure conditionResponse
Baseline capturedControls, rules and robots.txtPrevious state cannot be reconstructedDefer change
Preference expressedServed robots.txt and provider tokenMissing or conflicting directiveRestore prior configuration; investigate sync
Search paths reachableVerified requests, paths and statusesNewly denied intended search crawlerRestore changed controls; identify rule
Redirects workResponse chains and destinationsUnexpected or inaccessible destinationReverse responsible change
Other protections preservedWAF events and rule evidenceUnexplained broad bypass neededStop; avoid blanket exceptions
Provider gap recordedDated source and support statusPromised support marked implementedCorrect record; defer unsupported outcome

A copied Googlebot user-agent can inspect a response, but it cannot prove access by the real Googlebot. Use documented identity verification where available. Cloudflare's crawler-management documentation describes user-agent identification on the free plan and more thorough detection on upgraded plans. Label confidence accordingly.

The same documentation warns that unsuccessful requests can come from other rules or response errors. Diagnose the responsible control before changing training policy. Otherwise, a team can weaken the wrong rule and still leave the crawler problem unsolved.

Measure outcomes and retain rollback

The AI search measurement stack gives broader measurement context. This migration needs a narrower record that ties each observation to the specific setting change.

MeasurementEvidenceInterpretation limit
Preference deliveryTimestamped robots.txt and applicable directivesPublication is not compliance
Search reachabilityVerified requests by provider, path and statusUntested paths remain unknown
Crawl and index behaviorSearch Console and Bing Webmaster observationsRevisit timing and reporting lag matter
Search and summary outcomesSearch performance and separate visibility observationsChange does not establish causality
Commercial outcomesRelevant referrals and conversionsDemand and other releases confound attribution

Google includes AI Overviews and AI Mode in Search Console's overall Web reporting. Do not label an aggregate change as an isolated summary effect. Choose observation windows around actual crawler visits and processing, not an assumed immediate verdict.

If intended search access is newly denied, restore the recorded settings and recheck affected paths. Restoration cannot guarantee immediate recrawl or index recovery. The Cloudflare Radar evidence-trace test provides related evidence-handling context for a different product, not this control rollout.

Common mistakes when separating search from training

The central mistake is selecting Training Block while expecting a mixed-use crawler to continue searching. Inspect migrated values rather than relying on the old label's meaning. Search Allow does not cancel independent restrictions elsewhere in the configuration.

Agent settings cannot solve a missing training mechanism. An ad-only Disallow option cannot solve it either, because none is offered. Accountable status does not close Bing's gap, and a training directive is not a summary opt-out.

Be careful about what the evidence actually shows. A copied user-agent string tells you little about crawler identity, while robots.txt records a published preference without establishing that training stopped. Stable rankings can also be reassuring for the wrong reason: they do not tell you whether every important path remains reachable.

Avoid universal tag recipes. Microsoft's NOARCHIVE and NOCACHE interaction shows how combined directives can change the result. Verify product scope before deployment, especially before using removal tools.

Caveats: what this map cannot prove

Publishing a preference does not prove compliance. It also cannot remove content from existing models or establish how earlier collections were used. Evidence about crawl enforcement cannot, on its own, answer questions about identity or downstream data use.

Budget for implementation effort, changing documentation, cached responses, revisit timing and incomplete logs. Keep evidence retention and access limited where privacy requires it.

No setting guarantees rankings, traffic, citations or conversions. This article is based on public-source review, not a live Cloudflare configuration change, client deployment or controlled SEO experiment. Keep unsupported outcomes visible instead of calling the migration complete for every site.

Key Takeaways

  • 1Disallow AI Training combines robots.txt preferences for Accountable mixed-use crawlers with blocking of other training crawlers.
  • 2Selecting Training Block can deny mixed-use search crawlers, but existing settings generally migrate instead of automatically adopting the new blocking behavior.
  • 3Accountable includes capabilities and commitments; assess each provider mechanism separately.
  • 4Bing's robots.txt no-training support targets early 2027, so Disallow AI Training does not automatically convey that preference to Bing today.
  • 5Treat summaries separately, verify crawler access and retain rollback evidence without claiming proof of training non-use.

Conclusion

Keep search access open deliberately, with a recorded baseline and clear rollback triggers. Training Block can deny mixed-use search crawlers; Bing's robots.txt no-training support remains targeted for early 2027. Neither a dashboard label nor a provider commitment closes that gap. Set summary policy separately and leave unsupported outcomes marked unresolved. If useful, Optijara can help scope the settings map and measurement requirements before implementation.

Frequently Asked Questions

What does Cloudflare Disallow AI Training do?

It publishes applicable robots.txt training preferences through Bot Preference Sync for Accountable mixed-use crawlers and blocks other training crawlers. Provider support gaps remain; publication is not proof that training stopped.

Can Training Block prevent Googlebot or Bingbot from reaching pages?

Yes. The new Training Block semantics include mixed-use Googlebot, Bingbot and Applebot. Ad-only blocking can affect detected ad pages. Existing settings generally migrate, so this does not mean sites automatically lost search access.

Does disallowing Google-Extended or Applebot-Extended block traditional search?

These usage-control tokens are distinct from Googlebot and Applebot. Google and Apple document training controls separately from traditional search. Neither token is a universal AI summary opt-out.

Does Disallow AI Training send Bing a no-training preference today?

Not automatically through robots.txt. Cloudflare's September 15, 2026 release targets Bing support for early 2027. Historical NOARCHIVE guidance requires current-scope review; removal tools are not training-only substitutes.

Does Accountable mean every control is already available?

No. Accountable includes current capabilities and time-bound commitments. Verify each mechanism. Cloudflare's unified control over summary amounts is a future goal, not a feature shipped in this release. Record provider gaps, check verified crawler access and define rollback triggers rather than treating the designation as proof of implementation.

Sources

Share this article

Hamza Diaz

Written by

Hamza Diaz

Hamza Diaz is the founder of Optijara, where he builds practical AI agents, automation systems, and Copilot workflows for service businesses. He writes about AI operations, agent strategy, and real-world implementation for teams that want usable systems instead of hype.