Search Infrastructure

Cloudflare Separates Training Disallow From Blocking Mixed-Use Crawlers

By Kaleido Field Staff ยท September 17, 2026

A training preference and an access block do different jobs

Cloudflare's September 15 update distinguishes a training-disallow signal from blocking mixed-use crawlers. For its accountable crawler category, disallowing training can preserve search access; blocking can stop the same crawler from serving search. Operator support is not uniform, so the setting and the crawler's actual behavior need separate checks.

Citation-ready: Cloudflare's mixed-use crawler controls distinguish no-training instructions from network blocking, which can also interrupt search crawling.

Evidence boundary: Vendor documentation of controls and operator commitments. No crawler-compliance experiment, search-traffic effect or account configuration change was performed.

Cloudflare's official domain setup interface with separate search, training and agent settings
Image source: Cloudflare; official configuration screenshot, not a change to this site's settings. Used for editorial coverage of crawler access desk.

What happened and why it matters

A single crawler can have multiple downstream uses, so a publisher's intent cannot be accurately represented by treating every refusal as the same technical action.

Original source

Primary reference: Cloudflare September 15 mixed-use crawler announcement. Kaleido Field checked the event date and the article's attributed facts against this source.

Source check
Source dateSeptember 15, 2026
Checked by Kaleido FieldSeptember 17, 2026, CST
Source functionsearch infrastructure -> crawler purpose and discovery boundaries

The migration is intended to preserve practical behavior

Cloudflare says existing granular training-block settings move to the new disallow behavior for accountable mixed-use crawlers. A full block can include Applebot, Bingbot and Googlebot, with consequences for search access.

That distinction is operationally important: the label on a control is not enough to predict what requests it will reject. A configuration review should identify the actual crawler and the rule applied to it.

Commitment is not the same as present support

The post says Bing's robots-based no-training support is expected in early 2027; current controls include noarchive and webmaster settings. The accountable category therefore includes future commitments, not just behavior implemented uniformly today.

A publisher evaluating the change should record both dates: when it expressed a preference and when an operator supports the relevant mechanism. This avoids turning a announced transition into a completed enforcement claim.

Discovery and permission require separate evidence

Search visibility, training permissions and agent access are different questions even when they share network infrastructure. An allowed crawl does not prove that a page was indexed, and a stated restriction does not prove downstream compliance.

The answer-citation evidence map likewise separates discoverable content from an observed answer citation. This report adds the access-control layer; it does not report a gain or loss in Kaleido Field traffic.

Evidence boundary

Vendor documentation of controls and operator commitments. No crawler-compliance experiment, search-traffic effect or account configuration change was performed.

Reader briefing

Keep the source trail in view.

One concise email when a model, benchmark, or visual-intelligence claim materially changes.

FAQ

Does disallowing training necessarily block Google Search?

Not under the described accountable mixed-use disallow path; a full crawler block can affect search.