Search or AI Training: Cloudflare Makes Crawlers Accountable

Cloudflare's Accountable designation and Disallow AI Training split search indexing from AI training for the first time. Netics looks at the mechanism, the four criteria, and what publishers

Netics feature card for the Cloudflare search-or-AI-training article with the official Cloudflare logo
Netics editorial feature card using the official Cloudflare identity

TL;DR

  • Cloudflare launched "Disallow AI Training" and an "Accountable" designation for AI crawlers, splitting search indexing from AI training for website owners (September 15, 2026).
  • Mixed-use crawlers are 36.6% of verified crawler traffic on Cloudflare's network — the numbers that drive the tradeoff: fewer than 1% of sites block search, 17% restrict training.
  • Apple, Google and Microsoft are labeled Accountable; the four criteria cover training opt-out, AI-summary opt-out, URL-level visibility, and a no-ranking-penalty commitment.
  • Netics' take: this is the first time crawler access is treated as a contract with auditable terms — and that changes what "robots.txt" means for publishers.
  • Practical move: set the three controls explicitly, check your own crawler traffic, and treat AI summaries as part of your content strategy, not a side effect.
Official Cloudflare press-release card for the search-or-AI-training announcement
Official Cloudflare press-release image for "Cloudflare Helps End the Search-or-AI-Training Tradeoff"; source: https://www.cloudflare.com/press/press-releases/2026/cloudflare-helps-end-the-search-or-ai-training-tradeoff/

The thirty-year-old crawler deal just got new terms

On September 15, Cloudflare announced two things: an Accountable designation for AI crawling, and a "Disallow AI Training" setting that lets any website refuse AI training while staying fully present in search results. For nearly thirty years the deal between crawlers and website owners was simple, in Cloudflare's own framing: they crawl you, and they send traffic back. AI changed the terms, because one crawler can now collect the same page for two purposes — a search index and a training corpus — with no way for the site owner to separate them.

The numbers make the conflict concrete. Mixed-use crawlers, which collect content for both search and training, now make up 36.6% of verified crawler traffic on Cloudflare's network — the single largest category. Most sites still want to be found: fewer than 1% block search crawlers. But 17% already restrict AI training. Until now, refusing training often meant losing search visibility from crawlers that had no way to honor the split. That is the tradeoff Cloudflare is claiming to end.

Three toggles, not one blunt switch

The mechanism is the interesting part. Cloudflare is replacing the single "Block AI Bots" switch with three independent controls: one for search, one for AI training, one for AI agents. That split matters more than it looks. A content owner's answers to "may you index me", "may you train on me", and "may your agents operate on my behalf" are different decisions with different business consequences. Treating them as one switch forced publishers to choose between exposure and control.

Official Cloudflare blog hero: have it both ways
Official Cloudflare blog image for "Have it both ways: stay discoverable in search while disallowing AI training"; source: https://blog.cloudflare.com/accountable-mixed-use-ai-crawlers

Cloudflare also replaces "Managed Robots.txt" with "Bot Preference Sync": set your preferences once, and they apply across all supported crawlers automatically. We already argued that Bot Preference Sync defines a governance boundary — this announcement extends that boundary from Cloudflare's own products to the crawler market as a whole. Without a sync layer, a site owner would have to hand-edit robots.txt per crawler, which is exactly how the old convention fell apart.

Accountable is the word to argue about

The novel move is not the settings. It is the "Accountable" designation: Cloudflare now publishes criteria that crawler operators must meet to keep access, and names the operators that satisfy them. There are four requirements: a clear way to opt out of AI training via robots.txt or a comparable standard; a way to opt out of AI-generated search summaries; URL-level visibility into how your content is used for search and training; and a public confirmation that opting out of training will not affect ranking in traditional search.

Netics diagram: claim versus check on the four accountability criteria
Original Netics diagram: the four Accountable criteria as claim versus check — training opt-out, summary opt-out, URL-level visibility, and the no-ranking-penalty commitment; source: Cloudflare press release, September 15, 2026

Apple, Google and Microsoft are now labeled Accountable. Other mixed-use crawlers are blocked if site owners Block Training. Note the power shift hiding inside that sentence: a CDN that sits between websites and crawlers is now rating crawlers, and deciding which of them get to read the web. Whether that is a sensible referee or an uncomfortable concentration of power depends on your view of Cloudflare — but either way, it is a new form of gatekeeping that did not exist three months ago.

Cloudflare's own framing is worth reading carefully. Matthew Prince is quoted in the release: "This is how we make the Internet better: preserving the openness that makes search valuable while giving the people and businesses behind the web meaningful control over how their work is used." The first half is about openness; the second half is about control. Those pull in different directions, and the criteria are the place where the tension gets resolved.

What the numbers say about who blocks what

The useful stat is the gap: 36.6% of verified crawler traffic is mixed-use, but only 17% of sites restrict training. That gap says most site owners never noticed the change in the crawler deal, or never had a mechanism to act on it. The three independent controls fix the mechanism for Cloudflare customers, and the Accountable criteria give those customers a reason to trust the split. Google-Extended already lets sites opt out of training without leaving Search; other companies run separate crawlers for search and training, which lets Cloudflare block the training crawler without touching search. The designation is a way of making those existing behaviors legible and comparable.

Netics diagram: the crawler traffic numbers that drive the tradeoff
Original Netics diagram: 36.6% mixed-use crawler traffic, under 1% blocking search, 17% restricting training — and the three new toggles; source: Cloudflare press release, September 15, 2026

What a publisher should do this month

Take the recommended settings literally, because they encode an economics lesson. New ad-supported sites get: search crawling on, AI training disallowed, and agents blocked on pages that carry ads. Cloudflare's reasoning is transparent: ad revenue depends on a human actually seeing the page; training replaces that visit with an answer; agents fetch the page with nobody there to see the ads. Every other site gets all three allowed, matching how most non-ad-supported websites already operate.

For a publisher, the actionable reading is: know which of the three you are saying yes to. An ad-supported French media site should almost certainly disallow training and decide deliberately whether agents may read its pages. A documentation-driven SaaS business might reasonably allow everything, because its content is a lead-generation asset and AI answers can reference it. The one setting nobody should keep is "default" — that is how 36.6% of your crawler traffic ends up with two licenses it never signed.

Official Cloudflare settings table: recommended defaults for new domains
Official Cloudflare blog figure showing the recommended settings for new domains (search, training, agent toggles for ad and non-ad sites); source: https://blog.cloudflare.com/accountable-mixed-use-ai-crawlers

AI summaries are the next frontier in the same negotiation: every Accountable operator must give site owners a way to opt out, and Cloudflare's stated goal is site-level control over how much content is included, "in one place", by early next year. This is also where the IETF "ai-prefs" work becomes relevant — a standardized, portable way to express AI access preferences, which would turn today's Cloudflare-specific settings into a web-wide convention.

For a practical review of your site's crawler policy — what your robots.txt actually says, who is reading your content, and what the three toggles should be — start from the Netics homepage. The fix is cheap; the default is not.

Sources

Source: "Cloudflare Helps End the Search-or-AI-Training Tradeoff" — cloudflare.com, September 15, 2026.