
Cloudflare Splits AI Crawlers Into Three Switches — and Blocks Two by Default
Website owners can now separately control Search, Agent and Training bots, with agent and training crawlers blocked by default on ad-supported pages from September 15.
Cloudflare has overhauled how website owners manage AI traffic, replacing its blunt AI-bot blocker with three separate controls — Search, Agent and Training — and setting new defaults that push AI companies toward paying for content.
The three categories. Search covers traditional indexing, Googlebot-style crawls that feed search results. Agent covers bots fetching content live on a user's behalf, like a Claude or ChatGPT session pulling a page mid-conversation. Training covers bulk collection to build or fine-tune models, with no immediate user-facing task attached.
The new defaults. Starting September 15, AI training bots and AI agent crawlers will be blocked by default on ad-supported pages across Cloudflare's network unless a site owner explicitly re-enables them. Search crawlers stay allowed. The controls are available even to free-tier customers.
The catch. Cloudflare's data shows 36% of crawler activity now comes from mixed-use bots that blend search and training in a single crawler — a design that complicates blocking, since cutting off training may also cut off search indexing. That ambiguity is exactly what major search engines have relied on.
Why it matters. The move is the sharpest lever yet in publishers' fight to be paid for the content AI models consume. By classifying crawlers by purpose and opening the door to charging them directly, Cloudflare — which sits in front of a large share of the web — is trying to build the toll booth that the AI content economy has so far lacked.
Newsletter
Get Lanceum in your inbox
Weekly insights on AI and technology in Asia.


