Cloudflare splits AI bot controls into search, agents, and training
Introduction
Cloudflare has announced a major redesign of its AI crawler control system. Instead of a single switch that blocks all AI bots, customers will now see three independent categories: Search, Agent, and Training. According to the available material, the controls are available to all Cloudflare customers, including those on the free plan.
This is more than a small product settings change. The term “AI crawler” now covers several very different behaviors: search systems indexing content, AI agents browsing on behalf of users, and crawlers collecting data for model training. Treating all of them as one group is increasingly too blunt for publishers, developers, and site operators.
Key points
-
Search: indexing and answer retrieval
Search-oriented crawlers are meant to index pages and support responses when users ask questions. For many websites, this can create visibility, but it can also raise concerns when answers are surfaced elsewhere without sending users back to the original source. Separating this category allows site owners to decide whether AI-powered search access is acceptable. -
Agent: task-driven access
Agent traffic is different from traditional crawling. It is closer to an AI assistant visiting pages in response to a user request, such as reading, comparing, or helping complete a task. Because this traffic may be tied to immediate user intent, some websites may want to treat it differently from bulk data collection. -
Training: data collection for models
Training crawlers are associated with collecting content to train or improve AI models. This is often the most sensitive category, because once material becomes part of a training pipeline, the original site may have limited visibility into downstream use. Making Training a standalone option gives publishers a clearer way to express their preference.
Why it matters
The old one-click approach was simple, but it forced a difficult binary choice: allow AI bots or block them all. That made little sense for websites that might welcome search indexing but reject training use, or that might tolerate agent-based access while limiting broader scraping.
Cloudflare’s three-part model creates a middle ground. A site can set different rules for discovery, user-directed automation, and model-building. This is especially relevant for media sites, community platforms, blogs, and knowledge bases, where the central question is not whether all AI access is bad, but how content is being used and whether value flows back to the origin.
For AI companies, the change also increases pressure to describe crawler behavior more clearly. A generic “AI bot” label will become less useful as infrastructure providers and site owners ask: Is this crawler powering search? Is it acting on behalf of a user? Is it collecting data for training?
Cloudflare’s update does not settle the broader dispute over AI and web content. But it offers a practical framework for making the dispute more manageable. By separating AI access into Search, Agent, and Training, the company is giving website owners a more precise language for control—and possibly setting the stage for new norms around how AI systems interact with the open web.
Source: OSChina
Comments
Checking sign-in status...
Loading comments...