Back to articles
Industry News

Cloudflare reframes AI crawlers as Search, Agents, and Training traffic

4 min read

Introduction

Cloudflare is changing how website owners manage AI-related automated traffic. A year ago, the company’s message was centered on blocking AI bots and pushing back against uncompensated model training. Now the problem has become more complicated: search engines are turning into answer engines, AI assistants can browse on behalf of users, and the same crawler can serve several purposes at once.

Instead of asking whether a bot is simply “AI” or not, Cloudflare wants site owners to ask a more practical set of questions: What is this system doing on my site? What is it storing? How might it reuse or redistribute my content?

The new taxonomy

Cloudflare’s updated framework divides major AI-related traffic into three categories:

  • Search: systems that collect or index content so they can answer questions about it later. Cloudflare frames this as the category most naturally linked to referral traffic or other fair compensation.
  • Agent: automated behavior acting in real time on behalf of a person, such as chat fetch bots or browser-use agents that visit a website to complete a task.
  • Training: crawlers that take content to train or fine-tune models, where the data becomes part of the model’s underlying capabilities.

This distinction matters because not all automation creates the same risk or value exchange. A search index may help users discover a site. A training crawler may absorb the content without sending readers back. An agent may perform an action for a user while bypassing the original page experience.

More granular controls

Cloudflare says customers will be able to manage these three AI traffic types separately, including users on the free tier. This replaces the older, broader approach of simply blocking AI bots associated with model training.

The company is also pushing for more transparency from bot operators. If one company uses automation for search indexing, agentic browsing, and training, Cloudflare argues that those purposes should be represented separately. Multi-purpose crawlers should be tracked by all of their uses, not just one convenient label.

New defaults in 2026

Beginning September 15, 2026, Cloudflare plans to apply new defaults for newly onboarded domains. On pages that display ads, Training and Agent traffic will be blocked by default, while Search will remain allowed.

The reasoning is that an ad-supported page signals an expectation of human attention. If a bot trains on the content or an agent retrieves the answer without bringing the person to the page, the site owner may lose the monetizable visit. Search, by contrast, is treated as the automation category most likely to send visitors back.

Cloudflare also says multi-purpose crawlers that combine Search and Training will be governed by all applicable behaviors. If a customer blocks Training, then crawlers with both search and training purposes may be blocked under the most restrictive rule. The source specifically names Googlebot, Applebot, and BingBot as examples of multi-purpose crawlers affected by this logic. Site owners will still be able to opt out of the default configuration.

Why it matters

The broader story is that the old web bargain—crawl my pages and send me traffic—is under pressure. AI answer engines and agents can extract value from content while reducing the need for users to visit the original site. For small publishers, the dilemma is especially sharp: blocking automation may reduce discovery, but allowing everything may mean giving away content without return.

Cloudflare’s approach does not settle the compensation debate, but it does introduce a more precise control layer. It gives publishers a way to distinguish between indexing, real-time user delegation, and model training.

For enterprise customers, Cloudflare is also launching BotBase, a searchable database of known bots and agents inside the dashboard. Initially it focuses on visibility, with plans to expand into a control center for automated traffic.

The direction is clear: AI traffic governance is moving from blunt blocking toward purpose-based access rules. The next web crawling debate may be less about whether bots are allowed and more about what role they claim to play.

Source: Hacker News

Comments

Checking sign-in status...

Loading comments...

Related articles