Training crawlers are now half of crawler traffic. Should you block them?
Cloudflare says training crawlers made 52% of crawler requests in June 2026. Blocking them is a business choice; blocking search and user bots costs answers.

The short answer
Block training crawlers only if keeping your content out of future model training matters more than how models describe you. Never block the search crawlers and user-triggered fetchers that AI assistants use to answer questions, because that removes you from live answers. In Cloudflare's June 2026 data, training crawlers made 52% of crawler requests.
Key takeaways
- On Cloudflare's network, crawlers for AI training made 52% of crawler requests in June 2026, up from 22% in spring 2025.
- Training crawlers, search crawlers and user-triggered fetchers do different jobs, and AI companies such as OpenAI let you control them separately.
- Blocking a search crawler or user fetcher removes you from that assistant's live answers; blocking a training-only crawler such as GPTBot does not.
- Google's AI features in Search are governed by Googlebot, so a site cannot leave AI Overviews through robots.txt without leaving Search.
- Our view: for most companies that sell to buyers who use AI assistants, allowing all three types is the better default.
In this article
What Cloudflare reported
Cloudflare published a review on 1 July 2026, a year after it began blocking AI training crawlers by default on new domains. On its network, crawlers for AI training made 52% of crawler requests in June 2026, up from 22% in spring 2025.1 Mixed-use crawlers, which blend search, agent and training use, made up more than 36% of activity, and pure search crawling was a small and declining share.1
Two other figures frame the decision. Cloudflare says more than half of web traffic is now non-human, and that Google still accounts for about 88% of referral traffic.1 Crawling is growing fast; the visits that come back are still mostly from one search engine.
Read the source with its interest in mind. Cloudflare sells bot management and tools that let publishers charge AI companies for access, and it argues for blocking by default. The figures come from Cloudflare Radar and its own network, which carries a large share of the web but not all of it.
Share of crawler requests made for AI training
Three kinds of bot, three different jobs
The word "AI crawler" covers bots that do very different things, and the decision depends on which one you mean.
- Training crawlers collect pages that may be used to train future models. OpenAI's GPTBot is one; disallowing it tells OpenAI not to use your content for training.2
- Search crawlers index pages so an assistant can find and cite them in live answers. OpenAI says sites that opt out of OAI-SearchBot will not be shown in ChatGPT search answers.2
- User-triggered fetchers visit a page because a person asked the assistant to. OpenAI's ChatGPT-User is one, and OpenAI notes that robots.txt rules may not apply to such user-initiated requests.2
OpenAI states that each setting is independent, so a site can allow search and block training.2 Other AI companies publish similar splits in their own documentation, and the names change, so check each one before you edit robots.txt.
Google works differently
Google's AI features in Search are not controlled by a separate AI crawler. Google's guidance, last updated in December 2025, says robots.txt directives for Googlebot are the control for how a site is crawled for Search, AI features included.3 To limit what is shown, it points to snippet controls such as nosnippet and max-snippet, and it names Google-Extended as the control for training and grounding in some of Google's other systems.3
In practice, blocking Google-Extended does not remove you from AI Overviews or AI Mode, and blocking Googlebot removes you from Search entirely.3 Cloudflare's report criticises this combined approach.1 Whatever one thinks of it, it is the rule site owners work under today.
A robots.txt decision table
The table sets out what blocking each type costs in AI answers. It is our reading of the vendors' documentation, not a guarantee of how any engine behaves.
| Bot type | Example | If you block it | Our default |
|---|---|---|---|
| Training crawler | GPTBot | Content kept out of future training; ChatGPT search answers unaffected | Allow, unless content is the product |
| Google training and grounding control | Google-Extended | Out of training and grounding elsewhere at Google; Search unaffected | Allow, unless content is the product |
| AI search crawler | OAI-SearchBot | Dropped from that assistant's search answers | Allow |
| User-triggered fetcher | ChatGPT-User | Assistant may fail to read a page a user asked about | Allow |
| Search engine crawler | Googlebot | Out of Search, AI Overviews and AI Mode together | Allow; use snippet controls |
Sources: OpenAI crawler documentation; Google Search Central. Our defaults for companies that sell to buyers who use AI assistants.
Who should block training crawlers?
Publishers whose content is the product have the strongest case. A news site, a research firm or a database business sells the text itself, and training use without payment competes with that. For them, Cloudflare's default and its licensing tools make commercial sense.
Most companies are in a different position. A software vendor, a manufacturer or a clinic wants models to describe it accurately, and its pages are marketing, not inventory. Blocking a training-only crawler such as GPTBot will not remove it from ChatGPT's search answers, but it may leave a model's built-in knowledge of the company to older or third-party sources.
Many sites have the emphasis backwards. BrightEdge, which sells SEO software, reported in April 2026 that most companies it studied focused their bot policies on blocking training agents, while far fewer set rules for the search and user-facing agents that affect live answers.4 We have not verified BrightEdge's figures, so we describe its finding without them.
What to do this week
- Read your robots.txt and your CDN or firewall bot settings together; a block at the network level overrides an allow in the file.
- List every AI user agent you allow or block, and mark each as training, search or user-triggered.
- Allow search crawlers and user fetchers unless you have a specific reason not to.
- Decide on training crawlers as a business question, with whoever owns your content licensing.
- Recheck after changes, because new user agents appear and defaults change.
Our view: we allow search, user and training crawlers on sigzenai.com, because we want engines to describe us accurately. Crawler access is the cheapest item in AI visibility work, and as our review of schema, llms.txt and other citation tactics found, the hygiene items matter mainly when they are broken. A blocked search crawler is broken in the most expensive way. For how Google picks the pages it cites once it can read them, see our explainer on rankings and AI Overview citations.


