Explainer / Search and page content
Which AI crawlers should your site allow for agent search?
Choose AI crawler rules by purpose: keep agent search open, decide on training, and understand what robots.txt can control across four vendors.

If you want agents to find your public pages, keep the search crawlers open. OpenAI and Anthropic let you make a separate choice about training; Google's Gemini control combines training with grounding. Choose the use you want to permit before choosing which bot to block.
Consider a public API manual. Its owner wants a research agent to find the current documentation and send a reader to it. The owner also wants to withhold that material from model training. Those aims can coexist, but a rule that blocks every AI bot is too blunt to express them.
Can you block training while keeping AI search open?
Yes, where the vendor separates those purposes. A training crawler gathers material that may become part of a model. A search crawler helps a service find pages when answering a question. Refusing the first need not refuse the second.
OpenAI documents independent controls for GPTBot, which gathers potential training material, and OAI-SearchBot, which supplies ChatGPT search. The manual's owner can disallow GPTBot while leaving OAI-SearchBot open. Blocking the search bot excludes pages from ChatGPT search answers, although navigational links may remain.
Anthropic makes a similar split: ClaudeBot collects potential training content, and Claude-SearchBot supports search. Its policy describes a training restriction as excluding future material. A new rule is a choice about future collection, not a way to retrieve everything already collected.
That is the useful distinction for our hypothetical manual. It can refuse those training routes without closing those search routes. Whether an answer actually cites it is a further question.
Why is Google's choice different?
Google offers a control for two uses together. Google-Extended covers future Gemini training and grounding for Gemini Apps, alongside Vertex AI’s Grounding with Google Search. Grounding means supplying retrieved information to a model when it answers. Disallowing this token therefore gives up both uses in those services.
The token has no separate HTTP user-agent string. It governs uses of content collected through Google's existing crawlers, so the site's logs cannot count Google-Extended visits as though it were another bot.
Google Search is a separate decision. Google-Extended does not affect Search inclusion or ranking. AI Overviews and AI Mode use Googlebot's Search controls, with existing preview controls such as nosnippet and max-snippet governing what can be shown. Blocking Google-Extended does not opt a page out of those Search features.
For the manual's owner, that leaves a trade-off: keep Google Search available, but accept losing the named Gemini grounding uses to withhold content from Gemini training. Under that documented scope, the manual's owner makes both choices together.
What happens when a person asks an agent to open a page?
A fetch for one user's question is another route. It need not involve building a search index or gathering a training dataset, and the vendors describe different rules for it.
OpenAI's ChatGPT-User acts on user requests; its documentation says robots.txt rules may not apply. Anthropic's Claude-User also acts for users, but Anthropic says its bots honor robots.txt. Disabling that fetcher can prevent Claude from retrieving the manual for a question even if the search crawler remains allowed.
Perplexity distinguishes its search bot from its user fetcher. PerplexityBot supports search and is not used to crawl for foundation-model training. Perplexity-User visits pages for user requests and generally ignores robots.txt, according to its documentation.
The manual's owner wants this reading route open. An owner who wants to prevent access needs to look beyond a robots rule; a restriction that one fetcher follows is not a universal lock.
Which names belong to each purpose?
The table captures the four vendors' documented controls on October 9, 2026. It describes their policies, rather than measuring whether requests follow them.
| Vendor | Training control | Search control | User-request fetcher |
|---|---|---|---|
| OpenAI | GPTBot | OAI-SearchBot | ChatGPT-User; robots rules may not apply |
| Anthropic | ClaudeBot | Claude-SearchBot | Claude-User; honors robots rules |
| Google-Extended; also controls named Gemini grounding uses | Googlebot; includes Search AI features | Outside this comparison | |
| Perplexity | No training bot identified on the cited page | PerplexityBot | Perplexity-User; generally ignores robots rules |
Read across the row before copying a name into a block list. The product you want to exclude may share a control with a use you want to keep.
What does robots.txt actually enforce?
It gives cooperating crawlers instructions; it does not secure a page. RFC 9309 separates robots rules from access authorization and recommends real security measures, such as authentication, for content that must be protected. The file is public, so naming a private path there also advertises it.
For public pages, a specific rule can express a specific choice. This group requests that GPTBot stay out:
User-agent: GPTBot
Disallow: /
It says nothing about a different crawler's group. Under the standard, a crawler uses its matching group, or the wildcard group if none matches. Inspect the whole file before adding a rule: an existing default block can still close search routes you meant to leave open. Apply the intended policy to the host serving the manual, including its documentation subdomain.
What should you watch after changing the rules?
Watch the routes your policy intends to keep and close. For this manual, that means looking for training requests after the restriction and checking that search requests can still reach public documentation. This is our recommended check, not an assumption that every bot will arrive on schedule.
A name in a request is a declaration, not proof of origin. Google warns that user-agent strings can be spoofed; OpenAI and Perplexity publish bot IP ranges, and Perplexity recommends combining the name with the source IP when configuring its firewall rules. Keep traffic labelled declared, heuristic or unknown rather than treating a familiar name as verified identity.
Make the choice explicit: which training uses are refused, which search routes stay open, and whether the Google trade-off is acceptable. Then check the site's actual responses. A public manual meant for agents needs an open reading path, and an allow rule only helps the fetchers that read the file.
Frequently asked questions
Does blocking GPTBot remove a site from ChatGPT search?
OpenAI makes GPTBot and OAI-SearchBot independent controls. You can restrict training and leave the search crawler allowed; that does not guarantee a citation.
Does Google-Extended control AI Overviews?
No. AI Overviews and AI Mode use Google's Search controls, while Google-Extended governs training and grounding in the named Gemini services.
Can robots.txt keep an AI agent out of private pages?
It is not access authorization. Protect private content with authentication rather than relying on a crawler's willingness to follow robots rules.
Does allowing search crawlers guarantee inclusion in answers?
No. Google expressly says that meeting its requirements does not guarantee crawling, indexing or serving a page.
Read this article as Markdown Open in ChatGPT Open in Claude

