SEO, GEO & Ecommerce Tools

AI Bot robots.txt Policy Builder

Choose which AI training, search and user-triggered crawlers may access which paths, see what each choice does and does not change, and export the robots.txt section.

  • robots.txt section
  • Consequence summary
  • Merged file
Runs in your browser

Everything you paste, type or drop is processed in this browser tab. It is not uploaded, logged, stored or sent to analytics.

AI bot policy workspace

1 Choose a policy per bot

Start from:

Loading the crawler catalogue...

One per line, starting with /. * matches anything and $ anchors the end, as RFC 9309 defines.

2 Your current robots.txt (optional)

Paste it to get a merged file: groups for the bots you manage are replaced, everything else is kept as written. Nothing is uploaded.

Paths to test the result against

3 robots.txt and what it changes

Pick a starting policy, then adjust each bot.

What the AI Bot robots.txt Policy Builder does

This builder writes the robots.txt rules for AI crawlers - GPTBot, ClaudeBot, Google-Extended, PerplexityBot, CCBot and the rest - from a choice you make per bot, merges them into your existing robots.txt without disturbing anything else, and tells you in plain words what each choice changes and what it does not.

Every user agent in the list comes from its operator's own documentation, with the page linked and the date it was checked. The merged file is then tested with the matching rules of RFC 9309, the robots.txt standard, so you can see which rule decides each bot's access before you publish it.

How to use it

  1. Start from a preset - block training while allowing search, block all AI bots, or allow them all - and adjust individual bots. Googlebot is never changed by a preset.
  2. For any bot set to Block paths, list the paths it should stay out of, such as /account/ or /drafts/.
  3. Paste your current robots.txt. Groups for the bots you manage are replaced; your other groups, sitemaps and comments are kept exactly as written.
  4. Read what each choice does and does not change. Pay attention to user-triggered fetchers, which their operators say may not honour robots.txt.
  5. Check the test table: every managed bot against every test path, with the rule that decided it. Then copy or download the merged file and publish it at the root of each host.

Reading the results

Training crawlers collect content that may be used to train models (GPTBot, ClaudeBot, CCBot, Meta-ExternalAgent, Amazonbot). Search crawlers index pages so an AI search product can show and link them (OAI-SearchBot, Claude-SearchBot, PerplexityBot). User-triggered fetchers visit a page because someone asked an assistant to. Control tokens such as Google-Extended and Applebot-Extended never crawl; they only say how content fetched by the main crawler may be used.

The generated section combines bots with identical rules into one group, which RFC 9309 allows. Each bot then follows only its own group and ignores the "User-agent: *" group, so a bot you allow here is allowed even where your general rules say otherwise.

The test table uses RFC 9309 matching: the group naming the bot (case-insensitive) or, failing that, the * group; within it, the longest matching path wins and an equally long Allow beats Disallow. If a row says "As intended: No", another rule in your file is overriding the policy.

Worked example: an online shop that blocks training but stays in AI search

The shop's existing robots.txt blocks /cart/ and /search? for everyone, blocks GPTBot completely, and lists a sitemap. Choosing "Block training, allow search" sets 18 of the 19 documented agents: the five training crawlers and the two control tokens to Block all, and the search crawlers and user-triggered fetchers to Allow. Googlebot is left out.

The merge removes the old GPTBot group (its single Disallow: / is now in the new block), keeps the * group and the sitemap line untouched, and appends one block with two groups: seven agents under Disallow: / and eleven under Allow: /.

The test table shows GPTBot disallowed on /, /blog/robots-guide and /account/settings, OAI-SearchBot allowed on all three - including paths the * group would block, because it now has its own group - and every row as intended. The consequence list warns that Perplexity-User, ChatGPT-User, Meta-ExternalFetcher and Amzn-User may not honour the file at all.

Google-Extended does not keep you out of AI Overviews

Google documents Google-Extended as a control over whether content its crawlers fetch may be used to train future Gemini models and for grounding in Gemini Apps, and states that it does not affect inclusion or ranking in Google Search. AI Overviews and AI Mode are part of Search and are built on what Googlebot crawls, so blocking Google-Extended leaves them unchanged.

The controls Google documents for AI features in Search are page-level: nosnippet, data-nosnippet, max-snippet and noindex. Blocking Googlebot itself would take the paths out of Google Search altogether, which is why this builder marks it as risky and no preset touches it.

What robots.txt can and cannot do

RFC 9309 describes robots.txt as a set of rules crawlers are requested to honour; it is not access control. Operators such as OpenAI, Anthropic, Perplexity, Meta and Amazon document their background crawlers as respecting it, but several document that fetches made on a user's behalf may not. And no line in robots.txt reaches back in time: blocking a crawler stops future collection, not the use of what was already gathered.

Limitations: what the result does not prove

  • The catalogue lists only user agents with an official operator page, checked on the date shown. New agents appear and purposes change; re-check the linked page before relying on an entry.
  • It cannot verify that a request claiming to be a given bot really is. Operators publish IP ranges or reverse-DNS checks for that, which need server-side verification.
  • Each host (and each subdomain) needs its own robots.txt at its root. A file on www.example.com does not cover shop.example.com.
  • Blocking does not remove content from existing datasets, model training or search indexes, and user-triggered fetchers may ignore the rules entirely.

Privacy: where your data goes

Everything you paste, type or drop is processed in this browser tab. It is not uploaded, logged, stored or sent to analytics. Session recording and tag-manager scripts are switched off on this page.

Standards and sources

Frequently asked questions

How do I block GPTBot but keep appearing in ChatGPT search?

Block GPTBot and allow OAI-SearchBot. OpenAI documents them separately: GPTBot crawls content that may be used for training, while OAI-SearchBot surfaces sites in ChatGPT's search features. The "Block training, allow search" preset sets exactly this for every operator that separates the two.

Does blocking ClaudeBot stop Claude from reading my pages?

Only for training. Anthropic documents three agents: ClaudeBot for training data, Claude-SearchBot for search indexing and Claude-User for fetches a user asks for. Blocking ClaudeBot leaves the other two unaffected unless you block them as well.

Will blocking AI bots remove my content from models already trained?

No. robots.txt only affects future crawling. Content collected before the change, including copies in public datasets such as Common Crawl snapshots, is not withdrawn by editing the file.

Why are some bots marked as possibly not honouring robots.txt?

OpenAI, Perplexity, Meta and Amazon each document that their user-triggered fetchers may not follow robots.txt, because the visit is made on behalf of a person. A rule for them records your preference but should not be treated as an enforced block.

What is the difference between Google-Extended and Applebot-Extended?

Both are control tokens rather than crawlers. Google-Extended controls Gemini training and grounding use of content Google crawls; Applebot-Extended controls whether Apple uses Applebot-crawled data to train its foundation models. Neither affects the company's search results.

Why does an allowed bot ignore my User-agent: * rules?

Under RFC 9309 a crawler follows the group that names it and only falls back to the * group when no group names it. Once a bot has its own group, your general Disallow lines no longer apply to it, so repeat any paths it should still avoid.

Where does the merged robots.txt have to go?

At the root of each host, served as plain text at /robots.txt - for example https://example.com/robots.txt. RFC 9309 asks crawlers to read at least the first 500 KiB, so keep the file below that size.

Last reviewed by the A2Z.Tools team against the sources listed above.

Rate this tool

Was this tool useful? Your feedback helps us improve it.

No ratings yet — be the first to rate this tool.
Your rating (required)
0 / 2000

Please do not include passwords, payment details or other sensitive information.

Your feedback is sent privately to the A2Z.Tools team and will not be posted publicly.