If an AI search engine's crawler cannot reach your page, that engine cannot cite it. Most sites do not block AI crawlers on purpose. They block them by accident, through an old security rule, a copied robots.txt template, or a bot-management default nobody reviewed. This guide explains the main AI user agents and gives a default you can adapt.

Training bots, search bots, and user bots

AI companies now run several crawlers each, and they do different jobs. The distinction matters because you can allow one and block another.

  • Training crawlers collect pages to train future models. Blocking them keeps your content out of training data but does not, by itself, stop you from being cited in live answers.
  • Search crawlers build the index an answer engine searches when someone asks a question. Blocking these is what removes you from AI answers.
  • User-triggered fetchers visit a page when a person asks the assistant to read it or follow a link. They act on behalf of a user in real time.

The main AI user agents

OpenAI (ChatGPT)

  • GPTBot: training.
  • OAI-SearchBot: ChatGPT search results.
  • ChatGPT-User: fetches pages when a user's request requires it.

Anthropic (Claude)

  • ClaudeBot: training.
  • Claude-SearchBot: search results inside Claude.
  • Claude-User: fetches pages for a user's request.

Perplexity

  • PerplexityBot: Perplexity's search index.
  • Perplexity-User: fetches pages for a user's request.

Google

  • Googlebot: Google Search, including AI Overviews and AI Mode. Google's AI features in Search are built on its normal search index, so blocking Googlebot removes you from those too.
  • Google-Extended: a control token, not a separate crawler. It tells Google whether your content may be used for Gemini model training and grounding. Google states it does not affect inclusion or ranking in Google Search.

Microsoft, Apple, and others

  • Bingbot: Bing search, which also supplies Microsoft Copilot.
  • Applebot-Extended: controls use of your content for Apple's AI training. Regular Applebot powers Siri and Spotlight.
  • CCBot: Common Crawl, a public dataset widely used to train models.

A sensible default for publishers who want to be cited

If your goal is visibility in AI answers, allow the search and user agents. Whether to allow training crawlers is a business decision; many publishers allow search bots and block training bots.

# Allow AI search and user-triggered fetchers
User-agent: OAI-SearchBot
User-agent: ChatGPT-User
User-agent: Claude-SearchBot
User-agent: Claude-User
User-agent: PerplexityBot
User-agent: Perplexity-User
Allow: /

# Optional: opt out of model training
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: CCBot
Disallow: /

User-agent: *
Allow: /

Sitemap: https://example.com/sitemap.xml

Two technical notes. First, a crawler follows the most specific group that names it and ignores the * group, so repeat any private paths, such as an admin area, inside each named group. Second, robots.txt is a request, not a lock. Reputable crawlers honor it; to enforce rules, use your CDN or firewall.

Check your CDN and firewall, too

Robots.txt is only half the picture. Many CDNs and security tools now offer one-click AI bot blocking. If it is switched on, AI crawlers can be refused before they ever read your robots.txt. Review those settings alongside the file.

A five-minute audit

  1. Open yourdomain.com/robots.txt and search for each user agent above.
  2. Look for Disallow: / under User-agent: * with no exceptions. That blocks everything.
  3. Check your CDN or security dashboard for AI bot or "AI scraper" blocking.
  4. Review server logs for the user agents above to confirm they are getting 200 responses, not 403.
  5. Write the decision down, so the next site migration does not undo it.

AEO Sherpa allows all of the AI search, user, and training agents listed here. You can see our own file at aeosherpa.com/robots.txt.

Free newsletter

The Sherpa Brief

One email a week on what changed in AI search, what it means, and what to do about it. Written for marketers, SEOs, and publishers.

Free. Unsubscribe in one click.