Guide

The AI Crawler Guide: Configuring robots.txt for GPTBot, ClaudeBot & More

Which AI bots visit your site, what each one does, and how to allow AI search visibility while keeping control over model training.

A robots file connected to AI crawlers from OpenAI, Anthropic, Google Gemini and other bots

Key takeaways

  • AI companies run separate bots for model training, search indexing and user-triggered fetches.
  • You can allow search bots for visibility while blocking training bots if you prefer.
  • Google-Extended controls Gemini training use; it does not affect Google Search or AI Overviews.

Most AI providers now publish the user agents their crawlers use, and many separate training crawlers from search crawlers. That distinction matters: you can stay visible in AI answers without necessarily contributing your content to model training.

The main AI user agents

  • OpenAI: GPTBot (training), OAI-SearchBot (ChatGPT search results), ChatGPT-User (pages fetched on a user’s request).
  • Anthropic: ClaudeBot (training), Claude-SearchBot (search), Claude-User (user requests).
  • Perplexity: PerplexityBot (search index), Perplexity-User (user requests).
  • Google: Googlebot crawls for Search, including AI Overviews. Google-Extended is a control token for Gemini model use, not a separate crawler.

Example: allow AI search, block AI training

robots.txt
# AI search and user-triggered fetches: allowed
User-agent: OAI-SearchBot
User-agent: ChatGPT-User
User-agent: Claude-SearchBot
User-agent: Claude-User
User-agent: PerplexityBot
User-agent: Perplexity-User
Allow: /

# Model training: blocked
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: Google-Extended
Disallow: /

# Everyone else
User-agent: *
Allow: /
Disallow: /admin/

Sitemap: https://yourdomain.com/sitemap.xml

If you’re comfortable with your public content being used for training, and many businesses are, since it can reinforce how models describe them, simply allow all of the bots above.

Don’t forget your CDN and firewall

Some CDNs and security services block AI bots by default or through “bot fight” settings. A perfect robots.txt doesn’t help if requests are rejected before they reach your server. Check these settings alongside your robots file.

Verify it’s working

  1. Fetch /robots.txt in a browser and confirm the rules are live.
  2. Look for AI user agents in your server or CDN logs over the following weeks.
  3. Ask assistants questions your pages answer and see whether they cite you.

Ready to be the answer AI recommends?

Get your free AI readiness audit, or talk to our team about a full GEO & AEO strategy for your brand.