How To Block AI Bots (ChatGPT, Google, Anthropic & Perplexity)

If you prefer to limit or block AI bots like ChatGPT, Google’s AI, Anthropic and Perplexity from crawling your site, the easiest way to do this is to update your robots.txt file.

This guide provides step-by-step instructions to prevent these AI bots from accessing your content, helping you maintain control over your web presence and protect your data.

Blocking ChatGPT (OpenAI)

OpenAI documents three user agents relevant to publishers (also known as bots or web crawlers):

  1. GPTBot:
    • Purpose: Used for web crawling.
    • Usage: Crawls web pages to gather data that can be used to improve future AI models. It avoids sites requiring paywall access, are known to primarily aggregate personally identifiable information (PII), or have text that violates OpenAI’s policies.
  2. ChatGPT-User:
    • Purpose: Used when a ChatGPT user asks for a specific page.
    • Usage: Takes direct actions on behalf of ChatGPT users, answering live queries by accessing specific web content. It does not automatically crawl the web. OpenAI states that because a user triggers the fetch, robots.txt rules may not apply to it.
  3. OAI-SearchBot:
    • Purpose: Used to index pages for ChatGPT search.
    • Usage: Surfaces and links your pages inside ChatGPT search answers. It is not used for model training, so blocking it removes you from those answers without affecting training opt-out.

To block all three, add these lines to your robots.txt file:

User-agent: GPTBot
Disallow: /

User-agent: ChatGPT-User
Disallow: /

User-agent: OAI-SearchBot
Disallow: /

Blocking GPTBot keeps your content out of OpenAI’s training data, and blocking OAI-SearchBot removes your pages from ChatGPT search answers. The ChatGPT-User rule is a request rather than a guarantee, for the reason above.

Blocking Anthropic

Anthropic documents three crawlers. ClaudeBot collects training data, Claude-User fetches pages when a Claude user asks a question that needs the web, and Claude-SearchBot indexes pages for Claude’s search results. Anthropic says all three honour robots.txt, including the user-initiated one.

anthropic-ai and Claude-Web are older tokens Anthropic no longer documents. Listing them costs nothing, so the rules below cover all five:

User-agent: ClaudeBot
Disallow: /

User-agent: Claude-User
Disallow: /

User-agent: Claude-SearchBot
Disallow: /

User-agent: anthropic-ai
Disallow: /

User-agent: Claude-Web
Disallow: /

Blocking Google’s AI Bot

To stop Google using your content to train Gemini (formerly known as Bard), add this line to your robots.txt file:

User-agent: Google-Extended
Disallow: /

Google-Extended is a usage control rather than a crawler. Google states that it does not affect your inclusion in Google Search and is not a ranking signal, so blocking it costs you nothing in search.

Blocking Perplexity AI

Perplexity runs two agents. PerplexityBot indexes pages so they can be linked in Perplexity search results, and Perplexity-User fetches pages when someone asks a question. Perplexity states that Perplexity-User generally ignores robots.txt because a user triggered it.

User-agent: PerplexityBot
Disallow: /

User-agent: Perplexity-User
Disallow: /

Important Notes

  • These companies say they respect robots.txt, and at least one has been caught working around it. In August 2025 Cloudflare reported that Perplexity was fetching pages from sites that had blocked its declared bots, using an undeclared crawler with a generic Chrome user agent across rotating IP addresses, at 3 to 6 million requests a day. Cloudflare removed Perplexity from its verified bot list. Perplexity denies the claim.
  • These rules may not block plugins, extensions or agents built on top of these models that fetch URLs. Those tools send their own user agent, which the rules above do not cover.
  • robots.txt is an instruction, not access control. If you need to guarantee that a page stays out of an AI system, put it behind authentication.