Log in
Sign up with Google

AI training crawler · OpenAI

GPTBot is OpenAI’s AI training crawler

GPTBot is OpenAI's web crawler that collects content that may be used to train its generative AI foundation models.

Verifiable operator Respects robots.txt

Our take

Your call

It only collects training data, so allowing it is a content-licensing decision, not a visibility one.

User agent

Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; GPTBot/1.4; +https://openai.com/gptbot

01 · Operator

Who operates GPTBot

Company
OpenAI
Official docs
developers.openai.com

02 · Behavior

What GPTBot does

GPTBot crawls public web pages automatically to gather training data for OpenAI's models. It is not triggered by ChatGPT users. Disallowing GPTBot in robots.txt tells OpenAI that the site's content should not be used for model training.

03 · Impact

Why GPTBot matters for your site

If you allow it

  • Your content can inform future OpenAI models
  • Blocking it has no effect on ChatGPT search visibility (that is OAI-SearchBot)

If you block it

  • No direct referral traffic from training crawls
  • Keeps your content out of OpenAI model training
  • Saves crawl bandwidth

04 · Allow

How to allow GPTBot

robots.txt

User-agent: GPTBot
Allow: /

Cloudflare

# Security › WAF › Custom rules › Create rule
Expression: (http.user_agent contains "GPTBot")
Action:     Skip › All Super Bot Fight Mode rules

# Also check Security › Bots: "Block AI bots" can block it regardless of robots.txt.

WordPress

# WordPress serves a virtual robots.txt. Edit it with your SEO plugin:
# Yoast: SEO › Tools › File editor · Rank Math: General Settings › Edit robots.txt
User-agent: GPTBot
Allow: /

nginx

# nginx serves every user agent by default.
# Make sure no rule like this blocks it:
# if ($http_user_agent ~* "GPTBot") { return 403; }

Apache

# Apache serves every user agent by default.
# Make sure .htaccess has no rule like this:
# RewriteCond %{HTTP_USER_AGENT} GPTBot [NC]
# RewriteRule .* - [F,L]

05 · Block

How to block GPTBot

GPTBot follows robots.txt, so one rule is enough. Use a server or CDN rule only to stop spoofed copies.

robots.txt

User-agent: GPTBot
Disallow: /

Cloudflare

# Security › WAF › Custom rules › Create rule
Expression: (http.user_agent contains "GPTBot")
Action:     Block

WordPress

# WordPress serves a virtual robots.txt. Edit it with your SEO plugin:
# Yoast: SEO › Tools › File editor · Rank Math: General Settings › Edit robots.txt
User-agent: GPTBot
Disallow: /

nginx

# In your server { } block:
if ($http_user_agent ~* "GPTBot") {
    return 403;
}

Apache

# .htaccess
<IfModule mod_rewrite.c>
RewriteEngine On
RewriteCond %{HTTP_USER_AGENT} GPTBot [NC]
RewriteRule .* - [F,L]
</IfModule>

06 · User agents

User agents we see for GPTBot

The user agent published by the operator:

Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; GPTBot/1.4; +https://openai.com/gptbot

07 · Verification

Is it really GPTBot?

Anyone can copy a user agent. Real GPTBot requests come from the IP ranges published at openai.com/gptbot.json.

FAQ

Questions about GPTBot

Does blocking GPTBot remove my site from ChatGPT search?

No. OpenAI uses a separate crawler, OAI-SearchBot, for ChatGPT search results, and each bot is controlled with its own robots.txt token.

How do I block GPTBot?

Add "User-agent: GPTBot" followed by "Disallow: /" to your robots.txt. OpenAI states this signals that your content should not be used for training.

How can I verify that a request really comes from GPTBot?

Check the source IP against the ranges OpenAI publishes at openai.com/gptbot.json. The user agent alone can be spoofed.

Last reviewed Oct 6, 2026

Your site

See which pages GPTBot crawls on your website

Log Hero reads your server logs and shows every AI bot, every page, every day.

Sign up with Google

Free during early access