Log in
Sign up with Google

AI training crawler · Common Crawl

CCBot is Common Crawl’s AI training crawler

CCBot is the crawler of Common Crawl, a non-profit that publishes a free, open copy of the web for research and analysis.

Verifiable operator Respects robots.txt No JavaScript rendering

Our take

Your call

It feeds an open dataset commonly used for AI training, so allowing it is a content-licensing decision.

User agent

CCBot/2.0 (https://commoncrawl.org/faq/)

01 · Operator

Who operates CCBot

Company
Common Crawl
Official docs
commoncrawl.org

02 · Behavior

What CCBot does

CCBot crawls pages automatically and the results are released as the open Common Crawl dataset, which is widely used for AI training. It checks robots.txt before fetching, honors Crawl-delay and nofollow, and does not execute JavaScript. It is not triggered by users.

03 · Impact

Why CCBot matters for your site

If you allow it

  • Supports open research datasets
  • Your content may reach many downstream projects that use Common Crawl

If you block it

  • The dataset is widely used to train AI models by third parties
  • No direct referral traffic

04 · Allow

How to allow CCBot

robots.txt

User-agent: CCBot
Allow: /

Cloudflare

# Security › WAF › Custom rules › Create rule
Expression: (http.user_agent contains "CCBot")
Action:     Skip › All Super Bot Fight Mode rules

# Also check Security › Bots: "Block AI bots" can block it regardless of robots.txt.

WordPress

# WordPress serves a virtual robots.txt. Edit it with your SEO plugin:
# Yoast: SEO › Tools › File editor · Rank Math: General Settings › Edit robots.txt
User-agent: CCBot
Allow: /

nginx

# nginx serves every user agent by default.
# Make sure no rule like this blocks it:
# if ($http_user_agent ~* "CCBot") { return 403; }

Apache

# Apache serves every user agent by default.
# Make sure .htaccess has no rule like this:
# RewriteCond %{HTTP_USER_AGENT} CCBot [NC]
# RewriteRule .* - [F,L]

05 · Block

How to block CCBot

CCBot follows robots.txt, so one rule is enough. Use a server or CDN rule only to stop spoofed copies.

robots.txt

User-agent: CCBot
Disallow: /

Cloudflare

# Security › WAF › Custom rules › Create rule
Expression: (http.user_agent contains "CCBot")
Action:     Block

WordPress

# WordPress serves a virtual robots.txt. Edit it with your SEO plugin:
# Yoast: SEO › Tools › File editor · Rank Math: General Settings › Edit robots.txt
User-agent: CCBot
Disallow: /

nginx

# In your server { } block:
if ($http_user_agent ~* "CCBot") {
    return 403;
}

Apache

# .htaccess
<IfModule mod_rewrite.c>
RewriteEngine On
RewriteCond %{HTTP_USER_AGENT} CCBot [NC]
RewriteRule .* - [F,L]
</IfModule>

06 · User agents

User agents we see for CCBot

The user agent published by the operator:

CCBot/2.0 (https://commoncrawl.org/faq/)

07 · Verification

Is it really CCBot?

Anyone can copy a user agent. Real CCBot requests come from the IP ranges published at index.commoncrawl.org/ccbot.json. A reverse DNS lookup of a real request resolves to crawl.commoncrawl.org, and that hostname resolves back to the same IP.

FAQ

Questions about CCBot

Is CCBot used for AI training?

Common Crawl publishes its crawl as a free open dataset, and that dataset is widely used by others to train AI models.

Does CCBot respect Crawl-delay?

Yes. Common Crawl says CCBot honors Crawl-delay in robots.txt, for example "Crawl-delay: 2".

How do I verify CCBot?

Check that reverse DNS ends in crawl.commoncrawl.org and resolves back to the same IP, or match the IP against index.commoncrawl.org/ccbot.json.

Last reviewed Oct 6, 2026

Your site

See which pages CCBot crawls on your website

Log Hero reads your server logs and shows every AI bot, every page, every day.

Sign up with Google

Free during early access