01 · Operator
Who operates TerraCotta
- Company
- Ceramic AI
- Type
- AI search crawler
- Official docs
- github.com
02 · Behavior
What TerraCotta does
TerraCotta fetches public pages and adds them to Ceramic's search index. Ceramic offers this index as a web-scale search API for AI and LLM products. Its GitHub page says it respects robots.txt and publishes the IP addresses it crawls from.
03 · Impact
Why TerraCotta matters for your site
If you allow it
- Content can appear in AI products that use Ceramic's search API
- Operator documents its crawler and follows robots.txt
- Published IP list makes verification simple
If you block it
- Direct traffic benefit is not yet shown
- Your content feeds a commercial data product
- Adds crawl load
04 · Allow
How to allow TerraCotta
robots.txt
User-agent: TerraCotta
Allow: /
Cloudflare
# Security › WAF › Custom rules › Create rule
Expression: (http.user_agent contains "TerraCotta")
Action: Skip › All Super Bot Fight Mode rules
# Also check Security › Bots: "Block AI bots" can block it regardless of robots.txt.
WordPress
# WordPress serves a virtual robots.txt. Edit it with your SEO plugin:
# Yoast: SEO › Tools › File editor · Rank Math: General Settings › Edit robots.txt
User-agent: TerraCotta
Allow: /
nginx
# nginx serves every user agent by default.
# Make sure no rule like this blocks it:
# if ($http_user_agent ~* "TerraCotta") { return 403; }
Apache
# Apache serves every user agent by default.
# Make sure .htaccess has no rule like this:
# RewriteCond %{HTTP_USER_AGENT} TerraCotta [NC]
# RewriteRule .* - [F,L]
05 · Block
How to block TerraCotta
TerraCotta follows robots.txt, so one rule is enough. Use a server or CDN rule only to stop spoofed copies.
robots.txt
User-agent: TerraCotta
Disallow: /
Cloudflare
# Security › WAF › Custom rules › Create rule
Expression: (http.user_agent contains "TerraCotta")
Action: Block
WordPress
# WordPress serves a virtual robots.txt. Edit it with your SEO plugin:
# Yoast: SEO › Tools › File editor · Rank Math: General Settings › Edit robots.txt
User-agent: TerraCotta
Disallow: /
nginx
# In your server { } block:
if ($http_user_agent ~* "TerraCotta") {
return 403;
}
Apache
# .htaccess
<IfModule mod_rewrite.c>
RewriteEngine On
RewriteCond %{HTTP_USER_AGENT} TerraCotta [NC]
RewriteRule .* - [F,L]
</IfModule>
06 · User agents
User agents we see for TerraCotta
| User agent | Share | Last seen | Status |
|---|---|---|---|
TerraCotta 0.2 https://www.github.com/ceramicTeam/CeramicTerracotta |
100% | Unverified | |
TerraCotta 0.2 https://www.github.com/ceramicTeam/CeramicTerracotta |
0.5% | Verified |
07 · Verification
Is it really TerraCotta?
Anyone can copy a user agent. Real TerraCotta requests come from the IP ranges published at github.com/CeramicTeam/CeramicTerracotta/blob/main/ipList.txt.
IP addresses we saw from verified TerraCotta requests
| IP range | Share | Last seen |
|---|---|---|
54.193.6.55 |
100% |
Always check against the operator’s current list; ranges change.
FAQ
Questions about TerraCotta
Who operates the TerraCotta crawler?
TerraCotta is operated by Ceramic AI. The crawler's GitHub page gives crawler[at]ceramic[dot]ai as contact.
Does TerraCotta respect robots.txt?
Yes. Ceramic describes it as a crawler that respects robots.txt. Use the token TerraCotta in a User-agent line.
How can I verify TerraCotta requests?
Compare the request IP with the list in ipList.txt in Ceramic's CeramicTerracotta GitHub repository.
Last reviewed Oct 8, 2026