Log in
Sign up with Google

Other bot · Internet Archive

archive.org_bot is Internet Archive’s other bot

archive.org_bot is the web crawler of the Internet Archive. It saves public web pages for the Wayback Machine and Archive-It collections.

Respects robots.txt

Our take

Your call

It does not bring search traffic, but archiving has public value and the bot respects robots.txt.

User agent

Mozilla/5.0 (compatible; archive.org_bot +http://www.archive.org/details/archive.org_bot)

01 · Operator

Who operates archive.org_bot

Type
Other bot
Official docs
radar.cloudflare.com

02 · Behavior

What archive.org_bot does

archive.org_bot fetches complete pages to store snapshots for long-term preservation. Visit frequency depends on how popular a site is and how often it changes. Archive-It partner institutions also use it for their own collections.

03 · Impact

Why archive.org_bot matters for your site

If you allow it

  • Preserves your site's history in the Wayback Machine
  • Archived copies help users and researchers

If you block it

  • Old versions of pages remain publicly viewable
  • Adds crawl load on large sites

04 · Allow

How to allow archive.org_bot

robots.txt

User-agent: archive.org_bot
Allow: /

Cloudflare

# Security › WAF › Custom rules › Create rule
Expression: (http.user_agent contains "archive.org_bot")
Action:     Skip › All Super Bot Fight Mode rules

# Also check Security › Bots: "Block AI bots" can block it regardless of robots.txt.

WordPress

# WordPress serves a virtual robots.txt. Edit it with your SEO plugin:
# Yoast: SEO › Tools › File editor · Rank Math: General Settings › Edit robots.txt
User-agent: archive.org_bot
Allow: /

nginx

# nginx serves every user agent by default.
# Make sure no rule like this blocks it:
# if ($http_user_agent ~* "archive\.org_bot") { return 403; }

Apache

# Apache serves every user agent by default.
# Make sure .htaccess has no rule like this:
# RewriteCond %{HTTP_USER_AGENT} archive\.org_bot [NC]
# RewriteRule .* - [F,L]

05 · Block

How to block archive.org_bot

archive.org_bot follows robots.txt, so one rule is enough. Use a server or CDN rule only to stop spoofed copies.

robots.txt

User-agent: archive.org_bot
Disallow: /

Cloudflare

# Security › WAF › Custom rules › Create rule
Expression: (http.user_agent contains "archive.org_bot")
Action:     Block

WordPress

# WordPress serves a virtual robots.txt. Edit it with your SEO plugin:
# Yoast: SEO › Tools › File editor · Rank Math: General Settings › Edit robots.txt
User-agent: archive.org_bot
Disallow: /

nginx

# In your server { } block:
if ($http_user_agent ~* "archive\.org_bot") {
    return 403;
}

Apache

# .htaccess
<IfModule mod_rewrite.c>
RewriteEngine On
RewriteCond %{HTTP_USER_AGENT} archive\.org_bot [NC]
RewriteRule .* - [F,L]
</IfModule>

06 · User agents

User agents we see for archive.org_bot

The user agent published by the operator:

Mozilla/5.0 (compatible; archive.org_bot +http://www.archive.org/details/archive.org_bot)

07 · Verification

Is it really archive.org_bot?

archive.org_bot’s operator publishes no IP ranges or hostnames, so requests cannot be verified. Treat the user agent as a claim, and watch your logs for unusual request rates.

FAQ

Questions about archive.org_bot

What is archive.org_bot?

It is the Internet Archive's crawler. It captures web pages for the Wayback Machine.

Does archive.org_bot respect robots.txt?

Yes. Cloudflare Radar and Known Agents list it as following robots.txt rules addressed to archive.org_bot.

Does blocking archive.org_bot remove old snapshots?

Blocking only stops future crawls. Existing snapshots need a separate removal request to the Internet Archive.

Last reviewed Oct 7, 2026

Your site

See which pages archive.org_bot crawls on your website

Log Hero reads your server logs and shows every AI bot, every page, every day.

Sign up with Google

Free during early access