01 · Operator
Who operates Arquivo-web-crawler
- Company
- Arquivo.pt (FCCN)
- Type
- Other bot
02 · Behavior
What Arquivo-web-crawler does
The crawler is based on Heritrix, archiving software from the Internet Archive. It captures full pages and stores snapshots in the Arquivo.pt archive. It has documented a 10-second pause between requests to the same site.
03 · Impact
Why Arquivo-web-crawler matters for your site
If you allow it
- Preserves your site in a national public archive
- Polite crawl rate
If you block it
- No search traffic benefit
- Little relevance for sites outside Portugal
04 · Allow
How to allow Arquivo-web-crawler
robots.txt
User-agent: Arquivo-web-crawler
Allow: /
Cloudflare
# Security › WAF › Custom rules › Create rule
Expression: (http.user_agent contains "Arquivo-web-crawler")
Action: Skip › All Super Bot Fight Mode rules
# Also check Security › Bots: "Block AI bots" can block it regardless of robots.txt.
WordPress
# WordPress serves a virtual robots.txt. Edit it with your SEO plugin:
# Yoast: SEO › Tools › File editor · Rank Math: General Settings › Edit robots.txt
User-agent: Arquivo-web-crawler
Allow: /
nginx
# nginx serves every user agent by default.
# Make sure no rule like this blocks it:
# if ($http_user_agent ~* "Arquivo\-web\-crawler") { return 403; }
Apache
# Apache serves every user agent by default.
# Make sure .htaccess has no rule like this:
# RewriteCond %{HTTP_USER_AGENT} Arquivo\-web\-crawler [NC]
# RewriteRule .* - [F,L]
05 · Block
How to block Arquivo-web-crawler
Arquivo-web-crawler follows robots.txt, so one rule is enough. Use a server or CDN rule only to stop spoofed copies.
robots.txt
User-agent: Arquivo-web-crawler
Disallow: /
Cloudflare
# Security › WAF › Custom rules › Create rule
Expression: (http.user_agent contains "Arquivo-web-crawler")
Action: Block
WordPress
# WordPress serves a virtual robots.txt. Edit it with your SEO plugin:
# Yoast: SEO › Tools › File editor · Rank Math: General Settings › Edit robots.txt
User-agent: Arquivo-web-crawler
Disallow: /
nginx
# In your server { } block:
if ($http_user_agent ~* "Arquivo\-web\-crawler") {
return 403;
}
Apache
# .htaccess
<IfModule mod_rewrite.c>
RewriteEngine On
RewriteCond %{HTTP_USER_AGENT} Arquivo\-web\-crawler [NC]
RewriteRule .* - [F,L]
</IfModule>
06 · User agents
User agents we see for Arquivo-web-crawler
The user agent published by the operator:
Arquivo-web-crawler (compatible; heritrix/3.4.0-20200304 +https://arquivo.pt/faq-crawling)
07 · Verification
Is it really Arquivo-web-crawler?
Arquivo-web-crawler’s operator publishes no IP ranges or hostnames, so requests cannot be verified. Treat the user agent as a claim, and watch your logs for unusual request rates.
FAQ
Questions about Arquivo-web-crawler
Who operates Arquivo-web-crawler?
It is run by Arquivo.pt, the Portuguese web archive operated by FCCN.
Does Arquivo-web-crawler respect robots.txt?
Yes. Arquivo.pt states that site owners can exclude it with standard robots.txt rules.
How do I block Arquivo-web-crawler?
Add 'User-agent: Arquivo-web-crawler' and 'Disallow: /' to robots.txt. Existing snapshots are not removed by this.
Last reviewed Oct 7, 2026