Log in
Sign up with Google

Monitoring bot · OpenINTEL

OI-Crawler is OpenINTEL’s monitoring bot

OI-Crawler is the web crawler of OpenINTEL, an internet measurement project. It fetches landing pages for infrastructure research.

Respects robots.txt

Our take

Your call

It is a low-impact research crawler with no direct benefit to site owners.

User agent

OI-Crawler/Nutch (https://openintel.nl/webcrawl/)

01 · Operator

Who operates OI-Crawler

Company
OpenINTEL
Official docs
openintel.nl

02 · Behavior

What OI-Crawler does

OI-Crawler fetches the landing page of apex domains on OpenINTEL's list and follows redirects. It first reads and respects robots.txt. It uses an open-source crawler (Apache Nutch, per its user agent) and keeps request rates low. Traffic comes from a single IPv4 /26 prefix.

03 · Impact

Why OI-Crawler matters for your site

If you allow it

  • Supports public internet measurement research
  • Fetches only the landing page and respects robots.txt

If you block it

  • No search traffic or direct benefit
  • Hosts are disclosed only on request

04 · Allow

How to allow OI-Crawler

robots.txt

User-agent: OI-Crawler
Allow: /

Cloudflare

# Security › WAF › Custom rules › Create rule
Expression: (http.user_agent contains "OI-Crawler")
Action:     Skip › All Super Bot Fight Mode rules

# Also check Security › Bots: "Block AI bots" can block it regardless of robots.txt.

WordPress

# WordPress serves a virtual robots.txt. Edit it with your SEO plugin:
# Yoast: SEO › Tools › File editor · Rank Math: General Settings › Edit robots.txt
User-agent: OI-Crawler
Allow: /

nginx

# nginx serves every user agent by default.
# Make sure no rule like this blocks it:
# if ($http_user_agent ~* "OI\-Crawler") { return 403; }

Apache

# Apache serves every user agent by default.
# Make sure .htaccess has no rule like this:
# RewriteCond %{HTTP_USER_AGENT} OI\-Crawler [NC]
# RewriteRule .* - [F,L]

05 · Block

How to block OI-Crawler

OI-Crawler follows robots.txt, so one rule is enough. Use a server or CDN rule only to stop spoofed copies.

robots.txt

User-agent: OI-Crawler
Disallow: /

Cloudflare

# Security › WAF › Custom rules › Create rule
Expression: (http.user_agent contains "OI-Crawler")
Action:     Block

WordPress

# WordPress serves a virtual robots.txt. Edit it with your SEO plugin:
# Yoast: SEO › Tools › File editor · Rank Math: General Settings › Edit robots.txt
User-agent: OI-Crawler
Disallow: /

nginx

# In your server { } block:
if ($http_user_agent ~* "OI\-Crawler") {
    return 403;
}

Apache

# .htaccess
<IfModule mod_rewrite.c>
RewriteEngine On
RewriteCond %{HTTP_USER_AGENT} OI\-Crawler [NC]
RewriteRule .* - [F,L]
</IfModule>

06 · User agents

User agents we see for OI-Crawler

The user agent published by the operator:

OI-Crawler/Nutch (https://openintel.nl/webcrawl/)

07 · Verification

Is it really OI-Crawler?

OI-Crawler’s operator publishes no IP ranges or hostnames, so requests cannot be verified. Treat the user agent as a claim, and watch your logs for unusual request rates.

FAQ

Questions about OI-Crawler

What is OI-Crawler?

It is OpenINTEL's crawler. It fetches domain landing pages for internet measurement research.

Does OI-Crawler respect robots.txt?

Yes. OpenINTEL says it fetches and respects robots.txt before crawling.

How do I stop OI-Crawler?

Disallow OI-Crawler in robots.txt, or contact OpenINTEL using the details on its webcrawl page.

Last reviewed Oct 8, 2026

Your site

See which pages OI-Crawler crawls on your website

Log Hero reads your server logs and shows every AI bot, every page, every day.

Sign up with Google

Free during early access