Log in
Sign up with Google

Search engine crawler · Openindex B.V.

OpenindexSpider is Openindex B.V.’s search engine crawler

OpenindexSpider is the web crawler of Openindex, a Dutch search technology company. It crawls for research and development of search engines.

Respects robots.txt

Our take

Your call

It is well behaved but brings no direct visibility to most sites.

User agent

Mozilla/5.0 (compatible; OpenindexSpider; +https://www.openindex.io/saas/about-our-spider/)

01 · Operator

Who operates OpenindexSpider

Company
Openindex B.V.
Official docs
openindex.io

02 · Behavior

What OpenindexSpider does

It runs on Apache Nutch crawlers on a Hadoop cluster. Openindex uses the data to develop universal and focused search engines. It obeys robots.txt and Crawl-delay. It aims to request each host no more than once every few seconds.

03 · Impact

Why OpenindexSpider matters for your site

If you allow it

  • Obeys robots.txt and Crawl-delay
  • Polite crawl rate

If you block it

  • No public search engine that sends you traffic
  • Uses bandwidth for third-party research

04 · Allow

How to allow OpenindexSpider

robots.txt

User-agent: OpenindexSpider
Allow: /

Cloudflare

# Security › WAF › Custom rules › Create rule
Expression: (http.user_agent contains "OpenindexSpider")
Action:     Skip › All Super Bot Fight Mode rules

# Also check Security › Bots: "Block AI bots" can block it regardless of robots.txt.

WordPress

# WordPress serves a virtual robots.txt. Edit it with your SEO plugin:
# Yoast: SEO › Tools › File editor · Rank Math: General Settings › Edit robots.txt
User-agent: OpenindexSpider
Allow: /

nginx

# nginx serves every user agent by default.
# Make sure no rule like this blocks it:
# if ($http_user_agent ~* "OpenindexSpider") { return 403; }

Apache

# Apache serves every user agent by default.
# Make sure .htaccess has no rule like this:
# RewriteCond %{HTTP_USER_AGENT} OpenindexSpider [NC]
# RewriteRule .* - [F,L]

05 · Block

How to block OpenindexSpider

OpenindexSpider follows robots.txt, so one rule is enough. Use a server or CDN rule only to stop spoofed copies.

robots.txt

User-agent: OpenindexSpider
Disallow: /

Cloudflare

# Security › WAF › Custom rules › Create rule
Expression: (http.user_agent contains "OpenindexSpider")
Action:     Block

WordPress

# WordPress serves a virtual robots.txt. Edit it with your SEO plugin:
# Yoast: SEO › Tools › File editor · Rank Math: General Settings › Edit robots.txt
User-agent: OpenindexSpider
Disallow: /

nginx

# In your server { } block:
if ($http_user_agent ~* "OpenindexSpider") {
    return 403;
}

Apache

# .htaccess
<IfModule mod_rewrite.c>
RewriteEngine On
RewriteCond %{HTTP_USER_AGENT} OpenindexSpider [NC]
RewriteRule .* - [F,L]
</IfModule>

06 · User agents

User agents we see for OpenindexSpider

The user agent published by the operator:

Mozilla/5.0 (compatible; OpenindexSpider; +https://www.openindex.io/saas/about-our-spider/)

07 · Verification

Is it really OpenindexSpider?

OpenindexSpider’s operator publishes no IP ranges or hostnames, so requests cannot be verified. Treat the user agent as a claim, and watch your logs for unusual request rates.

FAQ

Questions about OpenindexSpider

How do I block OpenindexSpider?

Add "User-agent: OpenindexSpider" with "Disallow: /" to robots.txt. Openindex says it also responds to the name "Openindex".

Does OpenindexSpider respect Crawl-delay?

Yes. Openindex states the spider respects the Crawl-delay directive.

What software does OpenindexSpider use?

Openindex says it uses enhanced Apache Nutch crawlers running on an Apache Hadoop cluster.

Last reviewed Oct 8, 2026

Your site

See which pages OpenindexSpider crawls on your website

Log Hero reads your server logs and shows every AI bot, every page, every day.

Sign up with Google

Free during early access