Log in
Sign up with Google

Monitoring bot · Natural Language Processing Centre, Masaryk University (Brno)

SpiderLing is Natural Language Processing Centre, Masaryk University (Brno)’s monitoring bot

SpiderLing is a research crawler of the Natural Language Processing Centre at Masaryk University in Brno. It downloads web text for linguistic research.

Respects robots.txt

Our take

Your call

It is an academic crawler that obeys robots.txt and brings no direct traffic.

User agent

Mozilla/5.0 (compatible; SpiderLing (a SPIDER for LINGustic research); +http://nlp.fi.muni.cz/projects/biwec/)

Activity

How SpiderLing crawls

Requests per site per dayLAST 26 DAYS
Sep 12Sep 20Sep 29Oct 7
What it crawls30 DAYS
  • Content pages 100%
  • Other 0.2%

01 · Operator

Who operates SpiderLing

Official docs
nlp.fi.muni.cz

02 · Behavior

What SpiderLing does

SpiderLing collects texts from the web for the BiWeC project. The text is cleaned and annotated with part-of-speech tags and lemmas. The corpora are used to study language, not page content.

03 · Impact

Why SpiderLing matters for your site

If you allow it

  • Supports academic language research
  • Obeys robots.txt

If you block it

  • No traffic benefit
  • Your text may end up in linguistic corpora

04 · Allow

How to allow SpiderLing

robots.txt

User-agent: SpiderLing
Allow: /

Cloudflare

# Security › WAF › Custom rules › Create rule
Expression: (http.user_agent contains "SpiderLing")
Action:     Skip › All Super Bot Fight Mode rules

# Also check Security › Bots: "Block AI bots" can block it regardless of robots.txt.

WordPress

# WordPress serves a virtual robots.txt. Edit it with your SEO plugin:
# Yoast: SEO › Tools › File editor · Rank Math: General Settings › Edit robots.txt
User-agent: SpiderLing
Allow: /

nginx

# nginx serves every user agent by default.
# Make sure no rule like this blocks it:
# if ($http_user_agent ~* "SpiderLing") { return 403; }

Apache

# Apache serves every user agent by default.
# Make sure .htaccess has no rule like this:
# RewriteCond %{HTTP_USER_AGENT} SpiderLing [NC]
# RewriteRule .* - [F,L]

05 · Block

How to block SpiderLing

SpiderLing follows robots.txt, so one rule is enough. Use a server or CDN rule only to stop spoofed copies.

robots.txt

User-agent: SpiderLing
Disallow: /

Cloudflare

# Security › WAF › Custom rules › Create rule
Expression: (http.user_agent contains "SpiderLing")
Action:     Block

WordPress

# WordPress serves a virtual robots.txt. Edit it with your SEO plugin:
# Yoast: SEO › Tools › File editor · Rank Math: General Settings › Edit robots.txt
User-agent: SpiderLing
Disallow: /

nginx

# In your server { } block:
if ($http_user_agent ~* "SpiderLing") {
    return 403;
}

Apache

# .htaccess
<IfModule mod_rewrite.c>
RewriteEngine On
RewriteCond %{HTTP_USER_AGENT} SpiderLing [NC]
RewriteRule .* - [F,L]
</IfModule>

06 · User agents

User agents we see for SpiderLing

User agentShareLast seenStatus
Mozilla/5.0 (compatible; SpiderLing; +https://www.sketchengine.eu/crawler/) 100% Unverified

07 · Verification

Is it really SpiderLing?

SpiderLing’s operator publishes no IP ranges or hostnames, so requests cannot be verified. Treat the user agent as a claim, and watch your logs for unusual request rates.

FAQ

Questions about SpiderLing

Who operates SpiderLing?

The Natural Language Processing Centre of Masaryk University in Brno runs it for the BiWeC project.

How do I block SpiderLing?

Add "User-agent: SpiderLing" and "Disallow: /" to robots.txt. The operator states that it adheres to the Robots Exclusion Protocol.

What is the collected text used for?

It is used to build annotated text corpora for computational linguistics research.

Last reviewed Oct 8, 2026

Your site

See which pages SpiderLing crawls on your website

Log Hero reads your server logs and shows every AI bot, every page, every day.

Sign up with Google

Free during early access