01 · Operator
Who operates magpie-crawler
- Company
- Brandwatch
- Type
- Other bot
- Official docs
- brandwatch.com
02 · Behavior
What magpie-crawler does
magpie-crawler downloads public web pages for indexing and analysis by Brandwatch. It covers blogs, forums, news sites and social media content. The data supports Brandwatch's monitoring services.
03 · Impact
Why magpie-crawler matters for your site
If you allow it
- Lets brands see mentions on your site
- Respects robots.txt
- Only visits public pages
If you block it
- No direct traffic to your site
- Your content feeds a commercial analytics product
04 · Allow
How to allow magpie-crawler
robots.txt
User-agent: magpie-crawler
Allow: /
Cloudflare
# Security › WAF › Custom rules › Create rule
Expression: (http.user_agent contains "magpie-crawler")
Action: Skip › All Super Bot Fight Mode rules
# Also check Security › Bots: "Block AI bots" can block it regardless of robots.txt.
WordPress
# WordPress serves a virtual robots.txt. Edit it with your SEO plugin:
# Yoast: SEO › Tools › File editor · Rank Math: General Settings › Edit robots.txt
User-agent: magpie-crawler
Allow: /
nginx
# nginx serves every user agent by default.
# Make sure no rule like this blocks it:
# if ($http_user_agent ~* "magpie\-crawler") { return 403; }
Apache
# Apache serves every user agent by default.
# Make sure .htaccess has no rule like this:
# RewriteCond %{HTTP_USER_AGENT} magpie\-crawler [NC]
# RewriteRule .* - [F,L]
05 · Block
How to block magpie-crawler
magpie-crawler follows robots.txt, so one rule is enough. Use a server or CDN rule only to stop spoofed copies.
robots.txt
User-agent: magpie-crawler
Disallow: /
Cloudflare
# Security › WAF › Custom rules › Create rule
Expression: (http.user_agent contains "magpie-crawler")
Action: Block
WordPress
# WordPress serves a virtual robots.txt. Edit it with your SEO plugin:
# Yoast: SEO › Tools › File editor · Rank Math: General Settings › Edit robots.txt
User-agent: magpie-crawler
Disallow: /
nginx
# In your server { } block:
if ($http_user_agent ~* "magpie\-crawler") {
return 403;
}
Apache
# .htaccess
<IfModule mod_rewrite.c>
RewriteEngine On
RewriteCond %{HTTP_USER_AGENT} magpie\-crawler [NC]
RewriteRule .* - [F,L]
</IfModule>
06 · User agents
User agents we see for magpie-crawler
The user agent published by the operator:
magpie-crawler/1.1 (U; Linux amd64; en-GB; +http://www.brandwatch.net)
07 · Verification
Is it really magpie-crawler?
magpie-crawler’s operator publishes no IP ranges or hostnames, so requests cannot be verified. Treat the user agent as a claim, and watch your logs for unusual request rates.
FAQ
Questions about magpie-crawler
What is magpie-crawler?
It is Brandwatch's crawler. It downloads public pages so Brandwatch can index and analyze online conversations.
Does magpie-crawler respect robots.txt?
Yes. Brandwatch states the crawler only visits public pages and respects the robots.txt standard.
How do I block magpie-crawler?
Add "User-agent: magpie-crawler" and "Disallow: /" to your robots.txt file.
Last reviewed Oct 8, 2026