Log in
Sign up with Google

Scraper · Adrien Barbaresi (open-source project)

Trafilatura is Adrien Barbaresi (open-source project)’s scraper

Trafilatura is an open-source Python package and command-line tool for collecting web text and metadata. It is authored by Adrien Barbaresi.

Our take

Your call

It is a general scraping tool whose use depends on who runs it.

User agent

trafilatura/2.2.0 (+https://github.com/adbar/trafilatura)

Activity

How Trafilatura crawls

Requests per site per dayLAST 28 DAYS
Sep 10Sep 19Sep 28Oct 7
What it crawls30 DAYS
  • Content pages 100%

01 · Operator

Who operates Trafilatura

Type
Scraper
Official docs
github.com

02 · Behavior

What Trafilatura does

Trafilatura downloads web pages and extracts main text, metadata and comments. It also supports crawling and outputs several formats. Anyone can run it, so requests come from many different users and servers.

03 · Impact

Why Trafilatura matters for your site

If you allow it

  • Often used for research and text analysis
  • Requests usually target only article text

If you block it

  • Anyone can run it, so the purpose varies
  • Extracted text may feed datasets or AI models
  • Cannot be verified by IP or DNS

04 · Allow

How to allow Trafilatura

robots.txt

User-agent: trafilatura
Allow: /

Cloudflare

# Security › WAF › Custom rules › Create rule
Expression: (http.user_agent contains "trafilatura")
Action:     Skip › All Super Bot Fight Mode rules

# Also check Security › Bots: "Block AI bots" can block it regardless of robots.txt.

WordPress

# WordPress serves a virtual robots.txt. Edit it with your SEO plugin:
# Yoast: SEO › Tools › File editor · Rank Math: General Settings › Edit robots.txt
User-agent: trafilatura
Allow: /

nginx

# nginx serves every user agent by default.
# Make sure no rule like this blocks it:
# if ($http_user_agent ~* "trafilatura") { return 403; }

Apache

# Apache serves every user agent by default.
# Make sure .htaccess has no rule like this:
# RewriteCond %{HTTP_USER_AGENT} trafilatura [NC]
# RewriteRule .* - [F,L]

05 · Block

How to block Trafilatura

Start with robots.txt. If Trafilatura keeps showing up in your logs, block it at your CDN or web server.

robots.txt

User-agent: trafilatura
Disallow: /

Cloudflare

# Security › WAF › Custom rules › Create rule
Expression: (http.user_agent contains "trafilatura")
Action:     Block

WordPress

# WordPress serves a virtual robots.txt. Edit it with your SEO plugin:
# Yoast: SEO › Tools › File editor · Rank Math: General Settings › Edit robots.txt
User-agent: trafilatura
Disallow: /

nginx

# In your server { } block:
if ($http_user_agent ~* "trafilatura") {
    return 403;
}

Apache

# .htaccess
<IfModule mod_rewrite.c>
RewriteEngine On
RewriteCond %{HTTP_USER_AGENT} trafilatura [NC]
RewriteRule .* - [F,L]
</IfModule>

06 · User agents

User agents we see for Trafilatura

User agentShareLast seenStatus
trafilatura/2.2.0 (+https://github.com/adbar/trafilatura) 92% Unverified
trafilatura/2.0.0 (+https://github.com/adbar/trafilatura) 8.5% Unverified

07 · Verification

Is it really Trafilatura?

Trafilatura’s operator publishes no IP ranges or hostnames, so requests cannot be verified. Treat the user agent as a claim, and watch your logs for unusual request rates.

FAQ

Questions about Trafilatura

What is trafilatura in my server logs?

It is the user agent of the trafilatura Python package. Someone used it to download and extract text from your pages.

Who runs the trafilatura bot?

No single operator runs it. The software is open source under the Apache 2.0 license, and anyone can run it.

How do I block trafilatura?

Match the user agent token trafilatura at your server or firewall. Robots.txt handling is not documented on the project page.

Last reviewed Oct 8, 2026

Your site

See which pages Trafilatura crawls on your website

Log Hero reads your server logs and shows every AI bot, every page, every day.

Sign up with Google

Free during early access