Log in
Sign up with Google

Search engine crawler · Semantic Scholar (Ai2)

SemanticScholarBot is Semantic Scholar (Ai2)’s search engine crawler

SemanticScholarBot is the crawler of Semantic Scholar, an academic search engine based at Ai2. It looks for academic PDFs.

Our take

Allow

It indexes academic papers for a widely used scholarly search engine.

User agent

Mozilla/5.0 (compatible) SemanticScholarBot (+https://www.semanticscholar.org/crawler)

01 · Operator

Who operates SemanticScholarBot

Official docs
semanticscholar.org

02 · Behavior

What SemanticScholarBot does

The bot crawls certain domains to find academic PDFs. The papers feed Semantic Scholar's literature search. Semantic Scholar provides a contact option for questions about the crawler.

03 · Impact

Why SemanticScholarBot matters for your site

If you allow it

  • Makes your research papers discoverable in an academic search engine
  • Crawls focus on academic PDFs, not whole sites

If you block it

  • No benefit if your site has no academic papers
  • Robots.txt behavior is not documented on its crawler page

04 · Allow

How to allow SemanticScholarBot

robots.txt

User-agent: SemanticScholarBot
Allow: /

Cloudflare

# Security › WAF › Custom rules › Create rule
Expression: (http.user_agent contains "SemanticScholarBot")
Action:     Skip › All Super Bot Fight Mode rules

# Also check Security › Bots: "Block AI bots" can block it regardless of robots.txt.

WordPress

# WordPress serves a virtual robots.txt. Edit it with your SEO plugin:
# Yoast: SEO › Tools › File editor · Rank Math: General Settings › Edit robots.txt
User-agent: SemanticScholarBot
Allow: /

nginx

# nginx serves every user agent by default.
# Make sure no rule like this blocks it:
# if ($http_user_agent ~* "SemanticScholarBot") { return 403; }

Apache

# Apache serves every user agent by default.
# Make sure .htaccess has no rule like this:
# RewriteCond %{HTTP_USER_AGENT} SemanticScholarBot [NC]
# RewriteRule .* - [F,L]

05 · Block

How to block SemanticScholarBot

Start with robots.txt. If SemanticScholarBot keeps showing up in your logs, block it at your CDN or web server.

robots.txt

User-agent: SemanticScholarBot
Disallow: /

Cloudflare

# Security › WAF › Custom rules › Create rule
Expression: (http.user_agent contains "SemanticScholarBot")
Action:     Block

WordPress

# WordPress serves a virtual robots.txt. Edit it with your SEO plugin:
# Yoast: SEO › Tools › File editor · Rank Math: General Settings › Edit robots.txt
User-agent: SemanticScholarBot
Disallow: /

nginx

# In your server { } block:
if ($http_user_agent ~* "SemanticScholarBot") {
    return 403;
}

Apache

# .htaccess
<IfModule mod_rewrite.c>
RewriteEngine On
RewriteCond %{HTTP_USER_AGENT} SemanticScholarBot [NC]
RewriteRule .* - [F,L]
</IfModule>

06 · User agents

User agents we see for SemanticScholarBot

User agentShareLast seenStatus
Mozilla/5.0 (compatible) SemanticScholarBot (+https://www.semanticscholar.org/crawler) 100% Unverified

07 · Verification

Is it really SemanticScholarBot?

SemanticScholarBot’s operator publishes no IP ranges or hostnames, so requests cannot be verified. Treat the user agent as a claim, and watch your logs for unusual request rates.

FAQ

Questions about SemanticScholarBot

What is SemanticScholarBot?

It is the crawler of Semantic Scholar, an academic search engine based at Ai2. It crawls certain domains to find academic PDFs.

Why is SemanticScholarBot downloading PDFs from my site?

The bot looks for academic PDFs to include in Semantic Scholar. It targets domains likely to host research papers.

How do I contact Semantic Scholar about its crawler?

The crawler page at semanticscholar.org/crawler links to a contact form for questions or concerns.

Last reviewed Oct 8, 2026

Your site

See which pages SemanticScholarBot crawls on your website

Log Hero reads your server logs and shows every AI bot, every page, every day.

Sign up with Google

Free during early access