01 · Operator
Who operates SemanticScholarBot
- Company
- Semantic Scholar (Ai2)
- Official docs
- semanticscholar.org
02 · Behavior
What SemanticScholarBot does
The bot crawls certain domains to find academic PDFs. The papers feed Semantic Scholar's literature search. Semantic Scholar provides a contact option for questions about the crawler.
03 · Impact
Why SemanticScholarBot matters for your site
If you allow it
- Makes your research papers discoverable in an academic search engine
- Crawls focus on academic PDFs, not whole sites
If you block it
- No benefit if your site has no academic papers
- Robots.txt behavior is not documented on its crawler page
04 · Allow
How to allow SemanticScholarBot
robots.txt
User-agent: SemanticScholarBot
Allow: /
Cloudflare
# Security › WAF › Custom rules › Create rule
Expression: (http.user_agent contains "SemanticScholarBot")
Action: Skip › All Super Bot Fight Mode rules
# Also check Security › Bots: "Block AI bots" can block it regardless of robots.txt.
WordPress
# WordPress serves a virtual robots.txt. Edit it with your SEO plugin:
# Yoast: SEO › Tools › File editor · Rank Math: General Settings › Edit robots.txt
User-agent: SemanticScholarBot
Allow: /
nginx
# nginx serves every user agent by default.
# Make sure no rule like this blocks it:
# if ($http_user_agent ~* "SemanticScholarBot") { return 403; }
Apache
# Apache serves every user agent by default.
# Make sure .htaccess has no rule like this:
# RewriteCond %{HTTP_USER_AGENT} SemanticScholarBot [NC]
# RewriteRule .* - [F,L]
05 · Block
How to block SemanticScholarBot
Start with robots.txt. If SemanticScholarBot keeps showing up in your logs, block it at your CDN or web server.
robots.txt
User-agent: SemanticScholarBot
Disallow: /
Cloudflare
# Security › WAF › Custom rules › Create rule
Expression: (http.user_agent contains "SemanticScholarBot")
Action: Block
WordPress
# WordPress serves a virtual robots.txt. Edit it with your SEO plugin:
# Yoast: SEO › Tools › File editor · Rank Math: General Settings › Edit robots.txt
User-agent: SemanticScholarBot
Disallow: /
nginx
# In your server { } block:
if ($http_user_agent ~* "SemanticScholarBot") {
return 403;
}
Apache
# .htaccess
<IfModule mod_rewrite.c>
RewriteEngine On
RewriteCond %{HTTP_USER_AGENT} SemanticScholarBot [NC]
RewriteRule .* - [F,L]
</IfModule>
06 · User agents
User agents we see for SemanticScholarBot
| User agent | Share | Last seen | Status |
|---|---|---|---|
Mozilla/5.0 (compatible) SemanticScholarBot (+https://www.semanticscholar.org/crawler) |
100% | Unverified |
07 · Verification
Is it really SemanticScholarBot?
SemanticScholarBot’s operator publishes no IP ranges or hostnames, so requests cannot be verified. Treat the user agent as a claim, and watch your logs for unusual request rates.
FAQ
Questions about SemanticScholarBot
What is SemanticScholarBot?
It is the crawler of Semantic Scholar, an academic search engine based at Ai2. It crawls certain domains to find academic PDFs.
Why is SemanticScholarBot downloading PDFs from my site?
The bot looks for academic PDFs to include in Semantic Scholar. It targets domains likely to host research papers.
How do I contact Semantic Scholar about its crawler?
The crawler page at semanticscholar.org/crawler links to a contact form for questions or concerns.
Last reviewed Oct 8, 2026