Semantic Scholar
-
Semantic Scholar — Full Profile ( 2026-09)
Headquartered in Seattle, Washington, the Allen Institute for AI (AI2) is an independent 501(c)(3) nonprofit research institute founded on the endowment of late Microsoft co‑founder Paul G. Allen. First launched in 2015 under the AI2 umbrella, Semantic Scholar operates as a public-good AI primitive dedicated to academic literature intelligence. The institute’s funding primarily derives from the Allen estate and major public research grants, including a $2M NSF award in 2022 and a joint $152M NSF/Nvidia research grant secured in August 2025. Institutionally, AI2 is strictly non-VC-backed, non-bootstrapped, non-publicly-traded, and unacquired, and it is not classified as a Google/META platform primitive. As of current benchmarking, Semantic Scholar maintains no official first-party MCP server. All agent integrations rely entirely on community-built MCP bridges — such as the open-source smaniches/semantic-scholar-mcp — that wrap its public REST APIs, with zero vendor-hosted or officially supported MCP infrastructure. This positioning places Semantic Scholar firmly within the no-first-party MCP cluster alongside Motion, Todoist, Qualtrics and Documenso, clearly differentiating it from the first-party MCP cohort including Klaviyo, n8n, Databox, Typeform and Gmail.Website
Field Value Official Website https://semanticscholar.org · API: https://www.semanticscholar.org/product/api · docs: https://semanticscholar.org/product/api/tutorial · REST base: api.semanticscholar.org/graph/v1 · recommendations/v1 · datasets/v1 · AI2: allenai.org Company Allen Institute for Artificial Intelligence (AI2), Seattle; founded 2014 by Paul G. Allen (nonprofit); Semantic Scholar launched 2015; nonprofit-independent public-good (AI2 endowment + NSF + Nvidia grants; ~152M 2025 NSF/Nvidia 5-yr open-science-models grant); 236M+ papers / 2.49B citations / 79M authors (Sep 2026); no SOC 2/commercial-BAA (academic infra), GDPR-aware, open-data license; Cloud-hosted (AWS/GCP), no on-prem. Status Active, S2AG v1 stable, TLDR + citation-context + SPECTER2 embeddings live, Semantic Reader beta, no MCP server published by AI2 Category ai-productivity-data Pipeline Stage none Title
Field Value Product Semantic Scholar (web) + Semantic Scholar Academic Graph (S2AG) API Descriptive Title Semantic Scholar — AI-Enriched Academic Literature Search Engine & Open Research Graph (Allen AI nonprofit): 236M+ papers across STM+SSH, ML-extracted TLDR one-sentence summaries, citation-context classification (supportive/contrasting/mentioning), influence-weighted ranking (SPECTER2 embeddings), author/venue/citation graph, paper recommendations, bulk S2AG dataset downloads (monthly JSON snapshots) + REST API (Academic Graph / Recommendations / Datasets, 1 RPS authenticated intro, 300/100s shared unauth) + no first-party MCP (community Claude/Cursor bridges wrap REST only) + Semantic Reader beta (inline citation context, def lookup); permanently free no paid tier (AI2 nonprofit endowment, reverse Motion per-seat / reverse Qualtrics per-interaction / reverse Nanonets block-run / reverse Databox AI-credit); 2015 AI2 Seattle Paul-Allen nonprofit-independent public-good MCP-none-community-only One-line Positioning Researcher asks “scaling laws for transformers 2022-2024” → Semantic Scholar returns ranked papers with TLDR + citation-context chips + SPECTER2 neighbors → API consumer pulls same graph at 1 RPS with free key → agents mount community MCP (smaniches) that calls REST under the hood, no vendor mcp.semanticscholar.org; cost is $0 forever (grant-funded), true constraint is rate (1 RPS key / shared unauth pool), ownership AI2 nonprofit (distinct from vc-independent n8n/Databox, public-traded GOOGL/FB primitives, pe-private Qualtrics). Features
Field Value Search Core Keyword + ML semantic ranking over 236M+ papers, filters (year/venue/field/citation-count/open-access), author profile pages, venue pages, citation graph visualization links, saved searches/libraries (free account) AI Enrichment TLDR (one-sentence AI summary per paper), citation context (supportive / contrasting / mentioning tags on each citing span), influence-weighted ranking (SPECTER2 embeddings, not raw citation count), Auto-Abstract hints, related-paper recommendations, author alerts Semantic Reader (beta) Inline citation context popovers, definition lookup, related-work surfacing inside PDF view, annotation sync API Surface Academic Graph v1 (paper/author/venue/citation/reference/batch/details, fields param, SPECTER2 embeddings on select endpoints), Recommendations v1 (paper-to-paper similarity, positive/negative seeds), Datasets v1 (S2AG bulk JSON snapshots, monthly release, incremental updates API); REST only, no GraphQL, no gRPC MCP None first-party — AI2 publishes no mcp.semanticscholar.org; community servers (github.com/smaniches/semantic-scholar-mcp, Glama-listed) expose ~14 tools (paper_search / paper_details / author_batch / snippet_search / multi_recommend / status) by calling REST with env API key, 10 RPS client-side claim but upstream still 1 RPS AI2-enforced; → none-cluster, community-bridge-only, aligns Documenso/Qualtrics/MotionAuth / Rate No auth for basic search; API key free via request form → 1 RPS introductory across all endpoints, unauth shared pool ~300/100s effective lower at peak; higher RPS by use-case review (academic partner 10-100 RPS); 429 + exponential backoff; no per-call billing Data Sources arXiv/bioRxiv/medRxiv/PubMed/ACM/IEEE/DBLP/Springer Nature/Wiley/Taylor & Francis/SAGE/Unpaywall/Crossref; no Scopus/WoS native citation merge Pricing (2026-09, permanently free, no paid tier, rate-limited not credit-metered)
Field Value Web App $0 — full search, TLDR, citation context, libraries, alerts, Semantic Reader beta; no account needed for search, free account for saves API Unauthenticated $0 — shared pool ~300 req/100s across all anon users, throttled at peak, 429 heavy API Key (free) $0 — request form, email-delivered, 1 RPS introductory across all endpoints, higher on review; no quota dollars, no credit pool Bulk Datasets $0 — S2AG JSON snapshots monthly, incremental updates API, host yourself; no egress fee stated Type free-forever-no-tier-rate-limited; reverse Motion per-seat, reverse Qualtrics per-interaction, reverse Nanonets block-run, reverse Databox AI-credit, reverse Gmail 80M-unit threshold; cost lever = RPS not dollars Note AI2 $152M NSF/Nvidia 2025 grant funds infra; “free” is structural nonprofit mission, not growth-hacking freemium; no enterprise SLA, no priority support contract (email only) Reviews
Field Value G2 No standalone claimed “Semantic Scholar” G2 profile — AI Agent Index notes “G2 shows 0 verified reviews with an unclaimed profile” (academic user base not G2-rating); print no-rated-b2b-profile-standalone-2026(same rule as MonkeyLearn-standalone / FB-Ads-API / Gmail-API / Documenso / FFmpeg)Capterra None — no standalone Capterra page; academic tool not procured via Capterra Trustpilot No B2C page (B2B/academic infra) → no-rated-consumer-profileexcluded (aligns FFmpeg/MonkeyLearn/Qualtrics/Feedly/n8n/Databox/Documenso/Gmail-API)Editorial / AI Indices AI Agent Index 3.1/5 (Autonomy 2/5, Pricing clarity 5/5, Setup 5/5), APIs.io Kin 56.2/100 “developing”, agent-readiness 39/100 “agent aware” (MCP Server 0/12 pre-community-bridge), ToolRadar-style 4.4 academic-sentiment norm; dev consensus “REST production-grade for research, MCP absent first-party, 1 RPS key is slow for agents” Downstream Ecosystem Proof Elicit / Consensus / ResearchRabbit / Connected Papers / Litmaps / Sourcely all build on S2AG — de-facto academic-graph standard, stronger signal than star reviews Praise Free 236M-corpus + TLDR + citation-context + SPECTER2, no paywall unlike Scopus/WoS, REST clean + bulk dataset, nonprofit longevity (AI2 endowment+NSF), powers entire AI-research-tool ecosystem Gripes 1 RPS authenticated intro is slow for agentic crawls (must apply for higher), no first-party MCP (community bridge adds latency/opacity), no Scopus/WoS citation merge (undercounts impact), no Zotero/Mendeley native export, coverage bias toward CS/biomed (SSH thinner), no conversational synthesis (search not RAG-answer) Review Sources Breakdown
Field Value G2.com no-rated-b2b-profile-standalone-2026— unclaimed 0-review row, academic not bought-SaaS G2 patternCapterra none → no-standalone-capterra-profileTrustpilot none → no-rated-consumer-profileexcludedAI Agent Index / APIs.io / ToolRadar / dev blogs 3.1/5 + 56.2/100 + 4.4-norm directional, not G2-grade Ecosystem adoption (Elicit/Consensus/Connected Papers/Litmaps) strongest validity signal, not star Conclusion Prints “no G2/Capterra/TP standalone star, AI Agent Index 3.1 + APIs.io 56.2 dev-sentiment, ecosystem-adopted” — aligns MonkeyLearn-standalone/FB-Ads-API/Gmail-API/Documenso rule (nonprofit/academic infra, not G2-rated bought SaaS); do NOT fabricate 4.x Access
Field Value Web App semanticscholar.org — search, paper page, author page, library, alerts; no desktop app API REST v1 (graph/recommendations/datasets), Bearer API key optional, fields param, batch endpoints (paper/batch, author/batch), pagination next/cursor; no GraphQL, no gRPC, no SOAP MCP None first-party — community semantic-scholar-mcp(smaniches) wraps REST, Claude/Cursor mount via npx, env S2_API_KEY, 14 tools, upstream still 1 RPS; not AI2-supported, not in vendor MCP clusterIntegrations Zotero/Mendeley manual export only (no native), BibTeX/CSV/RIS export, OpenAthens/eduGAIN institutional login, GetFTR/LibKey full-text bridge; downstream Elicit/Consensus pull S2AG Login None for search; Google/Twitter/Facebook/email/institutional (OpenAthens) for library/alerts Self-host S2AG dataset downloadable to self-host graph; API itself Cloud-only, no on-prem endpoint Best For
Field Value Audience 1 Academic researchers / grad students doing literature discovery needing TLDR + citation-context without paywall Audience 2 Devs/data scientists building research tools (Elicit-style) on top of open 236M-corpus + SPECTER2 embeddings Audience 3 Systematic-review teams needing bulk S2AG snapshot locally (no per-seat, no WoS license) Audience 4 Agent builders willing to mount community MCP bridge for paper_search + snippet_search (rate-aware) Audience 5 CS/biomed NLP teams needing citation graphs / influence-weighted signals for training data Not For
Field Value Exclude 1 Teams needing conversational RAG answers (“summarize the field”) — Semantic Scholar searches/ranks, does not synthesize (use Elicit/Consensus/SciSpace on top) Exclude 2 Orgs requiring 50-100 RPS agentic crawl out-of-box — 1 RPS key intro, must apply academic-partner review (weeks) Exclude 3 Procurement needing G2/Capterra stars + SOC 2 + BAA + SLA — AI2 is nonprofit academic infra, none of those commercial artifacts Exclude 4 Fields outside CS/biomed/STM (SSH coverage thinner), or needing Scopus/WoS citation merge for tenure metrics Exclude 5 Buyers expecting first-party MCP with write/notify/server-push — none exists, community bridge only Exclude 6 Reference-manager replacement (no Zotero-sync, export-only) Competitors
Field Value Google Scholar Free, broadest index, no AI TLDR/citation-context/SPECTER2, no structured API (scrape-only), no bulk dataset PubMed / PubMed Central Biomedical gold standard, free, MeSH-tagged, but CS/SSH thin, no TLDR/embeddings, E-utilities API older Scopus / Web of Science Paywalled institutional, strongest citation completeness + field-normalized metrics, no AI enrichment, no free API Elicit / Consensus / SciSpace Build ON Semantic Scholar (S2AG) + add RAG synthesis/agents; SciSpace freemium 49/mo, Consensus $12/mo — complementary not pure competitors Connected Papers / ResearchRabbit / Litmaps S2AG-powered viz/discovery layers, narrower than S2 core OpenAlex Nonprofit open graph (237M works), REST+graphQL, more generous rate (10K/day), weaker AI TLDR/citation-context than S2, direct closest infra twin
