CLI tool
Future-House/paper-qa avatar
Future-House/paper-qa

PaperQA2: agentic RAG for scientific papers with grounded citations

High accuracy RAG for answering questions from scientific documents with citations

9,211 stars920 forksPythonApache-2.0

At a glance

What is it?
PaperQA2 is a Python package that answers questions from a local folder of PDFs, Office files and source code, attaching in-text citations to each claim. It is built for research workflows where a wrong answer without a source is worse than no answer.
Who is it for?
Adopt PaperQA2 if you already keep a folder of PDFs and need answers where every sentence carries a page-level citation back to the source, and if you accept that the default path sends chunks to OpenAI. Do not adopt it if you need a fully offline stack on day one, or if your corpus is web pages rather than papers.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 13 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 17, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The problem PaperQA2 targets: answers you can trace back to a page

Generic RAG pipelines are tuned for web text. They chunk a document, embed the chunks, retrieve the top matches and let a model write a paragraph. When the corpus is scientific literature, that pipeline breaks in specific ways. A methods section may contradict its own abstract. Two papers may use the same term for different constructs. A claim that appears in one paper may have been retracted. PaperQA2 is aimed at that setting. The README describes it as a package for high-accuracy retrieval augmented generation on PDFs, text files, Microsoft Office documents and source code files, with a focus on the scientific literature, and the example output it publishes attaches citations of the form Qian2011Neural pages 1-2 to individual claims rather than to the answer as a whole. That page-level attribution is the product. It is for researchers, literature reviewers and tool builders who need to check a statement against its source, not just read a plausible summary.

How the retrieval pipeline actually works

The README names the moving parts. Documents are parsed and cached into a full-text search index built on tantivy, a Rust search library, so keyword search runs locally rather than through a hosted service. Embeddings are stored in a Numpy vector database by default, with OpenAI embeddings and OpenAI models as the default providers. Retrieval is not a single pass. The README describes LLM-based re-ranking and contextual summarization, abbreviated RCS, alongside document metadata awareness in the embeddings, and it lists agentic RAG as a supported mode where a language agent iteratively refines queries and answers. Metadata comes from external providers: Semantic Scholar, Crossref and Unpaywall are named as dependencies, and the README states that citation counts are fetched with a retraction check. The CLI, pqa, ties these together. The Quickstart sequence is: point it at a folder of PDFs, let it fetch metadata, parse and cache the PDFs into the index, then ask a question. The LLM provider layer is LiteLLM, which the README says gives default support for all LiteLLM providers, so the model is a configuration choice rather than a hard dependency on one vendor.

Installing PaperQA2 and asking your first question

Installation is a single pip command, and the Quickstart uses the pqa CLI rather than a Python script. The block below creates a papers folder, downloads the PaperQA2 paper from arXiv as a test document, moves into that folder and asks a question. Run it from a directory where you are happy to have an index and a cache created, because the paper, its parsed text and its metadata all get written alongside it.

bash
pip install paper-qa
mkdir my_papers
curl -o my_papers/PaperQA2.pdf https://arxiv.org/pdf/2409.13740
cd my_papers
pqa ask 'What is PaperQA2?'

What you should see is an answer with in-text citations naming the source document and page ranges, in the style of the README's example output. The README also notes a rate-limit concern: it has a section on rate limits, and it warns that the default settings can be expensive or slow, so a first run against a large folder is worth doing with a small number of documents. For library use rather than the CLI, the README documents both an agentic path and a manual path for adding and querying documents, plus an async interface. It also documents specifying an embedding model, using local sentence-transformers embeddings, and reusing an index once it has been created, which is what you want on the second run so PDFs are not re-parsed.

Where the design costs you: cost, parsing and provider lock-in

The default configuration is not cheap. The README states that PaperQA2 uses OpenAI embeddings and OpenAI models by default, and the pipeline runs an LLM over retrieved chunks for re-ranking and summarization, not just once at the end. Every question therefore spends tokens on retrieval-side model calls as well as the final answer, and the README's own rate-limit section exists because that traffic can hit provider limits. The second cost is parsing. PDF text extraction is a known weak point in scientific publishing, and the project splits this out: pyproject.toml lists paper-qa-pypdf as a base dependency, while paper-qa-pymupdf, paper-qa-docling and paper-qa-nemotron appear as separate packages in the dev dependency group. That structure is a signal that the parser is swappable and that the default is not the only option, but it also means the quality of your index depends on a choice the README does not make for you. Third, the tool is the wrong shape for a web corpus. It is built around local files and paper metadata providers; if your sources are HTML pages or internal wikis, the metadata enrichment and the citation format do not map cleanly onto your content.

PaperQA2 compared with LlamaIndex and LangChain

The README addresses this directly in its FAQ, asking how PaperQA2 differs from LlamaIndex or LangChain. The distinction is scope. LlamaIndex and LangChain are general orchestration frameworks: they give you connectors, index abstractions and chains, and you assemble the retrieval logic yourself. PaperQA2 ships an opinionated pipeline for one document class, scientific papers, with the metadata providers, the re-ranking and summarization step, and the citation format already wired together. The trade-off runs in both directions. You get a working citation-producing pipeline without designing retrieval, and you give up the breadth of integrations and the ability to swap in arbitrary components that a general framework offers. If your problem is 'answer questions over papers with sources', the opinionated path saves real work. If your problem is 'build a retrieval system over mixed content that happens to include some PDFs', a general framework is the better starting point and PaperQA2 is a component you might call rather than a base you build on.

Maintenance, versioning and the Apache-2.0 licence

The repository is not archived, and the last push was on 2026-09-07, which is recent. Releases have moved to a calendar scheme: v2026.08.12, v2026.03.18 and v2026.03.12 are the three most recent, and the README has a section explaining that PaperQA2 went CalVer in December 2025 after previously following SemVer with major version bumps on breaking changes. For anyone pinning versions, that matters: a date-based version tells you when the release was cut, not how much changed. The project is Apache-2.0, which permits commercial use and modification with the usual attribution and notice requirements; the repository ships a LICENSE file and pyproject.toml declares the license by file reference. Note that the licence covers this code, not the models or the external metadata services it calls. Semantic Scholar, Crossref and Unpaywall each have their own terms, and the LLM provider you configure has its own pricing and data-handling policy. That is a configuration question for your own deployment, not a legal question this article can settle. The upgrade cost is mostly in the settings surface: the README keeps a settings cheatsheet, and the version history shows the project has been willing to change defaults across major versions.

Editorial conclusion

Adopt PaperQA2 if you already keep a folder of PDFs and need answers where every sentence carries a page-level citation back to the source, and if you accept that the default path sends chunks to OpenAI. Do not adopt it if you need a fully offline stack on day one, or if your corpus is web pages rather than papers. Before committing, run pqa ask against a handful of documents you already know well, check whether the cited pages actually support the claims, and confirm which model and embedding provider your settings resolve to.

Frequently asked questions

What is the difference between an article and a paper?

The README does not define the distinction. It describes PaperQA2 as a package for retrieval augmented generation on PDFs, text files, Microsoft Office documents and source code files, with a focus on the scientific literature, and its example citations reference named papers with page ranges.

What is a research paper called?

The README does not give a terminology definition. It refers to its sources as papers and scientific papers, and the metadata it fetches comes from Semantic Scholar, Crossref and Unpaywall.

What does paper mean in science?

The README does not define the word. It treats papers as the input documents: the Quickstart downloads a PDF from arXiv into a folder and runs pqa ask over it.

How do you explain a research paper?

The README does not cover explaining papers to a reader. The closest documented behaviour is that the agent answers a question over a folder of papers and cites the pages that support each claim.

Official sources

  1. Future-House/paper-qa on GitHub
  2. License: Apache-2.0
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/future-house-paper-qa.svg)](https://hysenlabs.com/projects/future-house-paper-qa)