# paperai: bulk LLM extraction over medical and scientific paper collections

> paperai is a Python application that runs configured LLM prompts across a corpus of medical and scientific papers and writes the answers into Markdown, CSV or annotated PDFs. It is a batch extraction tool, not a chat interface, and it depends on paperetl to prepare the data it reads.

**neuml/paperai** — 📄 🤖 AI for medical and scientific papers

- Repository: https://github.com/neuml/paperai
- Stars: 1,782 · Forks: 148
- Language: Python
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/neuml-paperai

## The problem paperai solves: turning a paper corpus into a filled-in table

Reading a few hundred papers to pull out one field from each is mechanical work. Sample size, study objective, journal, date: the same slots, over and over. paperai is built for that shape of task. The README describes it as an application that "goes through repositories of articles and generates bulk answers to questions backed by Large Language Model (LLM) prompts and Retrieval Augmented Generation (RAG) pipelines." The framing is deliberate. It is not a chatbot over your PDFs, and it is not a search engine. It is a configuration-driven batch job.

The intended user is someone with a corpus already in a database, a list of fields they want per paper, and enough local compute to run an LLM. The repository ships two example notebooks, one of which is a medical research project on young onset colon cancer, and an example report configuration named crc.yml. Those examples show the expected workflow more clearly than the README prose does.

If your task is "find me papers about X", paperai is the wrong layer. If your task is "for each of these 400 papers, tell me the sample size and the stated objective", it is aimed squarely at you.

## How the indexing and report pipeline actually fits together

There are two separable stages, and the split matters.

Indexing is handled by a module invoked as python -m paperai.index. The README states that paperai "indexes databases previously built with paperetl". So the input is not a folder of PDFs; it is a paperetl database. paperetl is a separate repository by the same author, and the README links to its Docker instructions rather than describing its schema. That is the first real constraint you hit: paperai does not ingest PDFs on its own.

The index itself is a txtai embeddings instance. The README says paperai "uses the default txtai embeddings configuration when not specified", and that an index.yml file "takes all the same options as a txtai embeddings instance". A minimal example given in the README sets path to sentence-transformers/all-MiniLM-L6-v2 and content to True. Because the configuration surface is txtai's, anything txtai supports for embeddings configuration is in scope, and paperai's own documentation does not restate it.

Querying happens against that index. The report configuration is a YAML file with a top-level name, an options block, and one or more named report sections. Each report section carries a query and a list of columns. A column can be a plain name, or a mapping with a name, a query and a question. The question is what gets substituted into the template; the query is what retrieves context from the index. The options block holds the llm path, a system prompt, a template with {question} and {context} placeholders, a context count, and params such as maxlength and stripthink.

That is the whole mechanism: retrieve context per column per paper, substitute into the template, run the LLM, collect the field value. Output goes to Markdown, CSV, or annotations on the PDF when the PDF is available.

## Installing paperai and running a first extraction

Installation is a single pip command. Python 3.10 or later is required, and the README recommends a virtual environment.

```bash
pip install paperai
```

There is also a from-source install for unreleased features:

```bash
pip install git+https://github.com/neuml/paperai
```

A Docker path exists too. The README gives these three commands, which download the project's Dockerfile, build an image tagged paperai, and start a container.

```bash
wget https://raw.githubusercontent.com/neuml/paperai/master/docker/Dockerfile
docker build -t paperai .
docker run --name paperai --rm -it paperai
```

The README notes that paperetl can be folded into the same image by building a paperetl image first and then passing BASE_IMAGE=paperetl and START=/scripts/start.sh as build arguments. That gives one image that both indexes and queries.

With the package installed, the first real step is building an index over a paperetl database. The command takes an input data path and an optional index configuration, where the configuration is either a vector model path or an index.yml file.

```bash
python -m paperai.index <path to input data> <optional index configuration>
```

If you want to pin the embedding model, write an index.yml first. The README's example is exactly this:

```yaml
path: sentence-transformers/all-MiniLM-L6-v2
content: True
```

Once the index exists, the fastest way to run queries is the bundled shell. Passing the model directory starts an interactive prompt where queries are typed directly into the console.

```bash
paperai <path to model directory>
```

For bulk work you do not use the shell. You write a report configuration. The README's example begins like this, with the LLM as a GGUF file path, a system prompt, a template containing {question} and {context}, a context count of 5, and params setting maxlength to 4096 and stripthink to True:

```yaml
name: ColonCancer
options:
    llm: Intelligent-Internet/II-Medical-8B-1706-GGUF/II-Medical-8B-1706.Q4_K_M.gguf
    system: You are a medical literature document parser. You extract fields from data.
    context: 5
    params:
        maxlength: 4096
        stripthink: True
```

The README's template text is also worth reading in full before you write your own, because it encodes the authors' assumptions about how the model should behave: keep it simple, extract only the data, never explain, and say "no data" when the field cannot be found in the context. Those instructions exist because the failure mode is a model that editorializes instead of filling a cell.

## Where paperai breaks down

The dependency on paperetl is the sharpest edge. paperai does not parse PDFs, so a corpus that paperetl cannot ingest is out of reach regardless of what paperai supports. The README does not document paperetl's accepted input formats, so that check has to happen in the other repository before you plan anything.

Cost scales multiplicatively. Every column in every report section is a separate question, and each one retrieves its own context and runs its own inference. A report with six columns over 400 papers is roughly 2,400 LLM calls. The README's own framing, "kicking off hundreds of ChatGPT prompts over your data", is accurate and should be read as a warning about runtime as much as a description of capability.

Extraction quality is not guaranteed by the tool. The template in the README instructs the model to output "no data" when a field is absent, which tells you the authors expect empty cells and want them marked rather than hallucinated. Whether your model obeys depends on the model. The configuration lets you point llm at any GGUF path, and the README offers no evaluation harness, no accuracy figure and no validation step. You are responsible for spot-checking the output.

The README also does not document rollback, incremental re-indexing, or how to update an existing index when new papers arrive. If your corpus grows continuously, that gap matters, and the documentation is silent on it.

Finally, the shell is a query console, not a search product. The repository's examples/search.py is described as an application to "search a paperai index", which suggests the shell is the intended interactive surface and anything richer is something you build.

## paperai compared with a general RAG framework

The obvious alternative is assembling the same pipeline from a general retrieval framework plus an LLM client. LangChain and LlamaIndex both cover retrieval and prompt orchestration, and both are far broader than paperai.

The difference in approach is the unit of work. A general framework gives you chains and retrievers as composable primitives, and you write the loop that walks your documents. paperai inverts that: the loop is fixed, and what you supply is a YAML file describing columns and questions. The retrieval layer underneath is txtai, which paperai depends on directly (txtai[api]>=8.5.0 in setup.py), so you are not getting a novel retriever. You are getting a batch runner with a report schema and three output formats.

That is a real trade. If your extraction job is genuinely tabular, with the same fields across many documents, paperai's configuration is less code than a hand-rolled chain. If your job needs branching logic, per-document routing, or a custom post-processing step, the fixed loop becomes the obstacle, and a general framework is the better fit. There is also txtai itself to consider: since paperai is a thin application on top of it, anyone already comfortable with txtai embeddings can build the retrieval half directly and keep only the parts of paperai they want.

## Maintenance, licensing and the cost of staying current

The repository is not archived and the last push was on 2026-07-14, two months before this writing. Releases have been uneven: v2.3.0 in December 2024, v2.4.0 and v2.5.0 in June and July 2025, and setup.py in the repository declares version 2.6.0, which is ahead of the newest listed release. That pattern is worth knowing if you pin versions.

The dependency list is the real maintenance surface. paperai pulls in txtai[api]>=8.5.0, staticvectors[train]>=0.2.0, skops>=0.9.0, txtmarker>=1.0.0, scikit-learn, networkx, PyYAML, regex, rich and text2digits. Several of these are lower-bound-only constraints, which means a fresh install can resolve to newer versions than the ones the project was tested against. The Makefile's test target runs unittest discovery over the test directory, and there is a separate coverage target, but the README does not publish a compatibility matrix for the dependencies.

Licensing is Apache-2.0, stated in setup.py and in the LICENSE file at the repository root. That is a permissive licence with an explicit patent grant, which matters for commercial deployment inside a company. It covers paperai's own code. It does not cover the models you point it at: the README's example uses Intelligent-Internet/II-Medical-8B-1706-GGUF, and the licence of any GGUF model you configure is a separate question that the paperai documentation does not address. Check the model card before shipping anything. This is not legal advice; the licence text is the authority.

## Conclusion

Adopt paperai if you already have a paperetl database and a defined set of fields to extract across hundreds of papers, and you are willing to tune the prompt template yourself. Do not adopt it if you want a search UI, a hosted service or a one-off question answered against a single PDF. Before committing, verify that paperetl can ingest your source format, that your machine can run the LLM you intend to configure, and that the report schema in your YAML actually produces the columns you need.

## FAQ

### What is paperai used for?

It runs configured LLM prompts over a corpus of medical and scientific papers and collects the answers as report fields. The README describes it as generating bulk answers to questions backed by LLM prompts and RAG pipelines, with output in Markdown, CSV or annotations on PDFs.

### How do I install paperai?

Install it from PyPI with pip install paperai, which requires Python 3.10 or later. The README also documents a GitHub install for unreleased features and a Docker build from the project's Dockerfile.

### Does paperai index PDFs directly?

No. The README states that paperai indexes databases previously built with paperetl, so PDF ingestion happens in paperetl rather than in paperai. The README links to paperetl's Docker instructions but does not describe its accepted input formats.

## Sources

- [Issues](https://github.com/neuml/paperai/issues)
- [License: Apache-2.0](https://github.com/neuml/paperai/blob/master/LICENSE)
- [neuml/paperai on GitHub](https://github.com/neuml/paperai)
- [README](https://github.com/neuml/paperai/blob/master/README.md)
- [Releases](https://github.com/neuml/paperai/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/neuml-paperai
