# MetaScreener sends one paper to four models and needs them all to agree

> An Apache-2.0 Python tool that automates the screening stage of a systematic review by asking several open-source models in parallel and calibrating their scores. The consensus logic is strict, the version in the tree is ahead of every published tag, and the cost tables are per paper rather than per review.

**ChaokunHong/MetaScreener** — AI-powered tool for efficient abstract and PDF screening in systematic reviews.

- Repository: https://github.com/ChaokunHong/MetaScreener
- Website: https://www.metascreener.net/
- Stars: 1,334 · Forks: 50
- Language: Python
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/chaokunhong-metascreener

## The version in the tree is ahead of every published tag

The packaging metadata declares one version:

```
version = "2.0.0a5"
```

The release history tells a different story. The most recent numbered tags are v2.0.0a4 and v2.0.0a3, both from February 2026, and the newest tag by date is not a version of the software at all: it is named for an audit protocol lock and is dated May 2026. So the code in the repository is one alpha ahead of anything published, and the most recent thing to happen in the release feed was not a release.

The classifier agrees with the suffix rather than the maturity. The package is marked Development Status :: 3 - Alpha and declares support for Python 3.11 and 3.12, while the interpreter floor is stated more loosely as 3.11 or newer, which leaves later interpreters permitted but undeclared. The last push to the repository is dated 11 June 2026.

None of this makes the tool unusable. It does mean that installing from the branch gives you a different build from the one behind the pip release, and that for a systematic review, where the screening record may be audited years later, the distinction is worth more than usual.

## Tier 1 auto-decisions require every model to agree

The decision router sorts papers into four tiers, and the first two are where the labour lands. Tier 0 is unambiguous:

```
| **0** | Hard-rule violation | Auto-exclude |
```

That tier is handled entirely by the rule engine, which auto-excludes anything violating a non-negotiable criterion such as a study in the wrong language, or an animal study where human evidence is required. No model is consulted, and a screening auditor can reproduce the exclusion from the rule alone.

Tier 1 is the opposite. It requires a high Element Consensus Score of at least 0.60 and full model agreement. Both conditions, not either one. A paper where the ensemble scores 0.95 on consensus but one model dissents does not qualify, and it goes down the review path instead.

The design is defensible for high-stakes screening, where a wrong inclusion costs more downstream than a slow review. It also means throughput is not a property of the tool alone. It depends on how much your chosen models diverge on your corpus, and a heterogeneous preset will send more papers to a human than a homogeneous one even when both are described as accurate.

## Consensus is scored per criterion element, not per paper

Most ensemble screening tools return one verdict per paper. This one asks each model for a per-element assessment of population, intervention, comparison and outcome, and the Element Consensus Score measures agreement element by element rather than on the aggregate.

The stated benefit is precision about where the disagreement is. If every model agrees the population matches but they split on outcome, the record shows that, instead of collapsing into a vague uncertain verdict with no explanation attached. For a reviewer deciding whether to spend ten minutes on a paper, knowing that the failure is confined to one element of PICO is the difference between a quick yes and a full read.

The aggregation layer applies post-hoc calibration to the raw scores, using Platt scaling or isotonic regression, and then blends by tier. That step matters because model-reported confidence is not a probability, and calibration is what turns it into something a threshold can be set against. The cost is that calibration requires labelled data, and the README does not say where the calibration set comes from, how large it needs to be, or whether it ships with the package. That gap is the first thing to close before trusting a 0.60 threshold in a real review.

## Real-time recalibration sits next to a claim of full reproducibility

Two rows of the same feature table pull in opposite directions. One promises a full reproducibility story, and it is specific about how: temperature set to 0.0, a fixed seed of 42, and an audit trail recorded on every decision. The other promises active learning, described as a human feedback loop that recalibrates model weights in real time.

Those are not contradictory once you read them as different kinds of reproducibility. What the first guarantees is that a decision can be replayed and attributed: which model said what, at what confidence, under which criteria. What the second does is change the weights themselves as reviewers override calls. The audit trail records the history of a review correctly, but it cannot make the earlier verdicts the ones the current weights would produce.

The practical consequence is about longitudinal studies rather than single reviews. Re-screening the same corpus next year, with a different set of reviewers and a recalibrated ensemble, will not reproduce the earlier inclusion list, and the export will not tell you which of those two reasons caused the difference. If a review protocol requires that every later change be documented, the override step needs a written policy around it.

## Installing from source needs Node.js, because the frontend is built separately

There are three ways in, and they differ in what they assume is already on the machine.

```
pip install metascreener
```

```
python -m metascreener
```

The first path assumes Python 3.11 or newer and nothing else. The third path assumes two toolchains:

```
git clone https://github.com/ChaokunHong/MetaScreener.git
cd MetaScreener
uv sync --extra dev        # Install Python dependencies
python run.py              # Start FastAPI + Vite dev servers
```

Node.js 18 or newer is listed as a prerequisite for the source route, because the frontend is a separate Vite application and run.py starts both the FastAPI backend on port 8000 and the Vite dev server on 5173. The pip wheel is built by hatchling from src/metascreener, so a wheel consumer never runs the JavaScript build. The lock file is uv.lock, which means the reproducible Python environment is a uv environment, and anyone pinning dependencies for a review should pin it the same way.

The Docker route sidesteps both toolchains and is the shortest to a running service, but it pulls a mutable tag:

## The Docker image is tagged latest, and the GPU extra is undocumented

The container path is three lines:

```
docker pull chaokunhong/metascreener:latest
```

```
docker run -p 8000:8000 \
  -e OPENROUTER_API_KEY="sk-or-v1-your-key-here" \
  chaokunhong/metascreener
```

The tag is latest, so two runs a month apart can execute different images, which is a poor property for a tool whose output feeds a review record. Pinning by digest costs one extra step and removes the question.

The optional dependencies contain a second, quieter surprise. Beyond the dev group there is a GPU group:

```
gpu = ["vllm>=0.6"]
```

That single entry implies a path where models are served locally rather than through an API, which matters for a screening tool because it moves per-paper cost from a paid API to your own hardware. Nothing in the README explains it: no instructions, no model list, no instructions for pointing the ensemble at a local endpoint. The documented route requires an OpenRouter key, so treat the GPU group as an unfinished affordance rather than a supported mode until someone writes the missing page.

Where the key lives is also worth knowing. It can be pasted into the settings page of the web interface, or supplied as an environment variable, and the README does not say where a key entered through the interface is persisted.

## The workflow runs well past screening into extraction and risk of bias

The project describes itself as a tool for abstract and PDF screening, and the description is accurate but narrower than the workflow. The web interface is laid out as eight numbered steps, and only the first three are about deciding whether to keep a paper.

Step 0 generates inclusion criteria from a research question, in PICO, PEO or SPIDER form, or accepts criteria you already have. Step 1 selects models and thresholds. Step 2 is title and abstract screening on an uploaded .ris, .bib, .csv or .xlsx export. Step 3 moves to full text, with what the documentation calls intelligent chunking, which is where the PDF parser earns its place in the dependency list.

Step 4 extracts structured data, tables and fields from the included PDFs. Step 5 runs risk-of-bias assessment against RoB 2, ROBINS-I or QUADAS-2. Step 6 produces performance metrics, calibration diagnostics and visualisations. Step 7 is the session audit trail.

So the honest description is a review assistant that starts at screening and continues into extraction and appraisal, with screening as the part that has been made automatic. That matters for scoping: if you only need the first pass, steps 4 to 6 are surface area you will not touch, and if you need the full pipeline, this is a different proposition from the one-line summary suggests. Exports are available as CSV, Excel, JSON and RIS, and a reviewer can override any decision in the interface.

## Cost is quoted per paper, so the ensemble multiplier is yours to compute

The pricing table gives three presets with an approximate cost per paper:

```
| Balanced | 4 models | ~$0.005 | Most reviews |
| Precision | 2 thinking + 2 large | ~$0.009 | High-stakes reviews |
| Budget | 1 anchor + 3 fast | ~$0.003 | Large-scale screening |
```

Two observations. First, the range in the feature table, roughly three cents to nine tenths of a cent per paper, is per paper and not per review, so it has to be multiplied by the size of your search export before it means anything. Second, all three presets name four models, so the per-paper figure already includes four calls, and the presets differ in which models those are rather than in how many.

That makes the threefold spread between Budget and Precision a statement about model choice, which is the intended trade: a preset with two reasoning-heavy models costs more and disagrees less, and by the tier rules disagreement is what determines your reviewer workload. Choosing between them on price alone is the wrong comparison, because the reviewer hours saved are not in the table.

The figures are described as an estimate on free-tier providers, so real cost depends on your provider plan and on how often the override path sends a paper back through the ensemble.

## Conclusion

MetaScreener is a sensible fit for a review team that already has inclusion criteria written down and wants to shrink the title-and-abstract pass without handing the decision to one model. Three things are worth checking first. The version in pyproject.toml is ahead of every published tag, so pin a release rather than the branch. Tier 1 auto-decisions require every model to agree, which means your human review workload depends on how much the ensemble actually diverges. And the calibration step needs a labelled set, which the project does not document, so do not assume the confidence numbers are trustworthy until you have checked them against your own judgements.

## FAQ

### What does MetaScreener need in order to run?

Python 3.11 or newer and an OpenRouter API key. The key can be pasted into the settings page of the web interface or supplied as the OPENROUTER_API_KEY environment variable. Running from source additionally requires uv and Node.js 18 or newer, because the frontend is a separate Vite application.

### How does MetaScreener turn model output into a decision?

Each paper goes to four or more models in parallel, raw confidence scores are calibrated with Platt scaling or isotonic regression, and blended by tier. A decision router then sorts papers into four tiers, where tier 0 is auto-excluded by a hard rule and tier 1 requires an Element Consensus Score of at least 0.60 plus full model agreement.

### Can I install MetaScreener with Docker?

Yes. The published image is pulled from chaokunhong/metascreener and run with port 8000 published and the OpenRouter key passed as an environment variable. The image is tagged latest, so pinning a digest is advisable if the screening record has to be reproducible later.

### Does MetaScreener do more than screen abstracts?

The web workflow has eight steps: criteria generation, settings, title and abstract screening, full text screening, structured data extraction, risk-of-bias assessment using RoB 2, ROBINS-I or QUADAS-2, evaluation with calibration diagnostics, and a session audit trail. Exports are available as CSV, Excel, JSON and RIS.

### How much does screening with MetaScreener cost?

The presets are quoted at roughly three thousandths of a dollar per paper for Budget, five thousandths for Balanced and nine thousandths for Precision. All three name four models, so the figures are per paper across the whole ensemble rather than per model call, and the estimates assume free-tier providers.

## Sources

- [ChaokunHong/MetaScreener on GitHub](https://github.com/ChaokunHong/MetaScreener)
- [License: Apache-2.0](https://github.com/ChaokunHong/MetaScreener/blob/main/LICENSE)
- [Project website](https://www.metascreener.net/)
- [README](https://github.com/ChaokunHong/MetaScreener/blob/main/README.md)
- [Releases](https://github.com/ChaokunHong/MetaScreener/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/chaokunhong-metascreener
