Library / SDK
scverse/scvi-tools avatar
scverse/scvi-tools

scvi-tools: probabilistic single-cell analysis with PyTorch and AnnData

Deep probabilistic analysis of single-cell and spatial omics data

1,698 stars476 forksPythonBSD-3-Clause

At a glance

What is it?
scvi-tools packages deep generative models for single-cell and spatial omics behind a Scanpy-compatible API. It is a research library, not a pipeline runner, and the README leaves most model-level detail to the user guide.
Who is it for?
Adopt scvi-tools if your work is single-cell or spatial omics and you want a generative model per task rather than a fixed pipeline: dimensionality reduction, integration, annotation, factor analysis, doublet detection and spatial deconvolution all sit behind the same AnnData-oriented API. Do not adopt it if you need a stable API across minor releases, a turnkey pipeline, or support for a Python older than 3.12.
Can I use it commercially?
Yes. BSD-3-Clause is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 1 day ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What scvi-tools solves, and for whom

Single-cell assays produce count matrices with technical noise that is not uniform across cells or batches. A batch effect, a difference in sequencing depth, or a doublet can look like biology if you analyse the raw matrix directly. scvi-tools addresses this by fitting a probabilistic model of the counts instead of transforming them with a fixed formula, then exposing the fitted model through a high-level API that reads and writes AnnData objects.

The README lists the tasks the bundled models cover: dimensionality reduction, data integration, automated annotation, factor analysis, doublet detection and spatial deconvolution. That list defines the audience. This is a library for computational biologists and method developers who already work in Python with Scanpy and AnnData, and who want a model they can train, save, reload and inspect. It is not aimed at bench scientists who want a web interface, and it is not a workflow manager.

The second audience is smaller but explicitly served. The README describes scvi-tools as containing "the building blocks to develop and deploy novel probabilistic models", powered by PyTorch Lightning and Pyro, and points to a skeleton repository as a starting point. So the package is both a collection of ready models and a framework for writing new ones, and those two roles pull the design in different directions.

How a scvi-tools model actually runs

The data flow starts with AnnData. Every model implementation is expected to accept an AnnData object, register the fields it needs (which layer holds counts, which column holds the batch label), and write its outputs back into the same object. That is why the README can say the API "interacts with Scanpy": the result of a fit is not a separate object graph you have to learn, it is slots on the object you already have.

Underneath, the model is a variational autoencoder trained with stochastic variational inference. The dependencies make the stack concrete: torch for the network and autograd, pyro-ppl for the probabilistic programming primitives, lightning for the training loop, torchmetrics for tracked quantities, and tensorboard for logging. Anndata and mudata handle the container, scanpy the surrounding single-cell operations, numba and sparse the numeric paths. Nothing in that list is optional in the base install.

The consequence is that a model is a trained artefact, not a function call. You configure it, you train it, you save it, you load it later. The README states that standard save and load functions and GPU acceleration are part of the high-level API. It does not describe how a saved model behaves when the AnnData schema changes, and the README does not document rollback or migration of saved models between versions. Treat a saved model as tied to the version that produced it.

Installing scvi-tools from conda-forge or PyPI

The README gives two install paths and no others. Conda users install from conda-forge:

bash
conda install scvi-tools -c conda-forge

pip users install from PyPI:

bash
pip install scvi-tools

The README then adds a warning that matters more than either command: "Please be sure to install a version of PyTorch that is compatible with your GPU (if applicable)." scvi-tools depends on torch but does not pin a CUDA build, so a plain pip install can leave you with a CPU-only wheel on a machine that has a GPU. Check the torch install before you file a bug about training speed.

One constraint is easy to miss. The pyproject.toml sets requires-python to ">=3.12" and classifies the package for Python 3.12, 3.13 and 3.14 only. If your environment is on 3.11, pip will refuse to resolve scvi-tools rather than install an older release.

For a first real use, the pattern the documentation describes is: load an AnnData object, register the model's required fields, instantiate, train, then query the fitted model. The exact model class and the exact field names differ per model and per dataset, so the user guide is the place to get them; inventing a registration call here would be wrong. What is stable across models is the shape of the sequence, and the fact that the object you pass in is the object you read results from afterwards.

Where scvi-tools is the wrong tool

The package classifies itself as "Development Status :: 4 - Beta" in pyproject.toml. That is the project's own label, and it should set expectations. If you are building a service that pins scvi-tools and expects the same call signature to survive a minor upgrade, the beta classifier and the release history are both telling you to pin exactly and read the CHANGELOG before moving.

The second limitation is platform-specific and stated in the build file. The autotune extra carries the comment "TODO: New autotune not supported on Windows yet", and its ray dependency is marked for non-Windows platforms only. Hyperparameter tuning is therefore a Linux and macOS story in this version. The base library still installs on Windows, but the tuning path does not.

The third is conceptual. A probabilistic model gives you a posterior, and a posterior has to be interpreted. If your analysis is a fixed sequence of normalisation, clustering and marker detection, and you do not want to reason about latent variables or training convergence, scvi-tools adds a layer you will not use. Scanpy alone does that work with fewer moving parts. The models here earn their cost when batch structure, count noise or spatial mixing genuinely distort the answer.

scvi-tools and Scanpy: what the split actually is

The obvious alternative is Scanpy on its own. The difference is not feature coverage, since scvi-tools depends on scanpy and calls into it. The difference is what each one treats as the object of analysis.

Scanpy operates on the matrix. Normalisation, highly variable gene selection, PCA, neighbourhood graphs and clustering are deterministic transformations of the values you give it. The result is reproducible from the input and the parameters, and there is no training step to converge or fail.

scvi-tools operates on a generative model of the matrix. The model has parameters that are estimated from the data, which means the output depends on the training run: initialisation, number of epochs, and whether optimisation converged. In exchange, the model can represent batch effects and count noise explicitly, and can produce quantities a fixed transformation cannot, such as a posterior over latent variables or a likelihood-based comparison between conditions. The README's task list (integration, automated annotation, doublet detection, spatial deconvolution) is exactly the set of problems where that trade is worth making.

A practical reading: use Scanpy for the parts of the analysis that are transformations, and reach for a scvi-tools model for the parts that are inference. The two share the AnnData container, so that split costs you nothing in plumbing.

Maintenance, releases and the BSD-3-Clause licence

The repository is not archived, and the last push was on 2026-09-10, the same day as the 1.5.1 release. Before that, 1.5.0 and 1.5.0.post1 landed on 2026-07-09. That is a steady cadence across the recent window, and the post-release suggests fixes are shipped rather than held. The upgrade cost that follows is real: a post release means the 1.5.0 tag is not the one you want, and pinning "1.5.0" in a requirements file gets you the earlier build.

Governance is documented in the README rather than implied. scvi-tools is part of the scverse project, which has a published governance page and is fiscally sponsored by NumFOCUS, with a donation link for developer time. For a lab deciding whether to build on it, that is a more useful signal than any download badge: there is an organisation with a stated role structure behind the package.

scvi-tools is BSD-3-Clause, and the README asks that you cite the Nature Biotechnology 2022 paper for the library plus the publication for the specific model you used, and the scverse paper for the ecosystem. The licence and the citation request are separate obligations. Permissive licensing does not remove the academic norm of citing the model, and the README is explicit that the model paper is a second citation, not an optional one. This is not legal advice; check how your institution handles both.

Editorial conclusion

Adopt scvi-tools if your work is single-cell or spatial omics and you want a generative model per task rather than a fixed pipeline: dimensionality reduction, integration, annotation, factor analysis, doublet detection and spatial deconvolution all sit behind the same AnnData-oriented API. Do not adopt it if you need a stable API across minor releases, a turnkey pipeline, or support for a Python older than 3.12. Before committing, check that PyTorch resolves against your GPU driver, confirm the model you intend to use is named in the user guide, and read the CHANGELOG entry for the version you pin.

Frequently asked questions

What is scvi-tools?

scvi-tools (single-cell variational inference tools) is a Python package for probabilistic modeling and analysis of single-cell omics data, built on PyTorch and AnnData. It bundles models for dimensionality reduction, data integration, automated annotation, factor analysis, doublet detection and spatial deconvolution, and also provides building blocks for developing new probabilistic models.

How does scVI integration work?

The README does not describe the internals of integration. It states that the models perform data integration among other tasks, that each model has a high-level API interacting with Scanpy, and that the implementations are probabilistic models built on PyTorch, Pyro and PyTorch Lightning. The user guide is where the model-level explanation lives.

What are the applications of scVI?

According to the README, the models cover dimensionality reduction, data integration, automated annotation, factor analysis, doublet detection and spatial deconvolution across single-cell, multi and spatial omics data. The README says the user guide provides an overview of each model.

How to install scvi-tools?

The README gives two commands: conda install scvi-tools -c conda-forge, or pip install scvi-tools. It adds that you should install a version of PyTorch compatible with your GPU if you have one. The package requires Python 3.12 or newer.

How does scVI work?

The README describes scvi-tools as a package for probabilistic modeling built on PyTorch and AnnData, whose model implementations are powered by PyTorch Lightning and Pyro and expose a high-level API that interacts with Scanpy. It does not lay out the model equations; the user guide and the cited publications are the sources for those.

Official sources

  1. License: BSD-3-Clause
  2. Project website
  3. README
  4. Releases
  5. scverse/scvi-tools on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/scverse-scvi-tools.svg)](https://hysenlabs.com/projects/scverse-scvi-tools)