Open-source project
scverse/scanpy avatar
scverse/scanpy

Scanpy: single-cell analysis in Python, and where it stops being the right tool

Single-cell analysis in Python. Scales to >100M cells.

2,576 stars784 forksPythonBSD-3-Clause

At a glance

What is it?
Scanpy is the scverse toolkit for preprocessing, clustering and visualising single-cell expression data on top of anndata. It is built for Python pipelines that need to scale past a million cells, and its documentation is explicit about which parts are still experimental.
Who is it for?
Adopt Scanpy if your single-cell work already lives in Python and you want preprocessing, clustering, UMAP and differential expression in one anndata-backed pipeline. Do not adopt it expecting a graphical interface or a stable internal API: the README states that internal APIs are not guaranteed, so anything imported from an undocumented module can move between releases.
Can I use it commercially?
Yes. BSD-3-Clause is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository received new commits within the last day.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The problem Scanpy solves for single-cell Python users

Single-cell expression experiments produce a matrix of cells against genes, and the analysis that follows is a fixed sequence: filter, normalise, select variable genes, reduce dimensions, cluster, then test genes for differences between groups. Scanpy packages that sequence into one library, and it is built jointly with anndata, which supplies the annotated data container that all the steps read from and write to. The README describes the scope as preprocessing, visualization, clustering, trajectory inference and differential expression testing. That list is the actual product boundary. If your work is single-cell gene expression and you want it in Python, this is the toolkit the scverse project maintains for it. The audience is computational biologists and data scientists who are comfortable writing Python, not analysts looking for a point-and-click application. The README points contributors at a contribution guide and users at a Discourse forum, which tells you where the project expects questions to land.

How Scanpy works: anndata as the single data structure

The architecture is unusual in how little of it is Scanpy. The heavy lifting of storage sits in anndata, a dependency pinned in pyproject.toml as anndata>=0.12.14. Scanpy functions take an AnnData object, modify it in place or return a new one, and the results of each step accumulate as annotations on that object rather than as separate arrays you have to track. Clustering results, embeddings and marker statistics all end up attached to the same container, which is why a Scanpy script reads as a chain of calls on one variable. The scaling claim in the README is that the Python implementation efficiently deals with datasets of more than one million cells. Above that, the README states that many Scanpy functions are now compatible with dask, and immediately flags this as experimental. That warning is the most important sentence in the README for anyone with a large dataset. Experimental means the interface can change and the coverage is partial, so a dask-backed pipeline is a bet on a moving target rather than a supported configuration. Scanpy does not implement its own numerical kernels for everything: it depends on the wider scientific Python stack, which is visible in the dependency list and in the fact that a Scanpy install pulls in a substantial set of packages.

Installing Scanpy and running a first analysis

The README links to the documentation and to PyPI and conda-forge badges, so both distribution channels exist. The package name is scanpy on both. A pip install is the shortest path:

bash
pip install scanpy

After that, the import name is scanpy. The repository requires Python 3.12 or newer, so an older interpreter will fail at install time rather than at import time. The README lists preprocessing, visualization, clustering, trajectory inference and differential expression testing as the covered steps, and the API section of the documentation is where the individual function names live. The README does not print a worked example, so the sequence to follow is the one the API section documents for those five steps. Functions are grouped by prefix: the pp namespace holds preprocessing, tl holds tools that compute results, and pl holds plotting functions. If you prefer conda, scanpy is published on conda-forge, which is the channel the README badge points at. For a notebook environment, the same pip command works inside a Jupyter session, and the pyproject classifiers list Framework :: Jupyter, so notebook use is an intended deployment rather than an afterthought.

The public API boundary is the real constraint

Most libraries discourage internal imports. Scanpy does something more specific and more useful: the README states plainly that the project cannot guarantee the stability of its internal APIs, whether that is the location of a function, its arguments, or something else. It gives a concrete example, from scanpy.logging import debug, and notes that this module has no leading underscore and still is not supported. That distinction matters because Python convention would lead you to assume an underscore-free module is fair game. Here it is not. The practical consequence is that a pipeline built on documented functions survives upgrades, while a pipeline built on convenience imports from submodules can break on a minor release. The README also offers a route out: if something you need is missing from the public API, open an issue. That is a real invitation, not a formality, and it is the correct move when a helper you rely on is undocumented. The trade-off is that Scanpy's stability promise is narrower than its import surface, and you have to know which side of the line each import sits on.

Where Scanpy is the wrong choice

Scanpy is a library, not an application. There is no graphical interface, no project file, and no saved session you can reopen and click through. If your analysis needs to be reproducible by someone who does not write code, Scanpy alone will not get you there. The second boundary is scale. The README claims efficient handling of more than one million cells, and then describes dask compatibility for datasets too large to fit into memory as experimental. Anyone whose data is comfortably above that mark is working with a feature the project itself labels as not settled. The third boundary is scope. Scanpy covers single-cell gene expression. Spatial transcriptomics is a different problem with a different data model, and the scverse ecosystem addresses it in a separate package rather than inside Scanpy. If your question is about tissue coordinates rather than cell-by-gene matrices, look at that package instead of trying to bend Scanpy to it. Finally, the internal API position cuts both ways: it keeps the maintainers free to refactor, and it means you cannot treat any undocumented function as a dependency without accepting churn.

Scanpy compared with Seurat

The most common comparison is with Seurat, the R toolkit for the same class of analysis. The difference is not which one computes a better UMAP. It is the surrounding ecosystem. Seurat lives in R, so adopting it means an R toolchain, R package management and R plotting conventions. Scanpy lives in Python, and its results sit in anndata objects that interoperate with the rest of the Python data stack, which is why the README can point at dask for out-of-core work without leaving the language. If your team already writes Python, Scanpy removes a language boundary. If your collaborators, lab protocols and existing scripts are in R, Seurat keeps you inside one environment instead of two. Neither choice is reversible at zero cost, because the data containers differ and moving between them is a translation step. The honest framing is that this is a toolchain decision first and an algorithm decision second. Scanpy's own documentation does not argue that it is better than Seurat, and anyone claiming a clear winner is arguing about something other than the libraries.

Maintenance, releases and what the licence permits

The repository is not archived. The most recent push recorded is 2026-09-10, and the release history shows 1.12.4 on 2026-08-27 alongside 1.13.0a2 on 2026-08-28 and 1.13.0a1 on 2026-07-24. The pattern is a stable line and a parallel alpha line, which means you can track a release series without adopting pre-release code. Pin to the stable version in production and treat the alpha tags as something to test against, not to deploy. Scanpy is licensed BSD-3-Clause, stated in pyproject.toml as BSD-3-clause and in the classifiers as an OSI-approved BSD licence. A permissive licence of this kind generally allows use, modification and redistribution with the licence text retained, but the terms that apply to your distribution are the ones in the LICENSE file at the repository root, and that is a question for your own legal review rather than something this article can settle. The project is part of scverse and is fiscally sponsored by NumFOCUS, which the README mentions in the context of donations. That structure matters for longevity: funding for developer time runs through a foundation rather than a single lab. The upgrade cost is mostly the API boundary described above. If you stay on documented functions, a version bump is usually a dependency resolution exercise; if you imported from submodules, read the release notes before upgrading.

Editorial conclusion

Adopt Scanpy if your single-cell work already lives in Python and you want preprocessing, clustering, UMAP and differential expression in one anndata-backed pipeline. Do not adopt it expecting a graphical interface or a stable internal API: the README states that internal APIs are not guaranteed, so anything imported from an undocumented module can move between releases. Before committing, verify three things against your own data: that your Python version satisfies the requires-python constraint of 3.12 or newer, that the dask-backed path is acceptable for your dataset given the README labels it experimental, and that the functions you plan to call appear in the published API section rather than only in the source tree.

Frequently asked questions

How do I install Scanpy?

Install it with pip install scanpy, or from the conda-forge channel, which the README links via a badge. The package requires Python 3.12 or newer, so check your interpreter version first.

How do I install Scanpy in a Jupyter notebook?

The same pip install scanpy command works inside a notebook session. The project classifiers list Framework :: Jupyter, so notebook use is an intended environment.

What is Scanpy used for?

The README describes it as a scalable toolkit for analyzing single-cell gene expression data, covering preprocessing, visualization, clustering, trajectory inference and differential expression testing. It is built jointly with anndata, which holds the data.

Is Scanpy better than Seurat?

Nothing in the project documentation supports a quality comparison. The practical difference is the ecosystem: Seurat is R, Scanpy is Python and stores results in anndata objects, so the choice usually follows the language your team and collaborators already use.

How can I remove doublets in Scanpy?

The README lists preprocessing among the covered steps but does not document a doublet-removal function. Check the API section of the documentation for the current function list.

Can Seurat be used in Python?

The project documentation does not describe running Seurat from Python. Scanpy is the Python-side toolkit for the same class of single-cell analysis, and it stores results in anndata objects rather than Seurat objects.

Official sources

  1. License: BSD-3-Clause
  2. Project website
  3. README
  4. Releases
  5. scverse/scanpy on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/scverse-scanpy.svg)](https://hysenlabs.com/projects/scverse-scanpy)