Open-source project
theislab/scarches avatar
theislab/scarches

scArches: mapping new single-cell data onto an existing reference atlas

Reference mapping for single-cell genomics

411 stars73 forksJupyter NotebookBSD-3-Clause

At a glance

What is it?
scArches wraps scvi-tools models in a surgery step so query single-cell data can be projected into a reference atlas instead of integrated from scratch. The package is a Python library for computational biologists who already work with AnnData objects, and its documentation, not its README, is where the actual usage lives.
Who is it for?
Adopt scArches if you already run scvi-tools models and need to project new query samples into a fixed reference without retraining the reference from scratch.
Can I use it commercially?
Yes. BSD-3-Clause is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 96 days ago.
What is it written in?
Mainly Jupyter Notebook, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The problem scArches targets: query data that has to fit an atlas that already exists

Integrating single-cell datasets usually means running a joint embedding over everything at once. That works when you control all the data. It stops working when a reference atlas is fixed, published, and annotated, and a new query sample arrives later. Re-running the joint integration changes the reference embedding, which invalidates the annotations attached to it.

scArches takes the other route. According to the README, it "allows your single-cell query data to be analyzed by integrating it into a reference atlas." The reference stays where it is; the query is mapped onto it. The README lists three concrete payoffs: transferring cell-type annotation from reference to query, identifying disease states by mapping to a healthy atlas, and imputing missing modalities or spatial locations.

The audience is narrow and specific. This is a library for people who already have an AnnData object, already know which reference they want to map against, and are willing to work inside the scvi-tools ecosystem. It is not a point-and-click tool, and the README does not pretend otherwise.

Architecture surgery: what the package actually changes in a trained model

The name is literal. The package name expands to "single-cell architecture surgery", and setup.py describes the project as "Transfer learning with Architecture Surgery on Single-cell data". The mechanism is transfer learning on a model that has already been trained on the reference.

A reference model is trained first. The surgery step then modifies the network so it can accept the query data, and the query is fitted into that modified architecture while the reference structure is preserved. The README frames the result as mapping into an integrated reference rather than a fresh joint embedding.

The implementation sits on top of scvi-tools. setup.py pins scvi-tools>=0.12.1 alongside torch>=1.8.0, scanpy[leiden]>=1.6.0, scHPL>=1.0.0 and muon. That dependency list is the clearest statement of scope in the repository: scArches is a layer over an existing probabilistic-model framework, not a self-contained one. If your reference model is not a scvi-tools model, the surgery step has nothing to operate on. The repository is mostly Jupyter Notebook by language, which matches a project whose primary artefacts are worked examples rather than a large library surface.

Installing scArches and running a first reference mapping

The README does not give install steps. It says, in a section headed "Usage and installation", to see the documentation and tutorials at scarches.readthedocs.io. The PyPI badge in the README points at the package name scarches, which is the distribution name on PyPI, while setup.py declares the package as scArches.

The dependency set is heavy enough that a dedicated environment is the sensible starting point, and the repository ships an envs/ directory for exactly that purpose. Installing from PyPI pulls the whole scvi-tools and torch stack:

bash
pip install scarches

Because setup.py sets SKLEARN_ALLOW_DEPRECATED_SKLEARN_PACKAGE_INSTALL to "True" with the comment that readthedocs otherwise fails, expect the dependency tree to include a deprecated sklearn shim. That environment variable is set inside setup.py itself, not something you normally need to export.

A first real use follows the two-stage pattern the project name implies: obtain a reference model, then map the query through it. The exact class names and call signatures are documented in the Read the Docs tutorials and in the notebooks/ directory of the repository, not in the README. Read one of those notebooks end to end before writing your own script, because the surgery call depends on which reference model you trained. What you should see after a successful mapping is a query object embedded in the reference space, at which point the README's stated use cases (annotation transfer, disease-state mapping) become possible.

Where scArches is the wrong tool

The first limitation is the reference itself. scArches maps onto an atlas; it does not build one. If no suitable reference model exists for your tissue, condition or modality, there is nothing to map into, and the package cannot help. The README's own examples assume a reference is already in hand.

The second is ecosystem lock-in. Every entry in install_requires points at the scvi-tools stack. A lab standardized on a different integration framework cannot use the surgery step without retraining its reference inside scvi-tools first, which is a larger project than the mapping itself.

The third is documentation surface. The README is short and routes everything to Read the Docs. It does not document rollback, it does not document version compatibility between scArches and specific scvi-tools releases, and it does not list the supported model classes. setup.py pins scvi-tools>=0.12.1, an open-ended lower bound, which means a newer scvi-tools release can change behaviour under a scArches version that was never tested against it.

Finally, the release history is thin. The most recent tagged release is v0.5.1 from 2023-06-27, described only as "adding Zenodo DOI". The last push to the repository was on 2026-06-26, so work continues on master, but a user pinning to a tagged release is pinning to something three years old. Treat master, not the tags, as the current state.

How scArches differs from joint integration with Scanpy and Harmony

The obvious alternative is joint integration: concatenate reference and query, then run a batch-correction method such as Harmony through Scanpy, or run the scVI model over the combined object. scanpy[leiden] is already a scArches dependency, so this is not a foreign workflow to the same users.

The difference is what happens to the reference. Joint integration recomputes the embedding over all cells, so the reference coordinates move. Any annotation, cluster label or downstream result tied to those coordinates has to be recomputed too, and results are not reproducible across query batches because each new query reshapes the space. scArches keeps the reference fixed and fits the query into it, which is why annotation transfer is meaningful: the labels still sit where the reference put them.

The cost of that stability is that you must commit to a reference up front. Joint integration degrades gracefully when the query is unlike the reference, because the query can pull the embedding. A fixed reference cannot adapt, and a query population absent from the reference has nowhere sensible to land. That trade-off, stability for adaptability, is the whole decision. If your reference is a curated atlas you intend to reuse across many query samples, scArches is the right shape. If you are integrating a handful of datasets once, joint integration is simpler and has no reference to maintain.

Maintenance, releases and the licence question

The repository is not archived. The last push was on 2026-06-26, which is recent enough that the project cannot be described as abandoned, but the release cadence tells a different story: v0.5.1 on 2023-06-27, v0.5.0 on 2022-02-08, v0.4.0 on 2021-08-05. Tagged releases are roughly annual to biennial, and the most recent one is a DOI addition rather than a feature release. Anyone tracking the project should follow master and the tests/ directory rather than waiting for a tag.

Upgrade cost is dominated by the transitive dependencies, not by scArches itself. A scvi-tools or torch major bump can break a working mapping script even when the scArches version is unchanged, because the lower bounds in setup.py are open-ended. Pin the full environment, not just scarches.

On licensing, the repository metadata states BSD-3-Clause, while setup.py declares license='MIT' and the classifiers list an MIT License classifier. Those two statements disagree, and the LICENSE file at the top level is the one to read. This is a discrepancy to resolve with your own review, not something to infer from the badge or the classifier.

Editorial conclusion

Adopt scArches if you already run scvi-tools models and need to project new query samples into a fixed reference without retraining the reference from scratch. Do not adopt it if you want a standalone pipeline with a self-contained README: the repository README defers to the Read the Docs site for installation and tutorials, and the last push was on 2026-06-26 with the most recent tagged release v0.5.1 dated 2023-06-27, so check the docs and the tests directory for the API you actually intend to call before committing a pipeline to it.

Frequently asked questions

Does scArches have installation instructions in its README?

No. The README's "Usage and installation" section points to the documentation and tutorials at scarches.readthedocs.io instead of listing steps, and the PyPI badge in the README points at the scarches package.

What Python packages does scArches depend on?

setup.py lists scanpy[leiden]>=1.6.0, scHPL>=1.0.0, numpy>=1.19.2, scipy>=1.5.2, scikit-learn>=0.23.2, matplotlib>=3.3.1, pandas>=1.1.2, torch>=1.8.0, scvi-tools>=0.12.1, tqdm>=4.56.0, requests, gdown and muon.

What can I do with scArches once my query data is mapped to a reference?

The README names three applications: transferring cell-type annotation from reference to query, identifying disease states by mapping to a healthy atlas, and imputing missing data modalities or spatial locations.

Is scArches still being developed?

The repository is not archived and the last push was on 2026-06-26, but the most recent tagged release is v0.5.1 from 2023-06-27, so tags lag behind the default branch.

What licence is scArches released under?

The repository metadata states BSD-3-Clause, while setup.py declares license='MIT' and includes an MIT License classifier, so the two disagree and the top-level LICENSE file is the one to read.

Official sources

  1. License: BSD-3-Clause
  2. Project website
  3. README
  4. Releases
  5. theislab/scarches on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/theislab-scarches.svg)](https://hysenlabs.com/projects/theislab-scarches)