Library / SDK
OmicsML/dance avatar
OmicsML/dance

DANCE: a deep learning library and benchmark for single-cell analysis

DANCE: a deep learning library and benchmark platform for single-cell analysis

387 stars37 forksPythonBSD-2-Clause

At a glance

What is it?
DANCE packages single-cell deep learning methods into one Python library with reproducible example scripts, so cell type annotation and clustering benchmarks can be rerun from a command line. Its strength is reproduction, not novelty.
Who is it for?
Adopt DANCE if you need to reproduce a published single-cell benchmark or compare methods under one preprocessing pipeline; the examples directory and pinned requirements.txt make that feasible. Do not adopt it expecting DANCE 2.0 features: the README lists the 2.0 codebase, web platform and API as unreleased checkboxes, and the newest release is v1.0.1 from 2023-12-04.
Can I use it commercially?
Yes. BSD-2-Clause is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 19 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What DANCE is for, and who it is actually for

Single-cell computational methods appear faster than anyone can compare them. Each paper ships its own preprocessing, its own train/test split and its own evaluation code, often in a different language or against incompatible library versions. The README states the problem directly: reproducing previous benchmarks is a key challenge, because different studies prepare data and evaluate differently, and the methods themselves may not be compatible.

DANCE answers that with a unified Python package that implements several popular computational single-cell methods and provides shared tooling for data downloading, data preprocessing and transformation (graph construction is named as an example), and model training and evaluation. The audience is therefore narrow and specific: computational biologists and machine learning researchers who want to rerun a published experiment, or who want to compare methods on identical inputs. If you are looking for a new model architecture, this is not that. DANCE 1.0 is positioned as a library and benchmark platform, and the README describes its main usage as providing readily available experiment reproduction.

The scope covers three modules: single-modality analysis (cell type annotation, clustering, gene imputation), single-cell multimodal omics (modality prediction and modality matching in DANCE 1.0, joint embedding), and spatially resolved transcriptomics (spatial domain identification, cell type deconvolution). That breadth is the point. One install gives you the preprocessing and evaluation scaffolding for all three.

How the library is organized: modules, examples and pinned dependencies

The repository layout is the clearest statement of intent. There is a dance/ package directory, an examples/ directory split by task, a tests/ directory, and a docker/ directory. The examples tree mirrors the three modules: examples/single_modality/, examples/multi_modality/, examples/spatial/, plus examples/atlas/, examples/tuning/ and examples/result_analysis/. A dataset_server.json file sits at the top of examples/, which is consistent with the README's claim that data downloading is handled by unified tooling rather than by each script.

The practical consequence is that a benchmark is a script plus command line arguments, not a notebook you adapt by hand. The README's worked example points at examples/single_modality/cell_type_annotation/scdeepsort.py and says the CLI options for reproducing a particular experiment are documented at the end of the script. That is an unusual documentation choice: the interface lives in the code, so you read the file's tail before you can run anything.

Dependencies are pinned hard. requirements.txt fixes scanpy==1.10.1, numpy==1.26.4, anndata==0.10.8, scib==1.1.5, pyro-ppl==1.9.0, torchnmf==0.3.5 and wandb==0.16.3, among others. It also carries a commented line noting that scib 1.1.4 requires pandas<2, which tells you the maintainers have already hit that conflict. Pinning this tightly is what makes reproduction possible, and it is also what will make the library awkward to drop into an environment that already has newer versions of scanpy or numpy.

Installing DANCE and running a cell type annotation benchmark

The README points to an Installation section and ships an install.sh at the repository root. The package is published on PyPI as pydance, which is the name in the PyPI badge. The README does not spell out the pip command in the text available here, so if you install from the repository, read install.sh first rather than guessing at flags.

The workflow the README describes has three steps after installation. Navigate to the folder holding the example script, read the CLI options at the end of that script, then run it and wait. For the scDeepSort cell type annotation example, the README gives this exact invocation for the Mouse Brain experiment:

bash
python scdeepsort.py --species mouse --tissue Brain --train_dataset 753 3285 --test_dataset 2695

The numeric arguments are dataset identifiers, not sizes: the README passes two training datasets and one test dataset by ID. Expect the script to fetch or locate those datasets through the shared data tooling, then train and evaluate. Because requirements.txt pins wandb==0.16.3, some scripts may attempt to log to Weights & Biases; the README does not document an offline or disabled mode, so that is something to check in the script before a long run.

If you prefer a container, the repository has a docker/ directory, though the README does not document the image names or build commands in the text available here. A development environment is also described in CONTRIBUTING.md, which the README references for automated quality controls and dev setup.

Where DANCE 1.0 stops: DANCE 2.0 is announced, not shipped

The README is candid about the gap between the published papers and the code. DANCE 2.0 is described as extending the toolkit with an automated preprocessing recommendation platform, moving preprocessing away from trial and error toward a systematic, data-driven and interpretable workflow. But the release schedule is a checklist with every box unchecked: open-source release of the DANCE 2.0 codebase, launch of the DANCE 2.0 web platform, and release of the DANCE 2.0 API.

That matters for adoption decisions. A bioRxiv paper on DANCE 2.0 exists, and the README links it, but nothing in the repository indicates that the code is available. Anyone reading the 2.0 paper and expecting a pip install will be disappointed. The newest release in the list is v1.0.1, dated 2023-12-04, described as bug fixes and new datasets; v1.0.0 landed on 2023-06-29. The last push to the default branch was on 2026-09-10, so the repository is not archived and work is happening, but the released artifact is still 1.0.1.

The second limitation is scope creep by design. DANCE implements other people's methods. If a method's original authors change it, or if a newer method beats it, DANCE does not automatically follow. The README points readers to a reproduced performance table for the implemented algorithms, which is the honest way to present this: the value is in the reproduction, and the reproduction is frozen at the versions in requirements.txt.

DANCE compared with assembling Scanpy and scIB yourself

The obvious alternative is to build the same pipeline from Scanpy for preprocessing and scIB for benchmarking, both of which DANCE already depends on. That approach gives you current versions and full control over every step. The difference is that you own the glue: dataset download, graph construction, train/test splitting and metric computation all become your code, and two people on the same team can end up with two different pipelines.

DANCE's approach is the opposite trade. It fixes the glue and pins the versions, so a benchmark is a script with dataset IDs on the command line. You get comparability at the cost of flexibility. If your dataset does not fit the expected format, or if you need a preprocessing step DANCE does not implement, you are working against the library rather than with it. The multi-modality and spatial modules sharpen this: modality prediction and modality matching exist only in DANCE 1.0, and joint embedding is the shared piece, so a multi-omics project that needs prediction is tied to the 1.0 line.

For a team that mainly wants to run its own method against baselines, assembling Scanpy and scIB is more direct. For a team that wants to reproduce a specific published number, DANCE removes exactly the work that makes reproduction fail.

Maintenance, licensing and the cost of upgrading

DANCE is BSD-2-Clause licensed, which is permissive and places few obligations on how you redistribute or modify it. The repository includes LICENSE, MANIFEST.in and setup.cfg, and builds through setuptools with a pyproject.toml that only declares the build backend and tool configuration. There is no separate legal consideration described beyond the licence itself.

The maintenance picture is mixed in a way worth stating plainly. The last push was on 2026-09-10, so the repository is being touched. The releases, however, stop at v1.0.1 on 2023-12-04. That combination usually means development is happening on the main branch ahead of any tagged release, and the README's unchecked 2.0 schedule supports that reading. If you depend on a released artifact, you are depending on code from 2023.

Upgrade cost is dominated by the pinned requirements. Moving to a newer scanpy or numpy means re-testing the example scripts, because the pins exist precisely to keep results comparable. The commented pandas constraint in requirements.txt is a preview of the kind of conflict you will meet. A reasonable approach is to keep DANCE in its own environment rather than merging it into an existing analysis environment, and to treat any version bump as a change that invalidates previous benchmark numbers until rerun.

Editorial conclusion

Adopt DANCE if you need to reproduce a published single-cell benchmark or compare methods under one preprocessing pipeline; the examples directory and pinned requirements.txt make that feasible. Do not adopt it expecting DANCE 2.0 features: the README lists the 2.0 codebase, web platform and API as unreleased checkboxes, and the newest release is v1.0.1 from 2023-12-04. Verify first that the example script you need runs against the pinned dependency set, since requirements.txt fixes versions such as scanpy==1.10.1 and numpy==1.26.4.

Frequently asked questions

How do I install DANCE?

The package is published on PyPI as pydance, and the repository also ships an install.sh at the root. The README points to an Installation section but does not spell out the pip command, so read install.sh if you install from the repository.

What can DANCE be used for?

The README lists three modules: single-modality analysis (cell type annotation, clustering, gene imputation), single-cell multimodal omics (modality prediction and matching in 1.0, joint embedding), and spatially resolved transcriptomics (spatial domain identification, cell type deconvolution).

Is the DANCE 2.0 codebase available?

No. The README's DANCE 2.0 release schedule lists the open-source codebase, the web platform and the API as unchecked items, and the newest release in the list is v1.0.1 from 2023-12-04.

How do I run a cell type annotation benchmark with DANCE?

The README's example navigates to examples/single_modality/cell_type_annotation, reads the CLI options at the end of scdeepsort.py, then runs the script with dataset IDs, for example python scdeepsort.py --species mouse --tissue Brain --train_dataset 753 3285 --test_dataset 2695.

What licence does DANCE use?

BSD-2-Clause, according to the licence badge and the LICENSE file in the repository.

Official sources

  1. License: BSD-2-Clause
  2. OmicsML/dance on GitHub
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/omicsml-dance.svg)](https://hysenlabs.com/projects/omicsml-dance)