CellRank 2: Unified Fate Mapping for Multi-View Single-Cell Data
CellRank: dynamics from multi-view single-cell data
At a glance
- What is it?
- CellRank is a Python library for studying cellular dynamics by building Markov state models from multi-view single-cell data. It is designed for computational biologists who want to infer differentiation direction, terminal cell states, and driver genes from RNA velocity, pseudotime, metabolic labels, or experimental time points without being locked into a single biological prior.
- Who is it for?
- CellRank is well suited for computational biologists working in the scverse ecosystem who need to infer cell fate from any combination of RNA velocity, pseudotime, or metabolic signals. It is not the right tool for researchers who need a single-prior, opinionated workflow: the modular design requires choosing and combining kernels, which adds configuration overhead compared to more constrained tools.
- Can I use it commercially?
- Yes. BSD-3-Clause is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 14 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What CellRank Solves and Who It Is For
Single-cell RNA sequencing captures a snapshot of cells at different stages of a biological process, but does not directly reveal which state each cell is moving toward. CellRank addresses this by modeling cell state transitions as a Markov chain, then using that chain to estimate the probability that a cell will eventually reach each terminal state. The result is a fate probability map over the full cell population.
The library is designed for researchers who work on differentiation, development, disease progression, or any process where cells move through distinct states over time. CellRank 2 was published in Nature Methods in 2024 and supports multiple biological priors for constructing the transition kernel: RNA velocity (from La Manno et al. 2018 and Bergen et al. 2020), pseudotime, developmental potential, experimental time points, and metabolic labels. Supporting multiple priors in a single framework is the primary design goal of version 2. The library scales to large cell numbers and is fully compatible with the scverse ecosystem, which means it integrates directly with scanpy and AnnData objects.
The Markov State Modeling Mechanism
CellRank's backend is built on pyGPCCA, an implementation of the Generalized Perron Cluster Cluster Analysis algorithm described in Reuter et al. (2018). pyGPCCA takes the cell-state transition matrix produced by CellRank's kernel layer and identifies macrostates: groups of cells that the Markov chain tends to enter and rarely leave. These macrostates correspond to biologically meaningful states such as progenitor populations or terminal cell types.
The kernel layer constructs the transition matrix from whichever biological prior the researcher chooses. A velocity kernel uses RNA velocity vectors to assign transition probabilities between neighboring cells. A pseudotime kernel uses gradient information from a precomputed pseudotime. A metabolic labeling kernel uses pulse-chase RNA labeling data. CellRank 2's key design contribution, as described in its paper, is making these kernels composable: a researcher can combine a velocity kernel with a pseudotime kernel by weighting their transition matrices, producing a consensus that neither prior alone would give.
Optional JAX acceleration is available via the `cellrank[jax]` extra dependency, which can speed up matrix computations when JAX and jaxlib are installed.
Installing CellRank and Running an Analysis
CellRank is available on PyPI. The standard install is:
pip install cellrankThe pyproject.toml requires Python 3.12 or later, which is a stricter lower bound than many single-cell tools. The core dependencies include anndata (0.11+), scanpy (1.11+), pyGPCCA (1.0.4+), numpy (2+), pandas (2.2+), and scipy (1.14+). Numba (0.60+) is required for JIT-compiled performance-critical paths. Optional extras for moscot integration are available via `pip install cellrank[moscot]`.
The README points to the documentation at cellrank.readthedocs.io for tutorials. It also notes that a Discourse forum at discourse.scverse.org is the primary community support channel for questions and bug reports. The v2.3.3 release was published on 2026-09-16, the date of the last repository push, indicating the release was the direct motivation for the update.
Key Outputs: Macrostates, Fate Probabilities, and Driver Genes
From a fitted model, CellRank can compute three categories of output. First, it identifies initial, terminal, and intermediate macrostates. Terminal macrostates are the absorbing states of the Markov chain, corresponding to cell fates the population is moving toward. Initial macrostates are the states the chain starts from.
Second, CellRank computes fate probabilities for each cell. These are continuous values between 0 and 1 that quantify how likely a given cell is to end up in each terminal state. Cells that sit at a branch point in the differentiation trajectory will show non-trivial probability mass toward multiple fates.
Third, CellRank can identify driver genes: genes whose expression is correlated with commitment to a specific fate. The README describes this as inferring driver genes for individual fates. CellRank can also visualize and cluster gene expression trends along pseudotime, showing how gene modules activate or deactivate as cells progress through the trajectory.
Limitations and Cases Where CellRank Is the Wrong Tool
CellRank's analysis depends entirely on the quality of the input prior. If RNA velocity estimates are noisy, as they often are in datasets with low splicing signal or high technical noise, the transition matrix will be unreliable and the inferred macrostates will not correspond to real biology. The library does not correct for poor velocity estimates; it propagates them.
The Python 3.12+ requirement is a practical constraint. Many established single-cell analysis environments were built on earlier Python versions and require careful dependency management to upgrade. The dependency stack is also heavy: scanpy, anndata, numba, scipy, and matplotlib are all required, which means the install footprint is larger than a minimal analysis environment.
CellRank is not designed for bulk RNA-seq data or spatial transcriptomics without single-cell resolution. It assumes the input is a cell-by-gene matrix in AnnData format, compatible with the scverse ecosystem. Data in formats designed for other frameworks will need to be converted first.
CellRank vs scVelo: Scope and Design Differences
scVelo is the most common tool that users compare to CellRank, as the search query "cellrank vs scvelo" reflects. The two tools address different problems. scVelo estimates RNA velocity: the direction and speed of transcriptional change at each cell, inferred from the ratio of spliced to unspliced RNA. It produces velocity vectors. CellRank takes those vectors (or other priors) as input and uses them to model population-level fate probabilities via Markov chains.
The practical consequence is that they are complementary rather than competing. A typical workflow runs scVelo first to estimate velocity, then feeds those estimates into CellRank to compute macrostates and fate probabilities. CellRank 2 also accepts priors that have nothing to do with RNA velocity, such as pseudotime or metabolic labels, which makes it usable in datasets where velocity estimation is unreliable or simply not available. scVelo does not provide fate probability outputs or driver gene inference in the same framework.
Maintenance, License, and the scverse Ecosystem
CellRank is maintained by Marius Lange (ETH Zurich) and Philipp Weiler, as listed in the pyproject.toml maintainers field. The v2.3.3 release on 2026-09-16 is the most recent, and the release history shows three releases in 2026 (v2.3.1 in June, v2.3.2 in June, v2.3.3 in September), indicating an active release cadence. The project uses semantic versioning and publishes release notes on GitHub.
The license is BSD-3-Clause, which permits commercial and academic use, modification, and redistribution with attribution and preservation of the license notice. CellRank is part of the scverse ecosystem, which also includes moscot for optimal transport, VeloVI for uncertainty-aware RNA velocity, and RegVelo for joint gene regulation and velocity learning. These related tools are all linked from the CellRank README and share AnnData as a common data structure.
Editorial conclusion
CellRank is well suited for computational biologists working in the scverse ecosystem who need to infer cell fate from any combination of RNA velocity, pseudotime, or metabolic signals. It is not the right tool for researchers who need a single-prior, opinionated workflow: the modular design requires choosing and combining kernels, which adds configuration overhead compared to more constrained tools. Before adopting CellRank, verify that your Python environment meets the 3.12+ requirement and that your AnnData object is formatted for scverse compatibility.
Frequently asked questions
what is cell rank
CellRank is a Python library that models cellular dynamics by building Markov chains from single-cell data. It uses biological signals such as RNA velocity, pseudotime, or metabolic labels to estimate which terminal cell states individual cells are likely to reach, and infers macrostates and driver genes from those estimates.
How does CellRank differ from scVelo?
scVelo estimates RNA velocity vectors from spliced and unspliced RNA ratios. CellRank takes those vectors (or other priors) as input and models population-level fate probabilities via Markov chains. The two tools are typically used together: scVelo produces velocity estimates, CellRank uses them to infer macrostates and fate probabilities.
Does CellRank require RNA velocity data?
CellRank 2 supports multiple biological priors beyond RNA velocity, including pseudotime, developmental potential, experimental time points, and metabolic labeling. Researchers can combine multiple priors by weighting their transition matrices, which is useful in datasets where RNA velocity estimates are unreliable.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/scverse-cellrank)