UMAP: non-linear dimension reduction for Python, and when it is the wrong tool
Uniform Manifold Approximation and Projection
At a glance
- What is it?
- lmcinnes/umap implements Uniform Manifold Approximation and Projection as a scikit-learn compatible transformer. It is a strong default for exploratory embedding, but the axes it produces are not measurements of anything.
- Who is it for?
- Adopt umap-learn if you need a fast, scikit-learn shaped non-linear embedding for exploratory work on tabular or single-cell style data, and you have numba and scikit-learn 1.6 or newer available. Do not adopt it if you need a stable coordinate system to compare across runs, or if you need an interpretable linear basis; PCA or a supervised model is the better choice there.
- Can I use it commercially?
- Yes. BSD-3-Clause is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 15 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 17, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What UMAP solves and who actually needs it
High-dimensional data is hard to look at. The README frames UMAP as a dimension reduction technique usable for visualisation similarly to t-SNE, and also for general non-linear dimension reduction. That second clause is the one that matters for engineering work. t-SNE was built for pictures; UMAP's README states it can be used as a drop in replacement for scikit-learn's t-SNE, but the package also exposes fit, transform and inverse_transform semantics that make it usable as a preprocessing step rather than only a final plot.
The intended audience is stated in the pyproject classifiers: Science/Research and Developers. In practice that means people working in Python who already have numpy, scipy and scikit-learn in their environment and want an embedding they can put inside a Pipeline. The algorithm rests on three assumptions listed in the README: the data is uniformly distributed on a Riemannian manifold, the Riemannian metric is locally constant, and the manifold is locally connected. The README is explicit that you do not need to worry about that, which is fair, but the assumptions are the reason the method behaves the way it does on sparse or clustered data.
How the fuzzy topological structure becomes a 2D embedding
The mechanism described in the README is a two-stage pipeline. First the data is modelled as a fuzzy topological structure derived from those three manifold assumptions. Then the embedding is found by searching for a low dimensional projection whose fuzzy topological structure is as close as possible to the original. That is a search, not a closed-form projection, which is the structural difference from PCA.
The practical dependency is pynndescent, listed as a required dependency in pyproject.toml. Nearest neighbour search is what makes the first stage tractable on large inputs, and pynndescent is the approximate nearest neighbour library that backs it. The second stage is a numba-compiled optimisation. The README says the original version used Cython, and that improved code clarity, simplicity and performance of Numba made the transition necessary. That choice explains the install constraints: numba is a hard dependency, and numba pins its own supported Python and numpy ranges, which is usually the first thing to break in a fresh environment.
densMAP is a separate mode bundled in the same package. The README states the densMAP algorithm augments UMAP to preserve local density information in addition to the topological structure, and points to the Nature Biotechnology 2021 paper by Narayan, Berger and Cho. If your question is about density rather than neighbourhood identity, that is the variant to look at, not plain UMAP.
Installing umap-learn and running a first embedding
Conda is the path the README highlights first, crediting the conda-forge team. The conda-forge packages are stated to be available for Linux, OS X and Windows 64 bit.
conda install -c conda-forge umap-learnIf you prefer pip, the README presumes you already have numba, scikit-learn and their requirements installed. Note that the distribution name on PyPI is umap-learn, not umap, while the import name is umap.
pip install umap-learnThe README documents three extras. The plot extra pulls in the visualisation stack, the parametric_umap extra installs a CPU-only Tensorflow, and the tbb extra is described as an optional CPU optimisation for x86 processors.
pip install umap-learn[plot]
pip install umap-learn[parametric_umap]
pip install umap-learn[tbb]The README also gives a manual route for when pip struggles with dependencies: install numpy and scipy, scikit-learn and numba through conda, then pull the package from pip. A manual source install downloads the master archive, unpacks it, and runs an editable install from the package directory.
conda install numpy scipy
conda install scikit-learn
conda install numba
pip install umap-learnFor a first real use, the README's own example is the shortest path. It loads the digits dataset and fits the default UMAP instance, which returns the embedding array directly.
import umap
from sklearn.datasets import load_digits
digits = load_digits()
embedding = umap.UMAP().fit_transform(digits.data)What you should see is a numpy array with one row per input sample and two columns by default. The README notes the package inherits from sklearn classes and drops in next to other sklearn transformers with an identical calling API, so the returned object slots into a Pipeline the same way a decomposition estimator would. Two parameters dominate the result. n_neighbors controls how many neighbouring points are used in the local approximations; the README says larger values preserve more global structure at the loss of local detail, that the range is often 5 to 50, and that 10 to 15 is a sensible default. min_dist controls how tightly points may be compressed together: larger values spread points more evenly, smaller values let the optimisation place points more accurately.
The embedding is not a coordinate system
The most common misuse of UMAP is treating the output axes as features. They are not. The optimisation searches for a projection matching topological structure, and the README's own parameter description makes the trade-off explicit: n_neighbors is a dial between global structure and local detail, and min_dist changes the geometry of the output without changing the data. Two runs with different values of either parameter produce different pictures of the same input, and neither is more correct.
A second limitation is environmental. numba is a required dependency and the README's install section repeatedly routes around dependency friction, including a fallback that installs numpy, scipy, scikit-learn and numba through conda before pulling umap-learn from pip. If you are shipping to a platform where numba wheels lag, or you are pinned to an older numpy, expect the install to be the hard part rather than the algorithm.
A third is that UMAP is the wrong tool when you need reproducibility across datasets. Fitting a fresh embedding on a new batch gives coordinates that are not comparable to the previous batch. The repository does include examples/mnist_transform_new_data.py, which is the transform path for new points, but the README does not document rollback or a stable global frame. If your downstream step compares coordinates across runs, UMAP is the wrong layer.
UMAP against PCA and t-SNE
PCA is a linear projection onto orthogonal directions of maximum variance. It is deterministic, it has a closed form, and each output axis has a stated meaning in terms of the input features. UMAP has none of those properties. If you need to say which input variables drive the separation, or you need the same coordinates tomorrow, PCA wins and UMAP cannot substitute for it.
t-SNE is the closer comparison, and the README positions UMAP as a drop in replacement for it. The difference in approach: t-SNE is built around pairwise similarity probabilities and is used almost exclusively for visualisation, while UMAP's README describes it as usable for general non-linear dimension reduction as well. The sklearn-compatible API is the operational difference. A UMAP object fits into a transformer chain and exposes transform for new data, which is why the repository carries examples like mnist_transform_new_data.py and inverse_transform_example.py.
Parametric UMAP is a third option inside the same project. It requires Tensorflow, listed in pyproject.toml as tensorflow >= 2.1 under the parametric_umap extra, and the README points to the Tensorflow install instructions as the recommended route. The trade-off is real: you take on a deep learning framework as a dependency in exchange for an embedding function that can be applied to new data without refitting.
Maintenance, licence and upgrade cost
The repository is not archived and the last push was on 2026-09-05. Recent releases are dated 2026-04-08 (release-0.5.12), 2026-01-12 (release-0.5.11) and 2025-12-11 (release-0.5.10.post2). pyproject.toml pins version 0.5.12 and requires Python 3.10 or newer, with classifiers for 3.10 through 3.13. The minimum dependency floors are numpy >= 1.23, scipy >= 1.3.1, scikit-learn >= 1.6, numba >= 0.51.2, pynndescent >= 0.5 and tqdm. The scikit-learn floor of 1.6 is the one that forces upgrades most often, because it moves faster than the others.
Licensing is BSD-3-Clause, and pyproject.toml declares license = {text = "BSD"}. That is a permissive licence, so the usual obligation is preserving the copyright notice and licence text in redistributions. Nothing here changes if you use it in a closed product. This is a description of the declared terms, not legal advice; check LICENSE.txt in the repository for the operative text.
One oddity worth flagging: the pyproject classifiers declare Development Status :: 3 - Alpha despite a 0.5.x version line and a JOSS paper. Treat the classifier as stale metadata rather than a signal about the code. Upgrade cost is dominated by the numba and scikit-learn floors, not by UMAP's own API, which the README presents as stable enough to sit alongside other sklearn transformers.
Editorial conclusion
Adopt umap-learn if you need a fast, scikit-learn shaped non-linear embedding for exploratory work on tabular or single-cell style data, and you have numba and scikit-learn 1.6 or newer available. Do not adopt it if you need a stable coordinate system to compare across runs, or if you need an interpretable linear basis; PCA or a supervised model is the better choice there. Before committing, verify that numba compiles on your target platform, check whether you need the plot, parametric_umap or tbb extras, and confirm that your downstream analysis does not treat embedding distances as metric quantities.
Frequently asked questions
What is UMAP and how does it work?
UMAP stands for Uniform Manifold Approximation and Projection. The README describes it as a dimension reduction technique that models the data as a fuzzy topological structure and then searches for a low dimensional projection with the closest equivalent structure.
How should I interpret UMAP results?
The README does not give interpretation guidance beyond noting that n_neighbors trades global structure against local detail and that min_dist controls how tightly points are compressed. Those two parameters change the picture without changing the data, so the output is best read as neighbourhood structure rather than as coordinates with units.
What is the difference between UMAP and t-SNE?
The README states UMAP can be used for visualisation similarly to t-SNE and also for general non-linear dimension reduction, and that it works as a drop in replacement for scikit-learn's t-SNE. The package inherits from sklearn classes, so it fits into a transformer chain with fit and transform rather than producing only a final plot.
What is UMAP and how does it work for dimension reduction?
The README says the algorithm rests on three assumptions about the data: it is uniformly distributed on a Riemannian manifold, the Riemannian metric is locally constant, and the manifold is locally connected. From those it builds a fuzzy topological structure and searches for a low dimensional projection with the closest equivalent structure.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/lmcinnes-umap)