Open-source project
pavlin-policar/openTSNE avatar
pavlin-policar/openTSNE

openTSNE: A Modular Python t-SNE Implementation With Embedding to Reference Space

Extensible, parallel implementations of t-SNE

1,625 stars177 forksPythonBSD-3-Clause

At a glance

What is it?
openTSNE is a Python library for t-SNE dimensionality reduction, built around FIt-SNE and Barnes-Hut and designed to scale to millions of points. The judgement: adopt it when you need to add new points to an existing embedding or want multiple t-SNE variants behind one API, and skip it when a single small dataset is all you have.
Who is it for?
Adopt openTSNE if you already have a t-SNE embedding and need to place new samples into it, or if you want to switch between FIt-SNE and Barnes-Hut without rewriting your pipeline. Do not adopt it if you need a documented rollback path, a pure-Python install on a machine with no C/C++ compiler, or a maintained release cadence tied to the latest Python versions.
Can I use it commercially?
Yes. BSD-3-Clause is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 22 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What openTSNE solves that scikit-learn's TSNE does not

The scikit-learn TSNE class is a single implementation with a single algorithm behind it. openTSNE is a package built around two: FIt-SNE, which the README describes as the default, and Barnes-Hut, which the README says is available as an alternative. The README states that openTSNE incorporates the latest improvements to t-SNE, including the ability to add new data points to existing embeddings, speed improvements that let t-SNE scale to millions of data points, and tricks to improve global alignment of the resulting visualizations.

The audience is narrow and identifiable. Anyone embedding single-cell transcriptomes, where the reference dataset is fixed and new batches keep arriving, gets the most out of the transform path. The README's own figure is a mouse retina embedding of 44,808 single cell transcriptomes from Macosko 2015, produced with the multiscale kernel trick to better preserve global alignment. That figure is the pitch: a dataset size and a problem shape that the stock t-SNE implementation handles poorly.

If you have 5,000 rows and one embedding to produce, openTSNE is more machinery than the task needs. The value shows up when the embedding is a long-lived artefact that new data must be projected into, not a one-off plot.

The two algorithms and how the transform path works

The repository is a Python package with compiled extensions. pyproject.toml declares the build requirements as setuptools, wheel, cython and numpy, and setup.py uses Cython's build_ext along with distutils compiler detection, including CompileError and LinkError handling. That layout tells you the numerical core is Cython plus C, not pure Python, which is why the README warns that openTSNE requires a C/C++ compiler to be available on the system.

The README states that openTSNE implements two efficient algorithms for t-SNE and asks you to cite the original authors of whichever one you use: FIt-SNE (the default) cites Linderman et al. 2019, and Barnes-Hut cites Van Der Maaten 2014 and Yang et al. 2013. Those are different approximations of the same objective with different cost profiles. Picking one is a modelling decision, not a performance toggle, and the README treats it that way by attaching separate citations.

The transform capability comes from the paper cited as reference [2], Poličar, Stražar and Zupan, Embedding to Reference t-SNE Space Addresses Batch Effects in Single-Cell Classification. The mechanism, as the title states, is embedding to a reference t-SNE space: you keep the original embedding and project new points into it rather than recomputing from scratch. The README does not spell out the numerical details of that projection, so treat the paper as the source for how it behaves under batch effects.

FFTW3 is the other moving part. The README says FFTW3 implements the Fast Fourier Transform, which is heavily used in openTSNE, and that if FFTW3 is not available openTSNE falls back to numpy's FFT, which is slightly slower. The README also scopes the difference: it is only noticeable with large data sets containing millions of data points. That is a useful boundary. On a dataset of tens of thousands of rows, the FFTW3 question is not worth your time.

Installing openTSNE and running a first embedding

There are two package routes, plus a source build. The README gives conda-forge first:

bash
conda install --channel conda-forge opentsne

That pulls a prebuilt package, so no local compiler is involved. The pip route is the second option:

bash
pip install opentsne

The README points at the PyPi package openTSNE for this. If neither wheel fits your platform, the source install is:

bash
pip install .

run in the root directory, which the README says installs the appropriate dependencies and compiles the necessary binary files. The README notes that openTSNE requires a C/C++ compiler to be available on the system, and that for multithreading the compiler must support OpenMP, which almost all compilers implement except older versions of clang on OSX systems. The README also suggests considering installing FFTW3 prior to installation, with the caveat above about when it matters.

The hello world in the README loads the iris dataset with scikit-learn and then fits a default TSNE:

python
from sklearn import datasets

iris = datasets.load_iris()
x, y = iris["data"], iris["target"]

from openTSNE import TSNE

embedding = TSNE().fit(x)

The README presents this as the whole getting-started story, and that is accurate as far as it goes: TSNE() with no arguments uses the defaults the library authors chose. What the README does not show in that snippet is the transform path, the multiscale kernel, or how to select Barnes-Hut, so a reader who wants those needs to go to the examples directory, which ships four notebooks covering simple usage, advanced usage, preserving global structure, and large data sets.

Where openTSNE is the wrong tool

The compiler requirement is a real deployment constraint, not a footnote. The README is explicit that openTSNE requires a C/C++ compiler to be available on the system. If your build environment is a slim container, a locked-down CI runner, or a managed notebook service without a toolchain, the pip and source routes can fail, and you are dependent on a prebuilt conda-forge or PyPI package matching your platform and Python version. The README states openTSNE can be installed on all supported versions of Python but does not enumerate them here; the link goes to the CPython devguide's version list, which is a moving target rather than a pinned support matrix.

The bigger gap is operational. The README does not document rollback, version pinning strategy, or any deprecation policy. The release history is sparse: v1.0.0 in May 2023, v1.0.1 in November 2023, v1.0.2 in August 2024. The last push to the repository was on 2026-09-08, so the code is being touched, but the tagged releases are not frequent. If your project needs a library whose API churn and release notes you can plan around, that cadence is thin.

There is also a conceptual limit that applies to t-SNE generally and openTSNE specifically. The library produces an embedding, and the README's own framing is visualization. Distances between clusters in a t-SNE plot are not a reliable measure of distance in the original space, which is why the README highlights the multiscale kernel trick as a way to better preserve global alignment. If your downstream step needs metric distances, openTSNE is not the component that gives them to you.

openTSNE against UMAP and PCA

UMAP is the comparison people actually make, and the difference is structural rather than a matter of quality. UMAP is built on a different objective and a different graph construction, and it typically produces embeddings faster on mid-sized data with a stronger tendency to preserve some global structure. openTSNE's counterargument is the transform path: the README states that openTSNE can add new data points to existing embeddings, which is the capability the reference [2] paper is about. If your workflow is fit-once, project-many, that is the axis to compare on, not the shape of the scatter plot.

PCA is the other comparison, and it is not really a competitor. PCA is a linear projection with a closed-form solution and an interpretable basis; t-SNE is a nonlinear embedding with no such basis. The README makes no claim that openTSNE replaces PCA, and it should not. A common pattern is PCA for preprocessing followed by t-SNE for visualization, and openTSNE fits into that as the second stage.

Within the t-SNE family, the honest comparison is against scikit-learn's TSNE. scikit-learn gives you one algorithm and a stable API from a large project. openTSNE gives you two algorithms, the transform path, and the multiscale kernel, at the cost of a compiled dependency and a smaller maintainer base. Pick based on whether you need the extras, not on which one is better in the abstract.

Licence, maintenance and upgrade cost

openTSNE is BSD-3-Clause, confirmed by the LICENSE file at the repository root and the badge in the README. That is a permissive licence: it allows use in closed-source products and modification, with the usual conditions around retaining the copyright notice and disclaimer. It does not carry the patent grant language of Apache-2.0, and it does not impose the copyleft obligations of GPL-family licences. The README does not discuss licence implications for downstream users, so this is a reading of the licence identifier rather than a statement from the project. This is not legal advice; check the LICENSE text against your own distribution model.

The citation requirement is not a licence term but is a practical obligation. The README asks that you cite the JSS paper, Poličar, Stražar and Zupan 2024, if you make use of openTSNE for your work, and separately asks you to cite the original authors of whichever algorithm you use: FIt-SNE (Linderman et al. 2019) or Barnes-Hut (Van Der Maaten 2014 and Yang et al. 2013). In an academic or published setting that means two or three bibliography entries, not one.

Upgrade cost is dominated by the compiled extensions. A version bump can require a rebuilt toolchain, a different OpenMP story on macOS clang, or a re-evaluation of the FFTW3 decision. Because the tagged releases are infrequent, upgrades are large jumps rather than small increments, and the README does not document a migration guide between them.

Editorial conclusion

Adopt openTSNE if you already have a t-SNE embedding and need to place new samples into it, or if you want to switch between FIt-SNE and Barnes-Hut without rewriting your pipeline. Do not adopt it if you need a documented rollback path, a pure-Python install on a machine with no C/C++ compiler, or a maintained release cadence tied to the latest Python versions. Before committing, verify that your platform's compiler supports OpenMP and decide whether to install FFTW3 first, since that choice affects whether openTSNE uses its own FFT or falls back to numpy's.

Frequently asked questions

What is the purpose of t-SNE?

The README describes t-SNE as a popular dimensionality-reduction algorithm for visualizing high-dimensional data sets, and openTSNE as a modular Python implementation of it. Its output is an embedding used for visualization, not a metric projection.

What are the differences between t-SNE and PCA?

PCA is a linear projection with an interpretable basis, while t-SNE is a nonlinear embedding used for visualization. The openTSNE README does not present the library as a replacement for PCA, and the two are often used in sequence.

What is the difference between t-SNE and UMAP?

The openTSNE README does not discuss UMAP. What it does state is that openTSNE can add new data points to existing embeddings and includes tricks to improve global alignment of the resulting visualizations, which are the capabilities to compare on.

How can t-SNE be used to visualize data?

The README's hello world loads the iris dataset with scikit-learn, then calls TSNE().fit(x) from openTSNE to produce an embedding. The examples directory ships four notebooks covering simple usage, advanced usage, preserving global structure and large data sets.

Official sources

  1. License: BSD-3-Clause
  2. pavlin-policar/openTSNE on GitHub
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/pavlin-policar-opentsne.svg)](https://hysenlabs.com/projects/pavlin-policar-opentsne)