# hdbscan: density clustering without a radius to tune

> The scikit-learn-contrib hdbscan package implements hierarchical density-based clustering in Python, with a scikit-learn compatible API, soft membership scores and GLOSH outlier detection. It suits exploratory work on unlabelled data, and fits badly when every point must be assigned.

**scikit-learn-contrib/hdbscan** — A high performance implementation of HDBSCAN clustering.

- Repository: https://github.com/scikit-learn-contrib/hdbscan
- Website: http://hdbscan.readthedocs.io/en/latest/
- Stars: 3,153 · Forks: 538
- Language: Jupyter Notebook
- License: BSD-3-Clause
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/scikit-learn-contrib-hdbscan

## What hdbscan is for and who reaches for it

HDBSCAN stands for Hierarchical Density-Based Spatial Clustering of Applications with Noise. The README describes it as performing DBSCAN over varying epsilon values and integrating the result to find the clustering with the best stability over epsilon. That single sentence explains the appeal: instead of picking a neighbourhood radius, you pick a minimum cluster size, and the algorithm searches the density hierarchy for the partition that survives across the widest range of scales.

The intended user is someone doing exploratory data analysis on data with no labels. The README states that HDBSCAN returns a good clustering straight away with little or no parameter tuning, and that the primary parameter, minimum cluster size, is intuitive to select. That claim is worth taking seriously but not literally. Minimum cluster size is easier to reason about than epsilon because it is expressed in points rather than distance units, but it still changes the answer, and the package exposes several other parameters that shape the result.

It is also a scikit-learn-contrib project, so it lives in the same ecosystem as the rest of your pipeline. The README says the package inherits from sklearn classes and drops in next to other sklearn clusterers with an identical calling API. If you already call fit_predict on KMeans or DBSCAN, the migration is close to mechanical.

## How the algorithm actually works in this implementation

The mechanism is a density hierarchy, not a single density threshold. The package builds a mutual reachability graph, computes a minimum spanning tree over it, and then condenses that tree into a hierarchy of candidate clusters. Each candidate is scored by persistence, which the README describes as a stability score over the range of distance scales present in the data. The clustering you get back is the set of clusters that are stable across the widest band of scales.

That structure is exposed rather than hidden. After fitting, the clusterer object carries the condensed cluster hierarchy, the robust single linkage cluster hierarchy, and the reachability distance minimal spanning tree, each with methods for plotting and for conversion to Pandas or NetworkX. This is the part of the package that distinguishes it from a black-box clusterer: you can inspect why two groups were merged or split, and you can see which points sit in low-density regions.

The implementation is not pure Python. The repository's setup.py declares Cython extensions for the tree, linkage, Boruvka, reachability, prediction and distance metric modules, so the heavy graph work runs in compiled code. The pyproject.toml build system requires setuptools, wheel, Cython 3.0.11 or later below 4, and NumPy 2.0 or later below 3. That build requirement is the first thing to check on a constrained platform.

## Installing hdbscan and running a first clustering

The package is distributed on PyPI and on conda-forge, both of which are shown as badges in the README. The straightforward install is:

```bash
pip install hdbscan
```

Because the package compiles Cython extensions, pip needs a working C toolchain and the build dependencies listed in pyproject.toml. On a machine without a compiler, the conda-forge build avoids that step:

```bash
conda install -c conda-forge hdbscan
```

The README gives this example for a first run. It generates a synthetic blob dataset, fits the clusterer with a minimum cluster size of 10, and returns a label per sample:

```python
import hdbscan
from sklearn.datasets import make_blobs

data, _ = make_blobs(1000)

clusterer = hdbscan.HDBSCAN(min_cluster_size=10)
cluster_labels = clusterer.fit_predict(data)
```

What you should see is an integer array of length 1000. Points that the algorithm considers noise carry the label -1, and the remaining points carry non-negative cluster ids. That -1 is not an error code. It is the algorithm telling you the point does not belong to any cluster at the chosen density, and the count of those points is the first diagnostic you should look at.

Input does not have to be a feature matrix. The README states that the package accepts an array, pandas dataframe or sparse matrix of shape (num_samples x num_features), or an array or sparse matrix giving a distance matrix between samples. The distance matrix route is useful when your data is not Euclidean, but it costs quadratic memory in the number of samples.

## Soft membership, persistence scores and GLOSH outliers

Hard labels are only one output. The clusterer object also exposes cluster membership strengths, which the README describes as enabling optional soft clustering at no further compute expense. That matters when a point sits between two clusters: instead of forcing a choice, you can keep the strength vector and decide later. The probabilities_ attribute is the usual way to reach this, and it is populated during the same fit.

Each cluster also receives a persistence score, described in the README as the stability of the cluster over the range of distance scales present in the data. This gives you a relative strength per cluster rather than per point, which is useful for filtering: a cluster with very low persistence is often an artifact of the density estimate rather than a real group.

Outlier detection is a separate algorithm sharing the same fit. The package supports GLOSH, and after fitting, the outlier scores are available through the outlier_scores_ attribute. The README notes that higher scores mean more outlier-like, and that selecting outliers via upper quantiles is often a good approach. Note the design consequence: outlier scores come from the same clustering structure, so changing min_cluster_size changes the outlier ranking too. These are not independent tools.

## Where hdbscan is the wrong tool

The noise label is the first limitation, and it is structural rather than incidental. HDBSCAN will leave points unassigned when they do not fit any density mode. If your downstream system expects every row to have a cluster id, you either need a fallback assignment step or you need a different algorithm. A database that cannot store a null cluster, or a report that groups by cluster, will break on the -1 values.

The second limitation is scale. The README's performance claims are about being faster than a Java reference implementation and faster than single linkage implementations in C and C++, with a specific note that performance on low dimensional data is better than sklearn's DBSCAN. Those claims come from the project's own benchmark notebooks, and they are about relative speed, not about complexity class. The core graph construction remains expensive as the number of samples grows, and the distance matrix input path is quadratic in memory. For very large or very high dimensional datasets, an approximate nearest-neighbour approach is the more realistic starting point.

The third limitation is interpretability of parameters beyond min_cluster_size. The README's selling point is that you barely need to tune anything, but the class exposes more than one knob, and the interaction between them is not something the README walks through. If you change several at once and the clusters move, the documentation will not tell you which one did it. Change one at a time.

## hdbscan against DBSCAN and k-means

The clearest comparison is with DBSCAN, which is also density-based and also produces noise labels. The difference is that DBSCAN uses a single global epsilon, so it can only find clusters of roughly one density. HDBSCAN runs DBSCAN over varying epsilon values and integrates the result, which is why the README says it can find clusters of varying densities and is more tolerant of parameter selection. If your data has a dense core and a sparse but real second group, DBSCAN with one epsilon will either merge them or discard the sparse one. HDBSCAN is designed for exactly that case.

Against k-means the difference is more fundamental. k-means assumes convex, roughly equal-sized clusters and requires you to choose k in advance; it assigns every point, including points that belong nowhere. HDBSCAN makes no convexity assumption, does not need k, and is allowed to reject points. If your clusters are elongated or nested, or if you genuinely do not know how many there are, HDBSCAN is the better starting point. If your clusters are well separated and you need speed and a fixed number of groups, k-means remains simpler to operate.

The package also ships robust single linkage, the algorithm of Chaudhuri and Dasgupta. The README describes this as a high performance version that outperforms scipy's standard single linkage implementation. If your problem is genuinely single-linkage shaped, that entry point exists inside the same package rather than requiring a separate dependency.

## Maintenance, licence and the cost of upgrading

The repository is not archived, and the last push was on 2026-06-12. Releases have been coming at a steady cadence through 2026: release-0.8.42 on 2026-03-27, release-0.8.43 on 2026-05-13, and release-0.8.44 on 2026-06-01. The current version in pyproject.toml is 0.8.44, and the project classifiers still mark it as Development Status 4 - Beta, which is worth knowing if your organisation treats beta classifiers as a procurement signal.

The licence is BSD-3-Clause, matching the LICENSE file and the BSD text in pyproject.toml. That is a permissive licence, so redistribution and modification are allowed with the usual attribution and disclaimer conditions. This is not legal advice; if you are embedding the package in a distributed product, have your own counsel read the LICENSE file rather than relying on the SPDX identifier alone.

Upgrade cost is dominated by the build requirements, not the Python API. The project requires Python 3.10 or later, scikit-learn 1.6 or later, and NumPy below 3. The build system pins Cython to 3.0.11 or later below 4. A NumPy major-version bump on your platform can force a rebuild of the Cython extensions, and a binary wheel built against one NumPy major version is not guaranteed to work against another. Pin your NumPy version in the environment that installs hdbscan, and rebuild rather than assuming an old wheel still loads.

## Conclusion

Adopt hdbscan when you are exploring unlabelled data and can accept a noise label, and when you want cluster stability scores rather than a fixed k. Do not adopt it if your pipeline requires every row to receive a cluster, or if you need to cluster millions of high-dimensional vectors, where approximate nearest-neighbour methods fit better. Before committing, run it on a sample of your own data and check the fraction of points labelled -1 at your chosen min_cluster_size, then read the condensed tree to confirm the clusters are the ones you expected.

## FAQ

### What is HDBSCAN used for?

It is used for density-based clustering of unlabelled data, where clusters may have different densities. The package also provides GLOSH outlier detection and soft cluster membership scores from the same fit.

### Is HDBSCAN better than DBSCAN?

The README states that HDBSCAN performs DBSCAN over varying epsilon values and integrates the result, which lets it find clusters of varying densities and makes it more tolerant of parameter selection. DBSCAN uses a single global epsilon, so it cannot separate groups of different densities.

### What is HDBSCAN in Python?

It is a Python package in the scikit-learn-contrib organisation that implements hierarchical density-based clustering. The README states that it inherits from sklearn classes and drops in next to other sklearn clusterers with an identical calling API.

### How to pip install hdbscan?

Run pip install hdbscan. The package compiles Cython extensions, so the build needs the dependencies declared in pyproject.toml, including Cython 3.0.11 or later below 4 and NumPy 2.0 or later below 3. A conda-forge build is also published.

### How to use hdbscan in Python?

Construct hdbscan.HDBSCAN with a minimum cluster size, then call fit_predict on your data. The README's example uses min_cluster_size=10 and returns an array of labels where -1 marks points treated as noise.

### What does HDBSCAN stand for?

Hierarchical Density-Based Spatial Clustering of Applications with Noise. The README expands the acronym in the opening section of the documentation.

## Sources

- [License: BSD-3-Clause](https://github.com/scikit-learn-contrib/hdbscan/blob/master/LICENSE)
- [Project website](http://hdbscan.readthedocs.io/en/latest/)
- [README](https://github.com/scikit-learn-contrib/hdbscan/blob/master/README.md)
- [Releases](https://github.com/scikit-learn-contrib/hdbscan/releases)
- [scikit-learn-contrib/hdbscan on GitHub](https://github.com/scikit-learn-contrib/hdbscan)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/scikit-learn-contrib-hdbscan
