# metric-learn: scikit-learn-compatible metric learning in Python

> metric-learn packages ten supervised and weakly-supervised metric learning algorithms behind the scikit-learn estimator API, so they drop into pipelines and model selection. The trade-off is a narrow dependency window and a release cadence worth checking before you commit.

**scikit-learn-contrib/metric-learn** — Metric learning algorithms in Python

- Repository: https://github.com/scikit-learn-contrib/metric-learn
- Website: http://contrib.scikit-learn.org/metric-learn/
- Stars: 1,430 · Forks: 232
- Language: Python
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/scikit-learn-contrib-metric-learn

## What metric-learn is for, and who it is for

Most classifiers assume a distance. k-nearest-neighbours, clustering, and any pipeline step that compares rows all behave according to whatever metric you hand them, and Euclidean distance on raw features is rarely the right one. Metric learning fits that distance from labelled data instead of guessing it. metric-learn collects implementations of ten such algorithms, including Large Margin Nearest Neighbor (LMNN), Neighborhood Components Analysis (NCA), Information Theoretic Metric Learning (ITML), Local Fisher Discriminant Analysis (LFDA), Relative Components Analysis (RCA), and Metric Learning for Kernel Regression (MLKR).

The audience is narrow and specific: Python users who already work inside scikit-learn and want a learned Mahalanobis-style distance that behaves like any other estimator. Because metric-learn is part of scikit-learn-contrib and its API follows scikit-learn, the README states that all scikit-learn routines for pipelining and model selection work with these algorithms through a unified interface. If your workflow lives in pandas, numpy and scikit-learn, this is the shortest path. If your features are images and your model is a CNN, it is the wrong library entirely.

## How metric-learn fits the scikit-learn estimator contract

The design choice that matters is that metric-learn does not introduce its own training loop, its own data container, or its own scoring convention. Each algorithm is an estimator that follows the scikit-learn API, which is why the README can claim compatibility for pipelining and model selection. In practice that means the learned object is fitted on feature arrays and a target, and the resulting transformation can be applied to new rows the same way you would apply any other fitted transformer.

The algorithms split by supervision. LMNN, NCA, ITML, LFDA, RCA and MLKR are supervised: they need labels. The README describes the set as covering supervised and weakly-supervised methods, so some entries accept weaker supervision than a full label per row. That distinction matters when you are choosing, because a weakly-supervised algorithm will accept constraints you can actually collect (same/different pairs, for instance) rather than requiring a clean label column.

One implementation detail is worth flagging. For Sparse Determinant Metric Learning (SDML), the README says the optional skggm package lets the algorithm solve problematic cases. That is an admission that the default path in this package does not cover every input for that one algorithm. If SDML is the algorithm you need, plan for the extra dependency rather than discovering it after a failed fit.

## Installing metric-learn and fitting a first metric

Three install routes are documented. The PyPI route is the shortest, and the README also gives a conda-forge command for Anaconda users. The package requires Python 3.6 or later, with numpy>=1.11.0, scipy>=0.17.0 and scikit-learn>=0.21.3. The README notes that the last version supporting Python 2 and Python 3.5 was v0.5.0, so anyone on an old interpreter has to pin backwards.

```bash
pip install metric-learn
```

Anaconda users can install from conda-forge instead:

```bash
conda install -c conda-forge metric-learn
```

For the latest code rather than a release, the README describes a manual install from a downloaded source repository, followed by the test suite:

```bash
python setup.py install
pytest test
```

The second command needs the pytest package present; the README says so explicitly. Running it is the fastest way to confirm that your numpy, scipy and scikit-learn versions are actually compatible with the code you just installed, rather than trusting the floors in the README.

For a first real use, the repository ships worked examples. The examples directory contains plot_metric_learning_examples.py and plot_sandwich.py, and the README points to the sphinx documentation for the API and further examples. Read those two files before writing your own fit call: the README lists algorithm names (LMNN, ITML, NCA and the rest) but does not print the corresponding class names or constructor arguments, and it does not document the shape or meaning of the learned transformation. The examples are where that information lives. Running the examples requires matplotlib, which the README marks as an optional dependency for examples only.

## Where metric-learn stops being the right tool

The clearest boundary is deep learning. The README describes metric-learn as efficient Python implementations of metric learning algorithms, with no mention of neural networks, GPU execution, or integration with PyTorch or TensorFlow. If your metric has to be learned jointly with an encoder over images, audio or text, this package cannot express that model. People searching for deep metric learning or PyTorch metric learning are looking at a different problem, and metric-learn is not a substitute for those frameworks.

The second boundary is the dependency window. The README pins scikit-learn>=0.21.3 as a floor, but floors are not ceilings, and the README does not state which scikit-learn versions have been validated. A codebase whose newest listed release is v0.7.0 from 2023-09-29 may or may not keep pace with scikit-learn's own deprecation cycle. The repository's last push was 2026-03-19, so there has been activity since that release, but the README does not document a compatibility matrix, and it does not document a rollback procedure if an upgrade breaks a fitted pipeline.

The third boundary is SDML specifically. Because the README ties its handling of problematic cases to an optional skggm install from a pinned commit, SDML is the algorithm most likely to fail on awkward data unless you add that dependency. That is a real constraint, not a footnote.

## metric-learn compared with a deep metric learning framework

The honest alternative is a deep metric learning library. The difference is architectural, not a matter of which one is better. A deep framework learns a metric implicitly: you define an embedding network and a loss over pairs or triplets, and the distance emerges from the learned embedding. metric-learn learns an explicit transformation over the feature vectors you already have. That means metric-learn has no encoder to train, no GPU requirement stated anywhere in the README, and no images to feed it.

That difference decides the choice. If your data is tabular or already vectorised into fixed-length features, metric-learn is the smaller, more direct tool, and it plugs into scikit-learn model selection so you can tune it with the same tooling you use elsewhere. If your data is raw pixels or text, the explicit-transformation approach has nothing to transform until something else extracts features, and you should be in a deep framework.

Within the scikit-learn world there is also a subtler alternative: skip metric learning and tune a kernel or use a tree ensemble that does not depend on a global distance. That works when your problem is prediction rather than retrieval or clustering. Metric learning earns its place when the distance itself is the deliverable, for example in nearest-neighbour retrieval, clustering, or kernel regression, which is why MLKR exists in this package at all.

## Release cadence, licence and what an upgrade costs

The release list is the most important maintenance fact here. The newest release shown is v0.7.0, dated 2023-09-29. Before that, v0.6.2 and v0.6.1 both landed on 2020-07-02. That is a long gap between the 0.6 line and 0.7, and it means the release history alone will not tell you how much has changed since. The repository's last push was 2026-03-19, so commits have continued after v0.7.0, but the README does not describe what is on master beyond the released version, and it does not document a changelog or migration notes for the 0.6 to 0.7 step.

That shapes upgrade cost. If you install from PyPI you get a released version and a stable target. If you install from source, as the README's manual instructions describe, you are on a moving branch with no documented release boundary, and you are responsible for running pytest test after each pull. For a library that sits inside a scikit-learn pipeline, that is a meaningful difference: a broken fit surfaces at training time, not at import time.

The licence is MIT, stated in the repository and shown in the README badge. MIT is permissive: it allows commercial and closed-source use, and it requires that the licence text and copyright notice be preserved. That is a summary of the identifier, not legal advice; if your organisation has licence review, the LICENSE.txt file in the repository root is the document to read.

One more cost that is easy to overlook. The README asks that scientific publications cite the JMLR paper by de Vazelhes et al., volume 21, number 138, pages 1 to 6, 2020, and it provides a BibTeX entry. That is a request, not a licence condition, but it is a real obligation if you publish results built on this library.

## Conclusion

Adopt metric-learn if you already build scikit-learn pipelines and want LMNN, NCA, ITML or LFDA without leaving that API, and you can live with the dependency floor the README states (Python 3.6+, numpy>=1.11.0, scipy>=0.17.0, scikit-learn>=0.21.3). Do not adopt it if you need a metric learned by a neural encoder over images or text; that work belongs in a deep metric learning framework, and this package has no neural components. Before you build on it, check the release history: the newest release listed is v0.7.0 from 2023-09-29, while the repository's last push was 2026-03-19, so confirm whether the fixes you need are in a release or only on master. Then verify on your own data which of the ten algorithms is exposed under the name you expect, because the README lists the algorithms but not the class names.

## FAQ

### What is metric learning and how does metric-learn implement it?

Metric learning fits a distance function from data instead of assuming one such as Euclidean distance. metric-learn provides Python implementations of ten such algorithms, including LMNN, NCA, ITML and LFDA, behind an API compatible with scikit-learn so they can be used in pipelines and model selection.

### How do I install metric-learn?

The README gives three routes: pip install metric-learn, conda install -c conda-forge metric-learn for Anaconda users, or a manual install from a downloaded source repository with python setup.py install. It requires Python 3.6 or later, numpy>=1.11.0, scipy>=0.17.0 and scikit-learn>=0.21.3.

### Which algorithms does metric-learn include?

The README lists Large Margin Nearest Neighbor (LMNN), Information Theoretic Metric Learning (ITML), Sparse Determinant Metric Learning (SDML), Least Squares Metric Learning (LSML), Sparse Compositional Metric Learning (SCML), Neighborhood Components Analysis (NCA), Local Fisher Discriminant Analysis (LFDA), Relative Components Analysis (RCA), Metric Learning for Kernel Regression (MLKR) and Mahalanobis Metric for Clustering (MMC).

### Does metric-learn work with scikit-learn pipelines?

Yes. The README states that because metric-learn is part of scikit-learn-contrib, its API is compatible with scikit-learn, which allows all scikit-learn routines for pipelining and model selection to be used with these algorithms through a unified interface.

### What does SDML need beyond the base install?

The README says that for SDML, using skggm allows the algorithm to solve problematic cases, and it gives a pip command to install a specific pinned commit of skggm from GitHub. That makes SDML the algorithm most likely to need an extra dependency.

### What licence does metric-learn use?

The repository and README badge identify the licence as MIT. The README also asks that scientific publications cite the JMLR paper by de Vazelhes et al., volume 21, number 138, 2020, and provides a BibTeX entry.

## Sources

- [License: MIT](https://github.com/scikit-learn-contrib/metric-learn/blob/master/LICENSE)
- [Project website](http://contrib.scikit-learn.org/metric-learn/)
- [README](https://github.com/scikit-learn-contrib/metric-learn/blob/master/README.md)
- [Releases](https://github.com/scikit-learn-contrib/metric-learn/releases)
- [scikit-learn-contrib/metric-learn on GitHub](https://github.com/scikit-learn-contrib/metric-learn)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/scikit-learn-contrib-metric-learn
