# DeepMatch: a TensorFlow library for training two-tower retrieval models and exporting embeddings

> DeepMatch collects seven published matching architectures (FM, DSSM, YoutubeDNN, NCF, SDM, MIND, ComiRec) behind a Keras-style fit/predict interface, so the output is a user and item vector you can index for ANN search. It is a research-oriented library with a narrow dependency story, and its release history is uneven.

**shenweichen/DeepMatch** — A deep matching model library for recommendations & advertising. It's easy to train models and to export representation vectors which can be used for ANN search.

- Repository: https://github.com/shenweichen/DeepMatch
- Website: https://deepmatch.readthedocs.io/en/latest/
- Stars: 2,436 · Forks: 541
- Language: Python
- License: Apache-2.0
- Published: 2026-09-28 · Updated: 2026-09-28 · Language: en
- Canonical page: https://hysenlabs.com/projects/shenweichen-deepmatch

## The retrieval stage problem DeepMatch targets

A recommender usually splits into two stages. A matching stage narrows a large item catalogue down to a few hundred candidates, and a ranking stage scores those candidates with heavier features. DeepMatch only addresses the first stage. The README describes it as a deep matching model library for recommendations and advertising, and the two things it promises are training a model and exporting representation vectors for users and items that can be used for ANN search.

The intended user is someone who has a candidate-generation problem and wants to try a published neural architecture rather than write one from scratch. The models table lists FM, DSSM, YoutubeDNN, NCF, SDM, MIND and COMIREC, each with the paper it comes from. That is a specific kind of value: the library is a curated set of reference implementations, not a framework with its own abstractions to learn. If your problem is ranking with dense cross features, DeepMatch is the wrong stage of the pipeline. The author maintains a separate project, DeepCTR, for that, and DeepMatch depends on it.

## How the fit and predict interface produces embeddings

The design choice that matters is that DeepMatch models behave like Keras models. The README states you can use any complex model with model.fit() and model.predict(). That means the training loop, callbacks, optimizers and metrics come from TensorFlow rather than from DeepMatch, and the library's job is to assemble the network and expose the right outputs.

The second half of the contract is the part that distinguishes it from a plain classifier. The library exports representation vectors for users and items, which is the artifact an approximate nearest neighbour index consumes. In practice this means a trained model is queried twice: once with user-side features to produce a user vector, once with item-side features to produce an item vector, and the two live in a shared space where inner product or cosine distance is meaningful. The repository layout supports this reading. There is a deepmatch package for the models and layers, a tests directory with tests.models and tests.layers subpackages, and an examples directory containing one runnable script per architecture: run_dssm_negsampling.py, run_dssm_inbatchsoftmax.py, run_youtubednn.py, run_sdm.py and run_ncf.py, alongside Colab notebooks for the same models on MovieLens 1M. Two DSSM scripts rather than one is the clearest signal that the negative sampling strategy is a first-class choice rather than a detail buried in a config.

## Installing DeepMatch and running a first model

The README is explicit that DeepMatch does not pin or install TensorFlow for you. You install a TensorFlow build that matches your Python, NumPy, CPU or GPU setup and operating system first, then install DeepMatch. The order is not cosmetic: setup.py declares only requests and deepctr~=0.9.4 as required packages, so pip will not pull a TensorFlow wheel on your behalf.

```bash
pip install tensorflow
pip install deepmatch
```

After that, the README points to the Quick Start page in the documentation and to the Colab notebooks under examples/ for a worked run. The example scripts are the more useful starting point because they run end to end from a terminal. They read examples/movielens_sample.txt, a small sample of MovieLens data, and examples/preprocess.py handles turning it into the input format the models expect. The README itself does not spell out the command line for an individual script, so check the script you intend to run before executing it.

Two compatibility notes from the README are worth applying before you write your own code. For Python 3.9 and above, DeepMatch and its dependencies allow h5py>=3.7.0. If TensorFlow reports a NumPy conflict, follow the NumPy requirement of the TensorFlow release you chose, for example numpy<2 where that release requires it. And use public tensorflow.keras APIs in your own code. The README warns against mixing tensorflow.python.keras with tensorflow.keras, because tensorflow.python.* is private TensorFlow API and can break model serialization or optimizer and metric loading across TensorFlow versions.

## Where DeepMatch will cost you time

The dependency story is the first real constraint. Because TensorFlow is installed separately and deepctr is pinned with a compatible-release operator at 0.9.4, the combination of Python version, TensorFlow version and NumPy version is yours to get right. The README acknowledges this by giving a NumPy fallback rather than a compatibility matrix, which means troubleshooting a broken environment is part of onboarding.

The second constraint is the release cadence. The most recent release listed is v0.3.2 on 2026-04-18, and the one before it was v0.3.1 on 2022-10-31. A gap of that length between releases tells you what to expect from bug reports and feature requests. The project is not archived, and the last push to the default branch was on 2026-04-18, so the repository is not frozen. But a library whose releases cluster around a few dates is one where you should read the code before you depend on an edge case.

The third is scope. DeepMatch trains and exports. The README does not describe an index, a serving component, or a feature store, so the ANN search it mentions is something you bring. If you need an end-to-end retrieval service rather than a model that emits vectors, you are assembling the rest yourself.

## DeepMatch against a general-purpose recommender framework

The natural comparison is with a broad recommender framework such as DeepCTR-Torch, which reimplements the DeepCTR family on PyTorch. The difference is not just the backend. DeepMatch is organized around the matching stage specifically: its model list is short and every entry is a retrieval architecture, and its documented output is a pair of embedding vectors for ANN search. A general framework tends to offer a wider catalogue of ranking and prediction models and leaves candidate generation to you.

That makes the choice mostly about where your bottleneck is. If you already have a ranking model and the problem is that scoring the full catalogue is too slow, a library whose stated purpose is exporting vectors for ANN search maps onto your problem directly. If your problem is that the ranking model is not accurate enough, nothing in DeepMatch's model list addresses it. The other axis is backend: the README's installation instructions and its warning about tensorflow.keras are TensorFlow-specific, so a PyTorch shop gets no benefit from the fit and predict interface.

## Maintenance, licence and what an upgrade costs

DeepMatch is Apache-2.0, as stated in setup.py and the LICENSE file at the repository root. That is a permissive licence, and the practical consequence is that you can embed the library in a commercial system. It is not a copyleft licence, so it does not require you to publish your own code. This is a description of the licence text, not legal advice; if the licence terms matter to your organisation, have someone qualified read them.

The upgrade cost is dominated by TensorFlow rather than by DeepMatch. The library's own version moved from 0.3.0 to 0.3.1 to 0.3.2, while the TensorFlow version you install is chosen independently and the README's guidance on h5py and NumPy exists precisely because those versions interact. The README's advice to stay on public tensorflow.keras APIs is the cheapest insurance available: code that avoids tensorflow.python.* is less likely to break when you move TensorFlow versions, because serialization and optimizer loading go through the public path. If you build on DeepMatch, treat the TensorFlow version as a pinned part of your environment and test the upgrade separately from any DeepMatch change.

## Conclusion

Adopt DeepMatch if you already run TensorFlow and want a reference implementation of a published matching model to compare against your own, with embedding export built into the same object you train. Do not adopt it if your stack is PyTorch, if you need a maintained serving path, or if you expect the library to resolve your TensorFlow and NumPy versions for you. Before committing, check that the model you want appears in the models table in the README, read the matching example script under examples/ for the input format it expects, and confirm that the TensorFlow build you intend to use is compatible with the deepctr version pinned in setup.py.

## FAQ

### Does DeepMatch install TensorFlow for me?

No. The README states that DeepMatch does not pin or install TensorFlow, and setup.py lists only requests and deepctr~=0.9.4 as required packages. You install a TensorFlow build matching your Python, NumPy and hardware first, then install DeepMatch.

### Which matching models does DeepMatch include?

The README's models table lists FM, DSSM, YoutubeDNN, NCF, SDM, MIND and COMIREC, each with its source paper. The examples directory has a runnable script or Colab notebook for most of them.

### What do I get out of a trained DeepMatch model?

The README describes the library as easy to train models and to export representation vectors for user and item which can be used for ANN search. The vectors are the artifact you feed into a nearest neighbour index; DeepMatch does not provide the index itself.

### What licence is DeepMatch released under?

Apache-2.0, according to setup.py and the LICENSE file at the repository root. That is a permissive licence rather than a copyleft one.

### Why does DeepMatch warn against tensorflow.python.keras?

The README says tensorflow.python.* is private TensorFlow API and can break model serialization or optimizer and metric loading across TensorFlow versions. It advises using public tensorflow.keras APIs in your own code and examples.

## Sources

- [License: Apache-2.0](https://github.com/shenweichen/DeepMatch/blob/master/LICENSE)
- [Project website](https://deepmatch.readthedocs.io/en/latest/)
- [README](https://github.com/shenweichen/DeepMatch/blob/master/README.md)
- [Releases](https://github.com/shenweichen/DeepMatch/releases)
- [shenweichen/DeepMatch on GitHub](https://github.com/shenweichen/DeepMatch)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/shenweichen-deepmatch
