Model or dataset
shenweichen/DeepCTR avatar
shenweichen/DeepCTR

DeepCTR: A TensorFlow Library for Click-Through Rate Prediction Models

Easy-to-use,Modular and Extendible package of deep-learning based CTR models .

8,050 stars2,214 forksPythonApache-2.0

At a glance

What is it?
DeepCTR is a Python package that provides modular, easy-to-use implementations of deep learning-based click-through rate prediction models with a Keras-like interface, covering architectures from DeepFM and Wide and Deep to xDeepFM and AutoInt. It targets researchers and engineers who want to experiment with published CTR architectures without reimplementing them from scratch, but its TensorFlow dependency and Python version constraints require attention before adopting it.
Who is it for?
DeepCTR is a good starting point for a data scientist or ML engineer who wants to run published CTR architectures against a dataset without reimplementing each paper. It is less suitable as a foundation for production recommendation systems that need PyTorch, or for teams on Python versions below 3.7 or above 3.12.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 90 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What DeepCTR Solves and Who It Is For

Click-through rate prediction is a core problem in online advertising and recommendation systems: given a user and a candidate item, estimate the probability that the user clicks. Dozens of deep learning architectures have been proposed for this problem at venues like KDD, IJCAI, and SIGIR. DeepCTR collects these architectures in a single Python package with a consistent interface, so a researcher can swap between DeepFM, Wide and Deep, DCN, and DIN without rewriting data pipelines or training loops.

The README describes DeepCTR as easy-to-use, modular, and extendible, with a package of deep learning-based CTR models and core component layers that can be used to build custom models. The Keras-like interface means any model in the package can be trained with `model.fit()` and evaluated with `model.predict()`, which aligns with the workflow most Python ML practitioners already know.

The target audience is data scientists and ML engineers at companies that run ad systems or recommender systems, and researchers who want a reliable baseline implementation of a published CTR architecture without spending time reading and debugging the original paper's code.

Models in the Package and Their Origins

The README lists over twenty model architectures, each traced to its original conference paper. The range covers a decade of CTR research. A representative sample:

Wide and Deep (DLRS 2016) learns a wide linear model alongside a deep neural network. DeepFM (IJCAI 2017) replaces the wide component with a factorization machine. xDeepFM (KDD 2018) adds an explicit feature interaction component called the Compressed Interaction Network. Deep and Cross Network (ADKDD 2017) introduces cross layers for automatic feature crossing. Deep Interest Network (KDD 2018) adds an attention mechanism over the user's browsing history to weight relevant past behaviors. AutoInt (CIKM 2019) applies a self-attention mechanism for automatic feature interaction.

Each model is listed with its source paper, venue, and year. This traceability is a deliberate design choice: if a user finds that a model behaves unexpectedly, they can go back to the paper and check the original formulation. The modular architecture also means that individual layers from these models can be reused in custom architectures, which is the extendibility the README references.

Installing DeepCTR and Running a First Model

DeepCTR does not install TensorFlow for you. The README is explicit: install a TensorFlow build that matches your Python, NumPy, CPU or GPU, and operating system first, then install DeepCTR:

bash
pip install tensorflow
pip install deepctr

For Python 3.9 and above, the README notes that modern h5py releases are supported with h5py version 3.7.0 or newer. The setup.py confirms: Python 3.7 or newer is required, and the package is tested on 3.7, 3.10, 3.11, and 3.12.

The README also warns about a common integration mistake: avoid mixing `tensorflow.python.keras` with `tensorflow.keras` in the same codebase. The `tensorflow.python.*` namespace is a private TensorFlow API and can break model serialization or optimizer and metric loading across TensorFlow versions. Use public `tensorflow.keras` APIs exclusively.

The examples/ directory in the repository contains working scripts for standard datasets. The `run_classification_criteo.py` script trains a model on the Criteo sample dataset. The `run_estimator_tfrecord_classification.py` script demonstrates the TFRecord-based Estimator interface for large-scale data. Running these scripts requires downloading or generating the corresponding data files, several of which are included in the examples/ directory as sample fragments.

The Keras Interface vs the Estimator Interface

DeepCTR offers two usage modes. The Keras interface wraps each model as a `tf.keras.Model`, providing the standard `model.compile()`, `model.fit()`, and `model.predict()` workflow. This is the straightforward path for experimentation: load a DataFrame, define the feature columns, instantiate a model, and call fit.

The Estimator interface targets large-scale data and distributed training. The README points to a separate quick-start guide for the Estimator path, which uses TFRecord files as input. The `run_estimator_tfrecord_classification.py` example in the examples/ directory shows the TFRecord-based setup. The Estimator interface requires more upfront configuration but integrates with TensorFlow's distributed training infrastructure.

For most research use cases, the Keras interface is the starting point. The Estimator interface becomes relevant when the training dataset does not fit in memory or when multi-GPU or multi-machine training is needed.

The package is compatible with both TensorFlow 1.15 and TensorFlow 2.x, which the README states explicitly. This compatibility range is wider than most TensorFlow-dependent projects maintain, which is useful for teams working with older infrastructure.

Related Projects: DeepCTR-Torch and DeepMatch

The README lists two related projects from the same author. DeepCTR-Torch, at github.com/shenweichen/DeepCTR-Torch, is the PyTorch equivalent of DeepCTR. It provides the same model architectures with a PyTorch backend, targeting teams that have standardized on PyTorch rather than TensorFlow.

DeepMatch, at github.com/shenweichen/DeepMatch, covers a different but related problem: candidate item retrieval or matching, which is the first stage in a two-stage recommendation pipeline. Where DeepCTR handles ranking (which item to click), DeepMatch handles retrieval (which items are candidates). The two projects are complementary in a full recommendation system.

A team building a full recommendation pipeline in TensorFlow would typically use DeepMatch for retrieval and DeepCTR for the ranking stage. A team on PyTorch would use DeepCTR-Torch for ranking. The separation keeps each package focused on its specific problem.

Limitations: TensorFlow Dependency and Compatibility Risks

DeepCTR's central limitation is its TensorFlow dependency. TensorFlow has had significant API changes across versions, particularly around Keras, and the intersection of TensorFlow version, Python version, NumPy version, and GPU driver compatibility is a known source of environment setup friction.

The README warns specifically about NumPy conflicts: if TensorFlow reports a NumPy conflict, follow the TensorFlow requirement for the selected TensorFlow release, for example using `numpy<2` when required by TensorFlow. This kind of version pinning requires attention when updating other packages in the same environment.

The package's last release is v0.9.4, released on 2026-04-16, with a last push to the master branch on 2026-07-02. The v0.9.3 release was in November 2022, nearly four years before v0.9.4, which indicates the pace of releases is not frequent. Teams adopting DeepCTR for production use should check the changelog for the specific models they plan to use and be prepared to maintain any custom modifications without relying on frequent upstream updates.

The package does not support JAX or any framework outside TensorFlow. Teams committed to PyTorch need to use DeepCTR-Torch instead.

License and Maintenance

DeepCTR is licensed under the Apache License 2.0, which permits use, modification, and distribution in both open-source and proprietary applications. There are no copyleft obligations: a company can use DeepCTR in a commercial product without releasing their application code. The Apache-2.0 license also includes a patent grant, which is relevant if any of the implemented architectures involve patent-covered techniques.

The project is maintained by Weichen Shen, with author email [email protected] listed in setup.py. The repository includes a contributing guide and a code quality badge from Codacy. CI runs on GitHub Actions, with coverage tracked via Codecov. The .readthedocs.yml file in the repository indicates that the documentation at deepctr-doc.readthedocs.io is built from the repo source.

Editorial conclusion

DeepCTR is a good starting point for a data scientist or ML engineer who wants to run published CTR architectures against a dataset without reimplementing each paper. It is less suitable as a foundation for production recommendation systems that need PyTorch, or for teams on Python versions below 3.7 or above 3.12. Before adopting it, verify TensorFlow compatibility with your NumPy version, check whether the model you need is in the models list, and decide whether the Keras or Estimator interface better fits your serving pipeline. The Apache-2.0 license allows use in proprietary applications without any copyleft obligation.

Frequently asked questions

Does DeepCTR work with TensorFlow 2.x?

Yes. The README states that DeepCTR is compatible with both TensorFlow 1.15 and TensorFlow 2.x.

Is there a PyTorch version of DeepCTR?

Yes. The README lists DeepCTR-Torch at github.com/shenweichen/DeepCTR-Torch as the PyTorch equivalent of DeepCTR.

What Python versions does DeepCTR support?

The setup.py lists Python 3.7 and newer as the requirement, with the package tested on 3.7, 3.10, 3.11, and 3.12.

Official sources

  1. License: Apache-2.0
  2. Project website
  3. README
  4. Releases
  5. shenweichen/DeepCTR on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/shenweichen-deepctr.svg)](https://hysenlabs.com/projects/shenweichen-deepctr)