Model or dataset
uber/causalml avatar
uber/causalml

causalml: Uber's Python package for uplift modeling and CATE estimation

Uplift modeling and causal inference with machine learning algorithms

6,000 stars876 forksPythonNOASSERTION

At a glance

What is it?
causalml estimates conditional average treatment effects from experiments or observational data. It is a strong fit for campaign targeting work, but the estimator families differ enough that picking one is the hard part.
Who is it for?
Adopt causalml if you already run experiments or hold observational data with a plausible treatment assignment, and you need per-user effect estimates rather than a single average lift. Skip it if you only need an average treatment effect, since statsmodels or a simple difference in means answers that with far less machinery.
Can I use it commercially?
Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
Is it still maintained?
Yes. The repository last received commits 40 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 17, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The problem causalml solves, and the people who have it

A/B tests tell you whether a treatment worked on average. They do not tell you which users to treat next time. That gap is where causalml lives. The README frames the package around estimating the Conditional Average Treatment Effect (CATE) of an intervention T on an outcome Y for users with observed features X, "without strong assumptions on the model form." Two use cases are named explicitly: campaign targeting optimization, where the goal is to find customers who respond favorably to an ad, and personalized engagement, where several treatment options compete and you want the best one per customer.

The audience is narrower than the topic suggests. If your question is "did the campaign lift sales overall," you do not need this library. causalml earns its place when the decision is per-unit allocation: which segment gets the ad, which message goes to which user, where a fixed budget produces the most incremental conversions. The repository sits under the uber organization and is described in its own disclaimer as "stable and being incubated for long-term support," with the caveat that it "may contain new experimental code, for which APIs are subject to change." That sentence matters more than most disclaimers: parts of the API surface are explicitly not frozen.

How causalml estimates a treatment effect: meta-learners, trees and neural backends

The package is organized as a family of estimators behind a common interface, so the same fit and predict pattern applies whether you choose a meta-learner, a tree model or a neural one. The repository layout shows this directly: causalml/inference/tree holds the tree-based implementations, and the Cython sources listed in setup.py reveal the internal structure, with separate modules for _criterion, _splitter and _builder, plus distinct groups for causal and uplift variants. Those compiled extensions are why installation is not a pure-Python pip install of source.

The meta-learner approach is the conceptual core. Rather than modeling the outcome directly, these estimators reframe the problem so that an ordinary supervised learner can be reused. The documentation's "Choosing an Estimator" page is presented as a decision path and capability matrix over the estimators, which is an admission that the choice is not obvious. The tutorial page goes further and trains one estimator per family, then shows how to decide which one to believe. That framing is honest about the discipline's central difficulty: there is no ground truth in real data, so validation has to come from the benchmark datasets (LaLonde, IHDP, Twins) and the metrics they support, named in the README as PEHE, ATE error and policy risk.

Optional backends extend the same interface to neural methods. The pyproject.toml defines extras for tensorflow, torch (with pyro-ppl) and jax (with flax, optax and orbax-checkpoint), so a neural estimator is installed by adding an extra rather than by changing the API.

Installing causalml and running a first estimate

The README does not inline installation steps; it points to the installation page in the ReadTheDocs documentation. The package metadata is more concrete. pyproject.toml declares requires-python >= 3.11, so an older interpreter will not work, and the dependency list pins some versions tightly (forestci==0.6, pathos==0.2.9) while allowing ranges on others (scikit-learn>=1.6.0, statsmodels>=0.14.5).

Because the tree estimators are Cython, the Makefile exposes an explicit build path. The build_ext target runs the extension build in place, and install depends on it:

bash
make build_ext
pip install .

Running make build_ext first compiles the .pyx sources listed in setup.py, including causalml/inference/tree/_tree/_criterion.pyx and the uplift and causal variants. If you skip this and the wheels for your platform are unavailable, imports of the tree module will fail rather than degrade gracefully.

For a first real use, the documentation's quickstart page is the intended entry point, and the tutorial is the one to read next because it compares estimator families rather than presenting a single recipe. The benchmark datasets are the safest place to start, since the README states they ship with SHA256-verified downloads and ground-truth metrics. A typical flow loads one of those datasets, fits an estimator, and scores it with PEHE or policy risk instead of guessing whether the model is right. The leaderboard notebook is documented as regenerating every published number end to end, which gives you a way to confirm your environment reproduces known results before you point the library at production data.

Where causalml is the wrong tool

The clearest limitation is the one the project states about itself. The disclaimer warns that experimental code may exist and that APIs are subject to change. Anyone building a long-lived pipeline on an estimator marked experimental should expect to revisit that code on upgrade.

A second constraint is environmental. The Cython extensions mean the package is not installable as pure Python everywhere, and the build path runs through setup.py. Teams on managed platforms that forbid compilation at install time will need to confirm wheel availability before planning around causalml.

The deeper limitation is statistical, not technical. CATE estimation from observational data depends on assumptions the library cannot check for you. The documentation's own structure reflects this: it devotes a page to choosing an estimator and a tutorial to deciding which estimate to believe, rather than presenting one default. If your treatment assignment is confounded in ways your features do not capture, no estimator in the suite will recover the true effect. Similarly, if the business decision only needs an average effect, the whole estimator-selection problem is unnecessary overhead. And if your sample is small, the heterogeneous estimates will be noisy in ways the benchmark metrics on large public datasets will not reveal.

causalml vs econml and other causal inference libraries

The most direct comparison is with EconML, which appears in causalml's own reference list as the co-subject of a KDD 2021 tutorial titled "Causal Inference and Machine Learning in Practice with EconML and CausalML." The two libraries overlap in purpose but differ in emphasis. causalml's README leads with uplift modeling and campaign targeting, and its benchmark suite is organized around uplift-style metrics: PEHE, ATE error and policy risk. EconML, by the Microsoft Research lineage, is generally described in terms of its orthogonal and doubly-robust estimation methods. The practical difference for a reader is which estimator families are first-class in each package and which metrics the maintainers treat as the scoreboard.

DoWhy occupies a different position again. Where causalml focuses on estimating an effect once you have decided what to estimate, DoWhy is built around stating a causal graph and testing assumptions. The two are complementary rather than competing: DoWhy-style reasoning about identification, then causalml-style estimation. If your problem is "I am not sure this comparison is even valid," causalml is not the first tool to reach for. If your problem is "I know the comparison is valid and I need per-user effects at scale," it is.

Maintenance, versioning and licence terms

The repository is not archived, and the last push was on 2026-08-20. Release cadence is visible in the tags: v0.17.0 on 2026-07-04, v0.16.0 on 2026-02-06, and v0.15.5 on 2025-07-09. That is roughly two releases a year, with the most recent arriving about six weeks after the prior minor version. The changelog is kept in docs/changelog.rst, and the README points to it as the place where versions and changes are documented, so upgrade planning has a written record to check.

The repository carries governance files that affect how changes land: GOVERNANCE.md, CHARTER.md, MAINTAINERS.md and STEERING_COMMITTEE.md, alongside ANTITRUST.md and TRADEMARKS.md. For an adopter, the practical implication is that the project has a defined decision process rather than a single maintainer.

On licensing, the README states the project is licensed under the Apache 2.0 License and points to the LICENSE file, and pyproject.toml carries the classifier "License :: OSI Approved :: Apache Software License." The repository metadata reports the licence as NOASSERTION, which is a classification artifact rather than a contradiction, but it is worth confirming against the LICENSE file if licence terms are a gating factor in your organisation. The Apache 2.0 text includes a patent grant and requires preservation of notices; how those terms apply to your distribution is a question for your own counsel, not something this article can settle.

Editorial conclusion

Adopt causalml if you already run experiments or hold observational data with a plausible treatment assignment, and you need per-user effect estimates rather than a single average lift. Skip it if you only need an average treatment effect, since statsmodels or a simple difference in means answers that with far less machinery. Before committing, verify that your environment satisfies requires-python >= 3.11 and that the Cython extensions build, because causalml.inference.tree ships .pyx sources compiled through setup.py rather than pure Python. Then run the tutorial notebook's per-family comparison on your own data before trusting any single estimator.

Frequently asked questions

What is causalml in machine learning?

causalml is a Python package that provides uplift modeling and causal inference methods built on machine learning algorithms. It estimates the Conditional Average Treatment Effect of an intervention on an outcome for users with observed features, and the README names campaign targeting optimization and personalized engagement as typical use cases.

How does causalml compare with DoWhy?

The two address different stages of the problem. DoWhy is not covered in causalml's own documentation, but causalml's focus is estimating a treatment effect once you have decided what to estimate, with benchmark metrics such as PEHE, ATE error and policy risk. If your concern is whether the comparison is valid in the first place, causalml is not the tool for that question.

How is causal ml different from traditional machine learning?

Traditional supervised learning predicts an outcome from features. causalml instead estimates the effect of an intervention T on outcome Y for users with features X, per the README, which is a different target than a prediction. The package documents benchmark metrics such as PEHE and policy risk precisely because accuracy on outcomes does not measure whether an effect estimate is correct.

What are the main three types of ML models?

causalml's documentation does not present a general taxonomy of machine learning models. It is organized around estimator families for treatment effect estimation, with a decision path and capability matrix on its "Choosing an Estimator" page.

Can you explain causal inference in a simple way?

causalml's own framing is that it estimates the causal impact of an intervention T on an outcome Y for users with observed features X, from experimental or observational data, without strong assumptions on the model form. The README's two named use cases are choosing which customers to target with an ad and choosing which treatment option suits each customer.

Is causal inference still relevant?

causalml's last push was on 2026-08-20 and v0.17.0 was released on 2026-07-04. The project's reference list also includes causal inference and machine learning workshops held at KDD in 2023, 2024 and 2025.

Official sources

  1. Issues
  2. README
  3. Releases
  4. uber/causalml on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/uber-causalml.svg)](https://hysenlabs.com/projects/uber-causalml)