Library / SDK
py-why/EconML avatar
py-why/EconML

EconML: heterogeneous treatment effects with machine learning that stays causal

ALICE (Automated Learning and Intelligence for Causation and Economics) is a Microsoft Research project aimed at applying Artificial Intelligence concepts to economic decision making. One of its goals is to build a toolkit that combines state-of-the-art machine learning techniques with econometrics in order to bring automation to complex causal inference problems. To date, the ALICE Python SDK (econml) implements orthogonal machine learning algorithms such as the double machine learning work of Chernozhukov et al. This toolkit is designed to measure the causal effect of some treatment variable(s) t on an outcome variable y, controlling for a set of features x.

4,800 stars828 forksJupyter NotebookNOASSERTION

At a glance

What is it?
EconML is the ALICE project's Python package from Microsoft Research, estimating heterogeneous treatment effects from observational data by combining econometrics with machine learning, implementing methods like double machine learning with a unified API, flexible effect models from forests to neural nets, valid confidence intervals, and policy learning, now under the PyWhy umbrella with v0.17.0 current.
Who is it for?
Use EconML when the question is what causal effect a treatment has on an outcome, and how that effect varies across a population, from observational data, provided you can defend the identifying assumptions, no unobserved confounders for most methods or a valid instrument Z for the others. It is not a substitute for experimental design where experiments are feasible.
Can I use it commercially?
Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
Is it still maintained?
Yes. The repository last received commits 5 days ago.
What is it written in?
Mainly Jupyter Notebook, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 27, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The problem statement, in variables

EconML exists to measure the causal effect of some treatment variables T on an outcome variable Y, controlling for a set of features X, W, and how that effect varies as a function of X, the heterogeneous part, since the same intervention helps different people differently. The methods apply even with observational, non-experimental or historical datasets, and the package is explicit about what that costs, some methods assume no unobserved confounders, no variable outside X and W affecting both treatment and outcome, while others assume access to an instrument Z that moves the treatment but has no direct effect on the outcome. Most methods provide confidence intervals and inference results, so the output is an effect estimate with uncertainty rather than a bare prediction, and that epistemic framing separates it from ordinary supervised learning. That structure, treatment, outcome, controls, and effect as a function of features, is the API users meet directly, and the documentation at pywhy.org walks each estimator through the same variables so a learned pattern transfers across methods.

ALICE, Microsoft Research, and PyWhy

The package was designed and built as part of the ALICE project, Automated Learning and Intelligence for Causation and Economics, at Microsoft Research, aimed at applying artificial intelligence concepts to economic decision making, and its four promises are stated in the documentation. Implement recent techniques from the literature at the intersection of econometrics and machine learning, maintain flexibility in modeling effect heterogeneity through random forests, boosting, lasso and neural nets while preserving the causal interpretation of the learned model, use a unified API, and build on standard Python packages for machine learning and data analysis. The project now lives under the PyWhy organization with PyWhy contributors listed as authors, a governance move that outlived the original research project framing, and the package carries the MIT license with Python 3.9 through 3.14 support.

What heterogeneous treatment effect estimation is

The package's own framing answers the question directly. One of the biggest promises of machine learning is to automate decision making in a multitude of domains, and at the core of many data-driven personalized decision scenarios is the estimation of heterogeneous treatment effects, what is the causal effect of an intervention on an outcome of interest for a sample with a particular set of features. The estimation is performed by methods that model the effect as a function of the features X rather than as a single average number, using machine learning for the nuisance components while preserving identification through the econometric structure, and the toolkit's estimate surfaces that function with confidence intervals, so a decision maker sees both the personalized effect and its uncertainty rather than a point value alone.

Double machine learning's contribution

The package's description names the double machine learning work of Chernozhukov et al. as the archetype of its orthogonal machine learning algorithms, and the contribution of that method to causal inference is precisely the orthogonalization idea. Machine learning predicts the treatment and the outcome from the control features, the residuals of those predictions carry the variation that identifies the effect, and the two-stage construction removes the bias that flexible ML models would otherwise inject into the estimate. EconML implements this family alongside related orthogonal methods, which is why its description leads with the term, and why the promise list pairs modern ML flexibility with causal interpretation as one requirement rather than two, the combination double machine learning made practical.

A dependency comment worth reading

The package metadata contains an unusual comment block explaining its dependency floors, lower bounds are deliberately conservative, roughly the newest release available at the end of 2024, capped so they never exceed the version actually exercised on the oldest supported interpreter, Python 3.9, in the lkg.txt file. The rationale is concrete, real lower bounds stop resolvers from silently back-tracking into ancient releases when a newer conflict appears, with a worked example, an unbounded scipy constraint allowed a resolver to select scipy 1.6.1, which has no wheels for modern Pythons and fails to build. Even numba is pinned explicitly as a transitive dependency to avoid incompatible pairings with numpy 2.4.x, and scikit-learn is capped below 1.10. The dependencies themselves, numpy, scipy, pandas, statsmodels, shap, sparse and joblib, are the standard scientific stack the promise list advertised. The comment also shows the cost of maintaining a package across six Python versions, where the oldest interpreter constrains every dependency decision, and where a naive requirements list can produce environments that resolve successfully but fail to build.

Cython with warnings as errors

Building from source runs a setup.py that compiles the package's performance-critical paths as native extensions. The mechanics show care, a glob finds pyx and c files recursively, and if both exist for a module the compiled c file is assumed up to date rather than recompiled, saving build time. Cython compilation sets Options.warning_errors to True, treating Cython warnings as errors, a discipline most projects reserve for linters, and the extensions include NumPy's headers with the NPY_NO_DEPRECATED_API macro set, forbidding use of the deprecated C API by construction. The primary language on GitHub is Jupyter Notebook rather than Python, reflecting that the repository's visible surface is notebooks and documentation examples, while the compiled core lives beneath the econml package directory. Zip safety is disabled in the setup call as well, a small requirement for packages that load native extensions and data files at runtime rather than through the import system alone.

A release history measured in years

The news section lists every release back to 2020, and the pattern is a slow, deliberate train, v0.6.1 in January 2020 through v0.14.0 in November 2022, then a pause to v0.14.1 in May 2023, v0.15.1 in July 2024, and v0.16.0 in July 2025, with v0.17.0 on July 31, 2026 as the current release. The repository supports the train with monte_carlo_tests and prototypes directories beside the main package, a lkg notebook recording the last-known-good environment, pre-commit configuration and a security policy. Installation is through PyPI as econml, with documentation at pywhy.org, and the last push landed 2026-09-25, so the slow release cadence is a curation choice, not inactivity.

Editorial conclusion

Use EconML when the question is what causal effect a treatment has on an outcome, and how that effect varies across a population, from observational data, provided you can defend the identifying assumptions, no unobserved confounders for most methods or a valid instrument Z for the others. It is not a substitute for experimental design where experiments are feasible. Before adopting, read the documentation's assumptions per estimator rather than treating the API as plug and play, note the dependency floor story in the package metadata, and if you build from source, expect Cython compilation with warnings promoted to errors.

Frequently asked questions

what is econml?

EconML is a Python package for estimating heterogeneous treatment effects from observational data via machine learning, built as part of the ALICE project at Microsoft Research and now maintained under PyWhy. It implements orthogonal machine learning algorithms such as double machine learning, measures the causal effect of treatments T on outcomes Y controlling for features X and W, and provides confidence intervals.

What is heterogeneous treatment effect estimation and how is it performed?

It is estimating the causal effect of an intervention on an outcome for a sample with a particular set of features, rather than one average effect. EconML performs it by modeling the effect as a function of the features X with flexible ML models like forests, boosting, lasso and neural nets, preserving causal interpretation through econometric structure, assuming no unobserved confounders or access to an instrument, and reporting confidence intervals.

How does double machine learning contribute to causal inference?

Double machine learning, the Chernozhukov et al. work EconML implements, contributes orthogonalization, using machine learning to partial out the control variables from both treatment and outcome so the residual variation identifies the causal effect with the bias that flexible models would otherwise inject removed from the estimate.

how to install econml?

Install the econml package from PyPI, which requires Python 3.9 through 3.14 and pulls the scientific stack including numpy, scipy, pandas, statsmodels, scikit-learn and shap, with conservative version floors documented in the package metadata. Building from source compiles Cython extensions with warnings treated as errors.

Official sources

  1. Issues
  2. Project website
  3. py-why/EconML on GitHub
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/py-why-econml.svg)](https://hysenlabs.com/projects/py-why-econml)