DoubleML: double machine learning for causal inference in Python
DoubleML - Double Machine Learning in Python
At a glance
- What is it?
- The Python implementation of the Chernozhukov et al. double machine learning framework wraps scikit-learn learners in Neyman-orthogonal score functions for PLR, PLIV, IRM and IIVM models. It is a good fit for economists and data scientists who already know their learners and need valid confidence intervals.
- Who is it for?
- Adopt DoubleML when you have a treatment, an outcome, a set of confounders and a learner you already trust, and you need confidence intervals rather than point predictions. Do not adopt it for prediction tasks, for panel or time-series designs, or as a substitute for a causal identification argument.
- Can I use it commercially?
- Yes. BSD-3-Clause is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 7 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The problem DoubleML addresses: regularisation bias in causal estimates
If you plug a machine learner into a treatment effect regression, the same flexibility that helps you fit the outcome also biases the coefficient you care about. Regularisation shrinks the estimate; overfitting pulls it the other way. The double machine learning framework of Chernozhukov et al. (2018) is the standard answer: split the sample, fit nuisance functions on one part, evaluate the score on the other, and use a Neyman-orthogonal moment condition so that small errors in the nuisance estimates do not translate into first-order bias in the target parameter.
DoubleML is the Python implementation of that framework. The README states it is built on top of scikit-learn and that it covers four model families: partially linear regression (PLR), partially linear IV regression (PLIV), interactive regression (IRM) and interactive IV regression (IIVM). The audience is narrow and that is a feature. If you are running an A/B test with a clean randomised assignment, you do not need this. If you have observational data, a binary or continuous treatment, high-dimensional confounders, and a reviewer who will ask for standard errors, this is the tool the econometrics literature points at.
How the object model splits nuisance estimation from inference
The architecture is documented as a single abstract base class, DoubleML, with the four model classes inheriting from it. The concrete classes implement two things: estimation of the nuisance functions through machine learning methods, and computation of the Neyman orthogonal score function. Everything else lives in the base class, which supplies the methods fit, bootstrap, confint, p_adjust and tune.
That split matters in practice. Swapping a learner, a resampling scheme or a score function is a constructor argument, not a rewrite. The README lists the extension points explicitly: new model classes whose score functions are linear in the target parameter, alternative score functions supplied as callables, and alternative resampling schemes. The repository also ships an OOP diagram at doc/oop.svg for anyone who wants to see the inheritance tree before reading source.
The dependency list in pyproject.toml shows where the weight sits: numpy, pandas, scipy, scikit-learn, statsmodels, joblib, plus optuna for tuning and matplotlib, seaborn and plotly for the diagnostic plots. There is an optional rdd extra that pulls in rdrobust, and a dev dependency group with pytest, xgboost and lightgbm. The package requires Python 3.10 or newer.
Installing DoubleML and fitting a partially linear regression
The README gives pip as the primary install path. The -U flag upgrades an existing install, which matters because the package ships frequently enough that stale versions are a real possibility.
pip install -U DoubleMLIf you want the source tree, the README documents an editable install. This is the path to take if you intend to add a score function or a resampling scheme, since the extension points are Python classes you subclass.
git clone [email protected]:DoubleML/doubleml-for-py.git
cd doubleml-for-py
pip install --editable .For development the README recommends uv, which resolves the dev dependency group in one step. Note that the extra flag in the documented command is rdd, not dev.
git clone [email protected]:DoubleML/doubleml-for-py.git
cd doubleml-for-py
uv sync --extra rddA first real use follows the standard workflow: construct a DoubleMLPLR with a data backend, a learner for the nuisance functions, and a score, then call fit. The README names fit, bootstrap, confint, p_adjust and tune as the inference surface, so a minimal script fits the model and then reads confint() for the interval. The documentation at docs.doubleml.org carries the worked examples with the exact argument names; the README does not reproduce them, so treat the docs as the reference for constructor signatures rather than guessing from the class list.
Where DoubleML is the wrong tool
The four model families are the boundary. If your design is a difference-in-differences with staggered adoption, a regression discontinuity (outside the rdd extra's scope), a synthetic control, or a mediation analysis, none of the four score functions matches it, and the README does not claim otherwise. You can add a model class with a score function linear in the target parameter, but that is a research contribution, not a configuration change.
The second limitation is statistical, not architectural. Neyman orthogonality controls first-order bias when the nuisance estimates converge fast enough. With a small sample and a flexible learner, that condition is an assumption you are making, not a property the package enforces. DoubleML will return a confidence interval regardless. The package does not diagnose whether your learner was good enough for the interval to be valid, and the README does not present a built-in check for that.
The third is scope creep. DoubleML estimates a parameter under a stated identification assumption. It does not test whether the assumption holds. If you cannot defend conditional ignorability or the instrument exclusion restriction in your setting, the software will happily produce a precise number for a quantity you have not identified.
DoubleML versus EconML: two answers to the same question
The closest alternative in Python is EconML from Microsoft Research. Both target causal inference with machine learning, and both build on scikit-learn estimators, but they answer different questions.
EconML centres on heterogeneous treatment effects: CATE estimation through meta-learners and causal forests, with the emphasis on how the effect varies across covariates. DoubleML centres on a low-dimensional target parameter, typically an average effect or a structural coefficient, with the emphasis on valid inference for that scalar. The README's model list is a structural one: PLR, PLIV, IRM, IIVM, each with a named score function. If your deliverable is a single number with a confidence interval and a defensible asymptotic argument, DoubleML is the more direct route. If your deliverable is a ranking of subgroups by predicted effect, EconML's framing is closer.
There is also an R twin. The README notes the Python package was developed together with an R package based on mlr3, available on GitHub and CRAN. Teams running both languages can keep the model specification aligned, which is unusual and worth knowing if your group splits across the two.
Maintenance cadence, licence and upgrade cost
The repository is not archived, and the last push was on 2026-08-18. Releases in the 0.11 line landed on 2026-01-19, 2026-05-22 and 2026-08-10, so a patch cadence of roughly one release every three to four months is visible in the release history. The README names two maintainers, @PhilippBach and @SvenKlaassen, and points bugs at the GitHub issue tracker. Funding from the Deutsche Forschungsgemeinschaft is acknowledged for Project Number 431701914 and Grant GRK 2805/1.
The licence is BSD-3-Clause, declared in pyproject.toml as a LICENSE file reference and in the classifiers as an OSI-approved BSD licence. That is permissive: you can use the package in commercial work, and redistribution requires keeping the copyright notice and licence text. The README asks for a citation if you use the package and links a CITATION.cff, which is a request rather than a licence term. For formal advice on how BSD-3-Clause interacts with your own distribution model, talk to counsel; nothing here is legal advice.
The upgrade cost is the one to watch. The dependency floor on scikit-learn is 1.6.0 and on numpy is 2.0.0, which means an environment pinned to an older scientific stack will not resolve. Because the public surface is a small set of methods on a base class, most minor upgrades should be mechanical, but the nuisance-learner interface is the part most likely to move when scikit-learn changes its estimator conventions. Pin the version in production and re-run your specification when you bump it.
Editorial conclusion
Adopt DoubleML when you have a treatment, an outcome, a set of confounders and a learner you already trust, and you need confidence intervals rather than point predictions. Do not adopt it for prediction tasks, for panel or time-series designs, or as a substitute for a causal identification argument. Before committing, check that the score function for your design is one of the four implemented families, verify that the nuisance learner can meet the convergence assumptions on your sample size, and confirm that the 0.11.x release you install is the one you tested against, since the package ships on a regular release cadence.
Frequently asked questions
What does double machine learning do?
It estimates a causal or structural parameter when nuisance functions are fitted with machine learning, using sample splitting and a Neyman-orthogonal score so that small nuisance errors do not create first-order bias in the target parameter. DoubleML implements this framework for PLR, PLIV, IRM and IIVM models.
How do I install DoubleML?
The README gives pip install -U DoubleML as the standard route. A source install uses git clone followed by pip install --editable ., and the README recommends uv with the rdd extra for development environments.
Which model classes does DoubleML provide?
Four: DoubleMLPLR for partially linear regression, DoubleMLPLIV for partially linear IV regression, DoubleMLIRM for interactive regression models and DoubleIIVM for interactive IV regression models. All other functionality, including fit, bootstrap, confint, p_adjust and tune, lives in the abstract base class DoubleML.
Which Python versions and dependencies does DoubleML require?
Python 3.10 or newer, according to the classifiers and the requires-python field in pyproject.toml. The runtime dependencies include numpy 2.0.0 or newer, scikit-learn 1.6.0 or newer, pandas, scipy, statsmodels, joblib, optuna, matplotlib, seaborn and plotly.
What licence does DoubleML use?
BSD-3-Clause, declared in pyproject.toml and listed in the classifiers as an OSI-approved BSD licence. The README also asks for a citation when the package is used in published work.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/doubleml-doubleml-for-py)