DoWhy: causal effect estimation in Python with graph-based assumptions and refutation tests
DoWhy is a Python library for causal inference that supports explicit modeling and testing of causal assumptions. DoWhy is based on a unified language for causal inference, combining causal graphical models and potential outcomes frameworks.
At a glance
- What is it?
- DoWhy is an MIT-licensed Python library that combines causal graphs and potential outcomes behind one API. It is strongest when you can write down your assumptions and want a built-in way to attack them.
- Who is it for?
- Adopt DoWhy if you can state a causal graph and want the identification step and refutation tests to be visible rather than buried in an estimator. Skip it if your problem is pure prediction, or if you cannot defend any graph at all: the four-step API will not invent the assumptions for you.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository received new commits within the last day.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The problem DoWhy addresses: assumptions you can write down and test
Most machine learning libraries answer associational questions well. Feed them features, get a prediction. The moment the question becomes "what happens to the outcome if I change this variable," that machinery gives no guarantee, because the fitted model has no representation of which variables are confounders and which are mediators. DoWhy is built for that second kind of question. The README frames it as decision-making: determining how a potential action affects an outcome, understanding what led to a current value, or simulating what happens if some variables are changed. The library is aimed at data scientists and researchers who already have domain knowledge about how variables influence each other and want that knowledge expressed in code. The README is explicit that a key feature is the refutation and falsification API, which can test causal assumptions for any estimation method, and it says this makes inference more accessible to non-experts. That is a defensible claim about workflow, not about correctness: the tests probe an assumption, they do not prove it.
Two frameworks, one pipeline: graphs for identification, potential outcomes for estimation
The architecture is the interesting part. DoWhy does not pick a side between graphical causal models and potential outcomes; it uses each where it is stronger. For effect estimation, the README states that it uses graph-based criteria and do-calculus to model assumptions and to identify a non-parametric causal effect, then switches to methods based primarily on potential outcomes for the actual estimation. That split matters in practice. Identification is a property of the graph, not of the data, so a wrong graph cannot be rescued by a better estimator. Estimation is a property of the data and the estimator, so a correct graph can still yield a noisy number. For questions beyond effect estimation, the library models the data generation process through explicit causal mechanisms at each node, which the README says unlocks attribution of observed effects to particular variables and point-wise counterfactuals. The four task families listed are effect estimation, quantifying causal influences (mediation, direct arrow strength, intrinsic causal influence), what-if analysis, and root cause analysis. Those are different APIs sharing conventions, not one function with flags.
Installing DoWhy and running a first estimation
The README states that DoWhy supports Python 3.8+, and that you can install with pip, poetry, or conda. The pyproject file constrains the package to Python >=3.9,<3.14, so the packaging metadata is narrower than the README line. The pip route is the shortest.
pip install dowhyThe conda route is documented as well, and the README notes a common failure: if you hit "Solving environment" problems with conda, it suggests running conda update --all and then installing dowhy. Poetry users get the third option.
poetry add dowhyOnce installed, the library's own documentation at py-why.github.io/dowhy carries the user guide and sample notebooks; the README points there rather than inlining a full walkthrough. The repository also ships an example_parallel_refutation.py at the top level and a tests directory, which is where to look for runnable code that matches your installed version. The README's case-study links cover effect estimation (hotel booking cancellations, customer loyalty programs, home visits on infant health), root cause analysis (an online shop, elevated latencies in a microservice architecture, a supply chain distribution change), and those notebooks are the practical starting point for the API shape.
Where DoWhy stops helping: unverifiable graphs and thin time-series support
The honest limitation is upstream of the code. DoWhy requires you to supply a causal graph or a set of assumptions, and the refutation API tests those assumptions rather than discovering them. If you cannot defend a graph, the library will still run and still print an estimate, and that estimate inherits every error in your structure. The refutation tests are a signal, not a verdict; the README describes them as testing assumptions, and passing a placebo or random-common-cause check does not establish that the graph is right. Time-series is the other soft spot. The v0.12 release notes describe support for time-series data as experimental, and no later release note in the list removes that label. If your data has temporal ordering, interference between units, or treatment that varies over time, this is not the tool to reach for first, and the documentation does not present a mature path for it. The project also classifies itself as Development Status 4 - Beta in pyproject.toml, which is worth weighing if you are embedding it in a production pipeline.
DoWhy compared with EconML: identification plus refutation versus estimation libraries
EconML is the natural comparison, and the difference is scope rather than quality. EconML focuses on estimating heterogeneous treatment effects with machine learning methods such as double machine learning and causal forests. DoWhy's README describes a broader pipeline: model assumptions as a graph, identify a non-parametric effect with do-calculus, estimate it, then refute the result. In other words, DoWhy wraps the estimation step in an explicit identification and falsification layer, while an estimation-focused library assumes the identification argument has already been made elsewhere. The two are not mutually exclusive, and the dependency notes in pyproject.toml mention numba as imported by econml, which indicates the projects have been used together. A reasonable division of labor: use DoWhy to write down and stress-test the causal claim, and reach for a specialized estimator when you need a particular heterogeneous-effect model that DoWhy's estimator list does not cover. The v0.14 release notes mention a new doubly robust estimator, so the estimator surface does keep growing, but the identification and refutation steps remain the reason to pick this library over a bare estimator.
Maintenance, licensing, and what an upgrade costs you
The repository is not archived, and the last push was on 2026-09-10, so it is being worked on. The release cadence visible in the notes is roughly two per year: v0.12 in November 2024, v0.13 in July 2025, v0.14 in November 2025. Each release has carried a compatibility or capability change rather than only bug fixes: Python 3.12 support and experimental time-series in v0.12, the Generalized Adjustment Criterion and missing data support in GCM in v0.13, Python 3.13 support and the doubly robust estimator in v0.14. That pattern means upgrades are not purely mechanical. If you pin a version, read the release notes for the version you are moving to before you move, because identification behavior and estimator availability have both changed across these releases. The licence is MIT, stated in both the LICENSE file and pyproject.toml, which permits commercial use and modification with the copyright notice retained; that is a factual description of the licence text, not legal advice, and your organization's own review still applies. The Python range in pyproject.toml is >=3.9,<3.14, so a Python 3.14 environment is outside the declared support envelope even though the README's 3.8+ line reads more permissively.
Editorial conclusion
Adopt DoWhy if you can state a causal graph and want the identification step and refutation tests to be visible rather than buried in an estimator. Skip it if your problem is pure prediction, or if you cannot defend any graph at all: the four-step API will not invent the assumptions for you. Before committing, read the identification output for your own dataset and check which estimators the installed version actually exposes.
Frequently asked questions
What is DoWhy?
DoWhy is a Python library for causal inference that supports explicit modeling and testing of causal assumptions. It combines graphical causal models and potential outcomes behind a unified interface for effect estimation, causal influence quantification, what-if analysis, and root cause analysis.
What are some Python libraries for causal inference?
DoWhy is one, and it is part of the PyWhy ecosystem, whose GitHub organization the README points to for related tools. The README also notes that econml imports numba, which is a sign the two libraries are used in the same environments.
How does DoWhy compare with EconML?
EconML concentrates on estimating treatment effects with machine learning methods, while DoWhy adds an explicit identification step using graph-based criteria and do-calculus, then a refutation API that tests the assumptions behind any estimation method. DoWhy's README describes this as combining graphical causal models and potential outcomes.
What is an example of a causal inference question DoWhy can answer?
The README lists effect estimation cases such as the effect of home visits on infant health and the effect of customer loyalty programs, plus root cause analysis cases like finding the root cause of elevated latencies in a microservice architecture. Each is a question about what happens when a variable is changed, not a prediction question.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/py-why-dowhy)