# pgmpy keeps twelve unbounded dependencies and a dev branch that ships the live examples

> An MIT-licensed Python toolkit for causal and probabilistic reasoning built on graphical models, released as 1.1.2 while the default dev branch moves ahead. Every runtime dependency has a lower bound and no ceiling, Python is capped below 3.15, the backend is a global switch set before imports, and all three quickstart samples end mid-statement.

**pgmpy/pgmpy** — Python Toolkit for Causal and Probabilistic Reasoning

- Repository: https://github.com/pgmpy/pgmpy
- Website: https://pgmpy.org/
- Stars: 3,350 · Forks: 1,181
- Language: Python
- License: MIT
- Published: 2026-09-24 · Updated: 2026-09-24 · Language: en
- Canonical page: https://hysenlabs.com/projects/pgmpy-pgmpy

## The default branch is dev and the license badge points at main

The repository's default branch is dev, and the latest tagged release is 1.1.2 from April 2026, which is also the version in the packaging metadata. The two agree, but the branch they sit on does not: dev was pushed at the start of October, so the code on the default branch is ahead of the newest release by several months of work. Two links in the header disagree about which branch is real. The open source row links the licence file through main, while the contributing guide, the examples tree and the Binder link all use dev. One of those is going to 404 for someone, and the licence row is the one a reader is most likely to click. Worth knowing that dev is where development happens rather than treating main as the trunk, since that reverses the usual assumption.

## Twelve runtime dependencies and not one upper bound

The runtime list is long for a probability package and every entry is an open-ended lower bound. huggingface_hub, networkx, numpy, scipy, scikit-learn, pandas, statsmodels, tqdm, pyparsing, joblib, opt_einsum and scikit-base. Nothing carries a ceiling, so a project that already pins NumPy and SciPy for its own reasons can end up in a resolution fight with this one, or silently get versions nobody tested. Two of those entries explain a lot about the design. huggingface_hub in the base dependency set is how datasets and models are fetched rather than shipped. scikit-base is the framework behind the unified composable API the README describes, which is what lets the estimators and algorithms share one interface and satisfy the scikit-learn conventions where possible.

## Python is capped below 3.15

The interpreter requirement is greater than or equal to 3.10 and less than 3.15, with classifiers for 3.10 through 3.14. Unlike the dependencies, this upper bound is deliberate and enforced by the resolver, so on a fresh 3.15 interpreter pip will refuse the package rather than install it and hope. For a scientific library that is the conservative choice, since the numerical stack underneath is the thing most likely to need work on a new interpreter, and the constraint is documented in the metadata rather than only in the classifiers. If you are on the newest Python and the package will not install, this bound is the reason, and there is no environment marker to work around it.

## Every quickstart sample stops mid-statement

There are three examples in the quickstart, one for discrete data, one for linear Gaussian data and one for mixtures with arbitrary relationships, and all three end in the middle of a line. The first finishes on a comment that begins Drop a colu. The second finishes on a comment about dropping a column and predicting with the model. The third ends on the variable name cpd_ before the functional conditional probability distributions are defined. The working part of each sample is complete, so you can follow the shape of the workflow from them, but none of them runs to the end as printed. The discrete one shows the sequence a first user needs: load a model, simulate data, learn a structure with the PC estimator, fit parameters, then read the conditional probability distributions back.

```python
from pgmpy.example_models import load_model

# Load a Discrete Bayesian Network and simulate data.
discrete_bn = load_model("bnlearn/alarm")
alarm_df = discrete_bn.simulate(n_samples=100)

# Learn a network from simulated data.
from pgmpy.estimators import PC

dag = PC(data=alarm_df).estimate(ci_test="chi_square", return_type="dag")
```

The linear variant is the same shape with a pearsonr conditional independence test and a linear Gaussian network class instead of the discrete one.

## The backend is a global switch set before the imports

The mixture example reaches for a config object from the package's global variables module and calls set_backend with torch on it, before importing a probabilistic programming library and then defining a functional network with functional conditional probability distributions. That ordering is the point. The backend is process-wide mutable state rather than a parameter passed to a constructor, so it has to be set before the pieces that capture it are imported, which means import order becomes part of your program's correctness. It also means two networks in one process cannot use different backends without resetting global state in between. The extras separate the requirement out: the torch extra pulls torch and a pyro distribution library, and the example only works with that extra installed.

## The Binder examples run the dev branch

The interactive tutorial link points at a Binder environment built from the dev branch of this repository with the examples directory as the starting path. So the notebooks a stranger is invited to click are running unreleased code on the development branch, not the version you would install from PyPI or conda-forge. That is a deliberate trade, and a common one, but it means a notebook can demonstrate an API that has changed since the last release, and it means the examples describe the code as it is being written. The notebooks themselves are the substantial part of the repository's examples directory, eighteen of them, covering discrete and linear network creation, defining conditional probability distributions, dynamic networks, expert knowledge, inference by junction tree, parameter learning over factor graphs, structure learning including Chow-Liu and TAN variants, simulating data, and one on extending pgmpy itself.

## Two tutorial repositories and a downloads row with no label

The header table and the resources list point at two different repositories for tutorial material, one named for tutorials and one named for notebooks, and nothing in the documentation says which is current. The same table has a row for download counts whose link text is empty, so the badge renders as a link with nothing in it. Around the edges, the project is unusually well organised for a research library: a governance file, a funding declaration file, a code of conduct, a citation file, a changelog, a pre-commit configuration and a committed test duration file that exists so the test splitter can balance its runs across machines. Funding is listed through several programmes and a scientific Python foundation, and the contribution guide points at a wiki page of mentored projects for newcomers rather than only at open issues.

## Conclusion

pgmpy is worth reaching for when your problem is a graph over variables rather than a table of rows, because it ships the model classes and the estimation, inference and simulation algorithms behind one composable API, and the estimators conform to the scikit-learn interface where that makes sense, so the pieces drop into a pipeline instead of needing a wrapper. Three things to decide first. Every runtime dependency is a lower bound with no ceiling, so the resolver, not the project, decides which NumPy, SciPy or pandas your environment gets. Python is capped below 3.15, which will eventually matter on a new interpreter. And the torch backend is selected through a global config object that has to be set before the imports, which is a sharp edge to know about before you wonder why a mixture model is not using your GPU.

## FAQ

### what is pgmpy

A Python toolkit providing the building blocks for causal and probabilistic reasoning with graphical models. It ships data structures for directed and other graph types, Bayesian networks, dynamic Bayesian networks and structural equation models, plus algorithms for causal discovery, causal identification, inference, validation, parameter estimation and simulation.

### what is pgmpy in python

The same toolkit, installed with pip or from conda-forge. Its algorithms follow a unified composable API and are scikit-learn compatible where possible, so they can be called directly, dropped into a scikit-learn pipeline, or used to build higher-level tools.

### How to install pgmpy?

From PyPI with pip install pgmpy, or from conda-forge with conda install conda-forge::pgmpy. Python 3.10 to 3.14 is supported, and the package is on both channels.

### how to install pgmpy package in python

The same two commands work depending on your environment. Working with the mixture or functional examples additionally needs the torch extra, which installs the torch backend.

## Sources

- [License: MIT](https://github.com/pgmpy/pgmpy/blob/dev/LICENSE)
- [pgmpy/pgmpy on GitHub](https://github.com/pgmpy/pgmpy)
- [Project website](https://pgmpy.org/)
- [README](https://github.com/pgmpy/pgmpy/blob/dev/README.md)
- [Releases](https://github.com/pgmpy/pgmpy/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/pgmpy-pgmpy
