# bnlearn: A Pipeline Wrapper for Bayesian Network Structure Learning in Python

> bnlearn packages structure learning, parameter learning, inference and sampling behind four fit() calls. It is a convenience layer over existing causal discovery code, and the ease of that layer is also where its limits start.

**erdogant/bnlearn** — Python package for Causal Discovery by learning the graphical structure of Bayesian networks. Structure Learning, Parameter Learning, Inferences, Sampling methods.

- Repository: https://github.com/erdogant/bnlearn
- Website: https://erdogant.github.io/bnlearn
- Stars: 645 · Forks: 61
- Language: Jupyter Notebook
- License: NOASSERTION
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/erdogant-bnlearn

## What bnlearn Is Trying to Fix

Probabilistic graphical models have a steep entry cost. Fitting a Bayesian network means choosing a structure learning algorithm, encoding the data correctly for that algorithm, estimating conditional probability tables, and then keeping the learned graph and the parameters in sync when you want to run inference. The README states the motivation plainly: because probabilistic graphical models can be difficult to use, bnlearn contains the most-wanted pipelines. The target reader is an analyst or applied machine learning engineer who wants a causal graph from a pandas-style table without assembling four separate libraries. The README leans on tutorial material, with Medium posts and podcasts linked beside most features, which suggests the intended path is guided rather than reference-driven. If you already know which scoring function and search strategy you want, the package is not aimed at you.

## The Four fit() Calls and What Sits Behind Them

The public surface is deliberately flat. Structure learning is bn.structure_learning.fit(), parameter learning is bn.parameter_learning.fit(), inference is bn.inference.fit(), and prediction is bn.predict(). Synthetic data comes from bn.sampling(). Around those, the README lists supporting functions: bn.independence_test() for edge strength, bn.knn_imputer() for imputation, bn.discretize() for binning continuous columns, bn.make_DAG() for building a graph by hand, and bn.check_model() for validating model parameters. The data flow implied by the API is linear: a DataFrame goes into structure learning, the returned graph object goes into parameter learning, and the fitted model goes into inference or sampling. The presence of bn.make_DAG() matters for a different workflow, where an expert supplies the graph and only the parameters are learned from data. That split, learn the graph from data or accept it from a human, is the main architectural decision the package exposes. The README does not describe the internal scoring functions or search operators, so the choice of algorithm is hidden behind keyword arguments you have to look up in the API documentation rather than in the README.

## Getting It Running: Install and First Calls

Installation follows the standard PyPI route: pip install bnlearn. The README does not show a conda command, so pip is the documented path. Once installed, the functions become available under the bn namespace, and the README lists them as bn.structure_learning.fit(), bn.parameter_learning.fit(), bn.inference.fit(), bn.predict(), bn.sampling(), bn.independence_test(), bn.knn_imputer(), bn.discretize(), bn.check_model() and bn.make_DAG(). Those names are the contract. What the README does not give is a worked call with arguments, so the concrete keyword names for the structure learning method, the scoring function, or the discretization bins are not visible in the repository front page. They live in the Sphinx site at erdogant.github.io/bnlearn, which the README links repeatedly and which is where you should read before writing code. The README does link Colab notebooks, and those are the closest thing to runnable examples in the material available. Treat the badges and the feature table as an index, not as documentation.

## Discretization Is Not Optional Housekeeping

The README lists discrete, continuous and mixed datasets as supported, and separately lists a discretize function with its own documentation page. That combination is worth reading carefully. Structure learning over discrete data and structure learning over continuous data are different problems with different scoring functions, and mixing the two in one table forces a decision about which columns get binned and how many bins they get. The package gives you bn.discretize() to make that decision explicit, and the README groups it with a comparison page, which hints that the library authors see discretization as a place where tools diverge. The practical consequence: the number of bins you choose changes the conditional probability tables that parameter learning later estimates, and changes which edges the structure search finds plausible. Two runs of the same pipeline with different bin counts can return different graphs. That is not a defect in bnlearn, it is a property of the underlying method, but a wrapper that makes the pipeline look like four uniform fit() calls can obscure it.

## Where the Wrapper Design Costs You

The value of bnlearn is that it hides orchestration. The cost is the same thing. When a learned graph comes back with an edge you did not expect, you need to know which scoring function and which search procedure produced it, and the README does not surface that. You will be reading the Sphinx API pages. A second limitation is the licence signal. The repository metadata reports NOASSERTION, while the README carries an MIT badge pointing at the LICENSE file on master. Those two statements do not agree, and for a package you intend to ship inside a product, that discrepancy is the first thing to resolve by opening the LICENSE file itself. A third constraint is scale. The README advertises exhaustive search alongside the other structure learning methods, and exhaustive search over parent sets grows with the number of variables in a way that makes it unsuitable past a modest column count. The README does not state a variable limit, so the honest position is that no threshold is documented and you should measure on your own data before assuming the default method is appropriate. Finally, the repository is primarily Jupyter Notebook, which means the notebook examples are likely to be ahead of or behind the packaged Python in ways that are hard to detect from the file listing alone.

## pgmpy and the Difference in Approach

The README's own causal inference row links to pgmpy.org, which makes pgmpy the obvious comparison. The difference is one of scope and posture. pgmpy is a modelling library: you construct a BayesianNetwork object, add nodes and edges, define conditional probability distributions, and then run inference. Control sits with you at every step, and the library assumes you know what graph you want or how to specify the search. bnlearn inverts that. It starts from a DataFrame and offers fit() calls that pick the mechanics for you, with bn.make_DAG() available as the escape hatch when you want to supply the structure yourself. If your work is exploratory, where you want a candidate graph quickly and are willing to iterate on discretization and method choices, bnlearn's shape is a better fit. If your work is confirmatory, where the graph is a hypothesis you must specify and defend line by line, pgmpy's explicit construction matches that intent more closely. Neither is a superset of the other, and the README itself treats pgmpy as a reference point rather than a rival.

## Maintenance, Releases and Upgrade Surface

The release cadence visible in the material is active: 0.14.0 in July 2026, 0.14.1 in late August 2026, and 0.14.2 on the same day as the last push, with the repository not archived. Patch releases at that interval usually mean bug fixes and small API adjustments rather than redesigns, but the material does not include a changelog, so the contents of each release cannot be confirmed from what is available here. The upgrade cost you should plan for is not the install, it is the reproducibility of your graphs. A structure learning result depends on the data, the discretization and the algorithm defaults. If a patch release changes a default, your previously learned graph may not reproduce byte for byte. Pinning the version in your requirements file and re-running bn.check_model() after an upgrade is the cheap safeguard. On licensing, the README badge says MIT and the repository metadata says NOASSERTION. I am not giving legal advice; the point is that you should read the LICENSE file on the master branch and confirm which of the two statements is authoritative before you depend on the package commercially.

## Conclusion

Adopt bnlearn if you want a single Python entry point for structure learning, parameter learning, inference and synthetic sampling, and you accept that the package is a convenience layer rather than a from-scratch implementation. Do not adopt it if you need a formally specified licence, since the repository metadata reports NOASSERTION while the README badge points to MIT, or if your data is high-dimensional and you have not decided on a discretization scheme first. Before committing, verify three things: which structure learning method each call defaults to, how the chosen discretization changes the learned graph, and the exact licence text in the LICENSE file on the master branch.

## FAQ

### What is Bayesian structure learning?

It is learning the graph structure of a Bayesian network from data or from expert knowledge, as opposed to fitting the numbers once the shape is known. In bnlearn the structure step is bn.structure_learning.fit(), and it is a separate call from bn.parameter_learning.fit(), which estimates the conditional probability distributions from observed data.

### Is a Bayesian network a dag?

It is represented as one. bnlearn exposes bn.make_DAG() to create a directed acyclic graph directly, and structure learning runs over that graph. The library also separates structural and parameter concerns in code, with bn.check_model() for validating a model and bn.get_parents() for reading parents back off the edges.

### How do I install bnlearn?

The package needs Python 3.10 or newer. Installing it pulls pgmpy between 1.1.2 and 1.2, networkx, matplotlib, numpy, pandas, scikit-learn, python-louvain and scipy at 1.14.0 or newer. The licence field in the package metadata reads MIT, and the default branch in the repository is master rather than main.

### What does the bnlearn command line do?

It is not a CLI for running models. The single console script maps bnlearn to bnlearn.skills_install:main, so invoking it runs a skills installer. The pipelines are called from Python instead, as bn.structure_learning.fit(), bn.parameter_learning.fit() and bn.inference.fit().

### Which causal methods does bnlearn cover?

Structure learning, parameter learning, inference, prediction, sampling and edge strength. The pipelines are bn.structure_learning.fit(), bn.parameter_learning.fit(), bn.inference.fit(), bn.predict(), bn.sampling() and bn.independence_test(). Causal inference is described as computing interventional and counterfactual distributions using do-calculus, and synthetic data generation is one of the listed features.

## Sources

- [erdogant/bnlearn on GitHub](https://github.com/erdogant/bnlearn)
- [Issues](https://github.com/erdogant/bnlearn/issues)
- [Project website](https://erdogant.github.io/bnlearn)
- [README](https://github.com/erdogant/bnlearn/blob/master/README.md)
- [Releases](https://github.com/erdogant/bnlearn/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/erdogant-bnlearn
