Library / SDK
process-intelligence-solutions/pm4py avatar
process-intelligence-solutions/pm4py

The README says Python 3.9, the manifest requires 3.11

Official public repository for PM4Py (Process Mining for Python) — an open-source library for exploring, analyzing, and optimizing business processes with Python.

1,045 stars357 forksPythonAGPL-3.0

At a glance

What is it?
A process mining library that came out of a research institute and is now maintained by a company, licensed under the AGPL with a commercial alternative. The interesting parts are the packaging details: three different statements about the supported Python floor, a dependency that is both required and optional, and a Docker image with a commented-out source build of the whole scientific stack.
Who is it for?
PM4Py is the right library to reach for if you are doing process mining in Python, and the combination of a permissive dependency set, an inductive miner that works in three lines, and extras that extend into machine learning, language model calls, object-centric logs and Polars is more than most alternatives in this space offer at one version. Two things to check before you pin it.
Can I use it commercially?
Yes, with strict conditions. AGPL-3.0 is a network copyleft licence: if people use a modified version over a network, for example as a hosted service, you must offer them its source code under the same licence.
Is it still maintained?
Yes. The repository last received commits 31 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 2, 2026, and from our analysis. They are not legal advice.

Editorial analysis

Three statements about which Python versions work

The installation section says the library can be installed on Python 3.9.x, 3.10.x, 3.11.x, 3.12.x, 3.13.x or 3.14.x, and then the command is a plain upgrade of the package from the index.

The project manifest says something different. The requires-python field is 3.11 or newer, and the classifiers list only 3.11, 3.12, 3.13 and 3.14. So the two statements overlap on four versions and disagree about two.

A third statement sits underneath both. The README adds that the library also runs on older Python environments with different requirement sets, and names one: Python 3.8, pinned at 3.8.10, with a requirements file kept in a third_party directory for older dependency sets.

So a reader on Python 3.9 or 3.10 has a README that says they are supported and a manifest that says they are not, and a reader on 3.8 has a requirements file that the manifest does not know about. The manifest is the one pip enforces, so the first two of those cases will fail at install time.

One dependency is both required and optional

The core dependency list contains an entry that is only installed on older Python:

code
cvxopt; python_version < '3.15'

A conditional requirement like that is an environment marker, so on any Python below 3.15 the solver package comes in with the base install. On 3.15 and above it does not.

The same package also appears in the optional dependency set, in a group called solvers, which installs cvxopt. That is not a mistake, since an extra is a way to add something you already have, but it means the phrase in the requirements documentation, that the default installation contains what is needed for mainstream use and additional integrations are grouped by feature, has one member that is in both.

The rest of the groups are cleaner and more informative. There is one for machine learning, one for large language model calls, one for object-centric event logs, one for a Polars dataframe backend, one for calendars, one for HTTP connectors, one for visualization, one for Windows interaction libraries, and one called stable.

That last one deserves attention. It pins exact versions of everything in the resolved set, and the README describes it as the reproducible installation using the exact dependency versions validated by the project, with the same versions used by continuous integration.

Three releases in seventy seconds, then five months of commits

The three most recent releases share a timestamp cluster. 2.7.20, 2.7.21 and 2.7.22 are all dated 2026-03-20, seventy-two seconds apart from first to last.

That looks like a fix being published, corrected and republished. Then the line stops: the last push to the repository is dated 2026-09-01, five and a half months later.

So the repository has had commits for months without a new tag, and the newest version anyone can install from the index is from March. The manifest does not carry a version number at all, it declares the version as dynamic, which is what lets the same file produce a different version per build.

The release notes live in a changelog file at the repository root, and that is the only place to find out what changed in those five months. There is also a coverage report committed at the root next to it, which means the test coverage figure for the current state of the tree is a file in version control rather than something a reader has to build to see.

The first example ships a placeholder file path

The example that is supposed to spark interest is four lines of real code and one placeholder:

python
import pm4py

if __name__ == "__main__":
    log = pm4py.read_xes('<path-to-xes-log-file.xes>')
    net, initial_marking, final_marking = pm4py.discover_petri_net_inductive(log)
    pm4py.view_petri_net(net, initial_marking, final_marking, format="svg")

The angle-bracketed file name has to be replaced with a real event log before anything runs, which is a reasonable thing for a four line teaser and a small obstacle for someone copying it into a terminal.

The three calls are the whole library in miniature. Reading an XES event log, discovering a Petri net with the inductive variant of the miner, and rendering the result as SVG. That last call is why graphviz appears in the core dependency list rather than in a visualization group, since the diagram is produced through it.

The default branch is named release rather than main, which is worth knowing before you clone and read something you expected to be the development line.

AGPL-3.0 in the text, or-later in the manifest

The licensing section states that the open source version available on GitHub is licensed under version 3 of the GNU Affero General Public License. The manifest's SPDX expression is AGPL-3.0-or-later.

Those are not the same grant. The README describes version 3, and the or-later suffix permits relicensing under later versions of the licence. It is a small difference, and it is exactly the kind of detail that matters when a legal team reads a manifest instead of a README.

The rest of the licensing story is commercial. The maintainers offer a separate version of the library for commercial use in closed-source environments under a different licence, and point at their own site for the options. That is the arrangement you would expect from a library that started inside a research institute, and the README says exactly that: development started at the Fraunhofer Institute for Applied Information Technology, and the project is now managed and developed by Process Intelligence Solutions, a spin-off.

The third party licensing picture is handled in the repository rather than in prose. A third_party directory lists the licences of the direct dependencies, and a separate file there holds the full list of transitive dependencies with their licences. For a library with a dependency set this wide, having that inventory in the tree is worth more than a paragraph.

The Docker image carries a commented-out build of the entire stack

The container file is not minimal. It starts from a pinned Python image, 3.13.0 on bookworm, then installs a long list of system packages: graphviz and tini and a ping utility, an ODBC development package, a full compiler toolchain with flex, bison, cmake and the automake set, and a large block of scientific libraries including OpenBLAS, LAPACK, the Boost family, jemalloc, Thrift and FFTW, plus clang and llvm at version 16.

Then there is a run of lines that are all commented out, and they are the interesting part. They clone numpy, pandas, scipy, lxml, matplotlib, duckdb and the Apache Arrow C++ library, update submodules, and build each from source, with a cmake invocation for Arrow that sets the compute, CSV, dataset and filesystem options and installs into a prefix that the layer then puts on the library path.

That is a record of a different approach to the same problem, left in the file. Nothing uses those layers as written, and a reader looking for how the image is actually built has to skip seven lines of commented experimentation to get there.

The Python patch level is also pinned to 3.13.0 while the manifest supports 3.11 through 3.14, so the image is one patch release behind and two minor versions narrower than the supported range.

Ten extras, a flat examples directory, and an unexplained folder

The requirements section enumerates ten optional feature groups, and two of them are meta rather than technical: all, which installs every optional feature supported on the current platform, and stable, the pinned set the project validates in CI. A third, windows, exists for interaction libraries that only apply on that platform.

Development installation is two commands, an editable install of the checkout, and a dependency group install for tooling, with the lint group given as the example. The manifest's build requirements are setuptools and wheel, which is a plain choice for a project that ships wheels rather than an extension.

The repository layout is wide: a package directory, tests, documentation, notebooks, a files directory for data, a third_party directory for licences, and an examples directory that is flat. The visible portion of that examples directory is a long alphabetical list of standalone scripts, from a check for missing items through activity conversion, several approximation and alignment variants, batch detection, BPMN conversion, case overlap statistics, concept drift, and correlation mining. There is no per-topic subdirectory, so the example for the thing you need is a search away.

One entry has no explanation anywhere in the visible documentation: a directory called safety_checks at the repository root. Its name suggests dependency scanning, and its presence next to the third_party licence inventory is consistent with that, but nothing in the README says what it is.

Editorial conclusion

PM4Py is the right library to reach for if you are doing process mining in Python, and the combination of a permissive dependency set, an inductive miner that works in three lines, and extras that extend into machine learning, language model calls, object-centric logs and Polars is more than most alternatives in this space offer at one version. Two things to check before you pin it. Which Python you are on, because the README advertises 3.9 and 3.10 while the manifest requires 3.11, and the only reliable answer is the manifest. And which version you install, because the newest release tag is from March even though commits have continued since, so read the changelog at the root rather than assuming the index is current. If the licence matters to you, read the manifest rather than the README paragraph, because the SPDX expression grants or-later while the prose says version 3, and the commercial closed-source edition is a separate conversation with the vendor.

Frequently asked questions

How do I install PM4Py?

pip install -U pm4py. Optional features come from extras, for example pip install -U "pm4py[polars,ml]", with pm4py[all] installing everything the current platform supports and pm4py[stable] pinning the exact dependency versions the project validates in CI. The README lists 3.9 through 3.14 while the project manifest requires Python 3.11 or newer.

What is PM4Py?

A Python library for process mining, developed at the Fraunhofer Institute for Applied Information Technology and now managed and developed by Process Intelligence Solutions, a spin-off from that institute. The first example in the documentation reads an XES event log, discovers a Petri net with the inductive miner and renders it as SVG.

Can I use PM4Py in a closed-source commercial product?

The version published on GitHub is licensed under the GNU Affero General Public License, expressed in the project manifest as AGPL-3.0-or-later, while the README states version 3 of the AGPL. The maintainers separately offer a version for commercial use in closed-source environments under a different licence, with the options described on their own site.

What optional features do the PM4Py extras add?

Ten groups, including machine learning with an earth mover distance package and scikit-learn, language model calls with openai and requests, object-centric event logs with jsonschema and pyarrow, a Polars dataframe backend, calendars, HTTP connectors, visualization, a Windows-only group for window and input libraries, plus an all group and a stable group that pins the versions used by continuous integration.

Which Python versions does PM4Py officially support?

The project manifest requires Python 3.11 or newer and its classifiers cover 3.11, 3.12, 3.13 and 3.14, while the README text says 3.9 through 3.14. The README also points at a separate requirements file for Python 3.8 at version 3.8.10 for older environments with different dependency sets.

How is PM4Py cited in academic work?

The documentation asks for a specific journal article: Alessandro Berti, Sebastiaan van Zelst and Daniel Schuster, PM4Py: A process mining library for Python, published in Software Impacts, volume 17, article 100556, with a DOI and a BiBTeX entry in the README. The article is dated 2023 while the current version is 2.7.x.

Official sources

  1. License: AGPL-3.0
  2. process-intelligence-solutions/pm4py on GitHub
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/process-intelligence-solutions-pm4py.svg)](https://hysenlabs.com/projects/process-intelligence-solutions-pm4py)