Model or dataset
zenml-io/zenml avatar
zenml-io/zenml

ZenML: A Python Pipeline Layer That Tracks Runs and Sits on Top of Your Existing Stack

ZenML 🙏: One AI Platform from Pipelines to Agents. https://zenml.io.

5,582 stars657 forksPythonApache-2.0

At a glance

What is it?
ZenML wraps Python functions into pipelines and records what each run produced, then runs that pipeline on whatever infrastructure you already have. It is aimed at ML and AI engineers in company settings, and its value depends on whether you want a metadata and orchestration layer rather than a training framework.
Who is it for?
Adopt ZenML if you already have Python training or agent code and want run tracking, artifact lineage and a stack abstraction without rewriting it, and if you are willing to run a server component for anything beyond local work. Skip it if you need a scheduler with long-standing operational depth or a pure experiment tracker with no pipeline concept.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 13 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 17, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The gap ZenML targets: Python code that runs, but leaves no record

Most teams already have working Python. A script loads data, trains a model or calls a model provider, evaluates the result, and writes a file. The problem starts the second time someone asks which data produced that artifact, which parameters were used, or why last week's number was better. ZenML's answer is to make the function the unit of work: you write steps, connect them into a pipeline, and the framework records each run with metrics, logs and metadata. The README describes the audience plainly as ML or AI engineers working on traditional ML, LLM workflows or agents "in a company setting", which is a narrower claim than general-purpose workflow tooling. If you are a solo researcher running one notebook, the tracking layer adds ceremony with little return. If several people share models, prompts or datasets, the record is the point.

Pipelines, stacks and snapshots: how the pieces fit together

The README states that ZenML lets you write workflows (pipelines) that run on any infrastructure backend (stacks), and that you can embed any Pythonic logic inside them, such as training a model or running an agentic loop. The operational layer around that is described as five behaviours: automatically containerizing and tracking your code, tracking individual runs with metrics, logs and metadata, abstracting away infrastructure complexity, integrating existing tools such as MLflow, Langgraph, Langfuse, Sagemaker and GCP Vertex, and supporting fast iteration with an observable layer in development and in production.

The architecture is explicitly client-server, with an integrated web dashboard in a separate repository, zenml-io/zenml-dashboard. The README splits the two modes: local development via pip install "zenml[local]" runs both client and server on your machine, while production means deploying the server separately and connecting a slimmer client with pip install zenml followed by zenml login <server-url>. That split is the design decision worth understanding before you start. It means the interesting state lives in a server process and a database, not in your working directory, which is what makes shared run history possible and also what makes the local-to-production transition a real step rather than a flag.

The quickstart example is the one the README recommends first, and it is said to demonstrate pipelines, steps, artifacts, snapshots and deployments. Snapshots and artifacts are the vocabulary to internalise: a step produces an artifact, and a snapshot captures the pipeline definition at a point in time. If those words do not map onto how your team reasons about its work, the abstraction will feel imposed.

Installing ZenML and running a first pipeline

The README gives a five-minute path. Three commands install the server-capable package, initialise a repository, and start or connect to a server.

bash
pip install "zenml[server]"  # pip install zenml will install a slimmer client
zenml init
zenml login

The comment on the first line is from the README and matters: the bare package is a client, and the extra installs server capabilities. zenml init creates the repository that ZenML expects in your project, and zenml login brings up a local server or connects you to a remote one. After that you should have a running server and a dashboard to look at, since the architecture section states the dashboard is integrated.

From there the README points at the examples directory, recommending the quickstart example first, and at the documentation for a guided build. The docs list a "Your First AI Pipeline" page and a starter guide described as going from zero to production in 30 minutes. If you prefer to skip the guided path, the examples directory in the repository is the faster survey: agent_comparison, deploying_ml_model, llm_finetuning, optuna_hyperparameter_tuning and mlops_starter are all present as runnable directories. Note what the README does not give: there is no published step-by-step for defining a step and pipeline in the README itself, so the first real pipeline definition comes from the docs or the examples rather than from the front page.

For a self-hosted server rather than the local one, the repository ships a docker-compose.yml. It defines a MySQL 8.4 service named db on port 3306 with MYSQL_ROOT_PASSWORD set to password, and a zenml_server service mapped from port 80 on the host to 8080 in the container. The compose file also carries admin credentials through an anchor:

yaml
x-admin-creds: &admin-creds
  ZENML_DEFAULT_USER_NAME: admin
  ZENML_DEFAULT_USER_PASSWORD: password

Those defaults are development values. Anyone deploying that compose file unchanged to a reachable host is publishing a server with a known admin password, which is a configuration decision the file leaves to you.

Where ZenML is the wrong tool

The clearest limitation is structural: ZenML is a Python framework with a client-server backend, not a scheduler with a decade of operational scar tissue. If your organisation already runs Airflow for general data engineering and the team is comfortable with it, adding ZenML means two orchestration systems and a question about which one owns the DAG. The README does not claim to replace a general scheduler, and the comparison is a real fork in the road rather than a feature gap.

A second boundary is the server dependency. Local mode is genuinely local, but the production path the README describes requires deploying a server and connecting clients to it. Teams that want a library with no long-running process will find that a mismatch, and the README does not document rollback or downgrade procedures for server-side state, so plan your upgrade path with that silence in mind.

Third, the integration list is a list of names, not a compatibility matrix. The README names MLflow, Langgraph, Langfuse, Sagemaker and GCP Vertex as examples of tools it integrates with, without stating supported versions. Before you commit, check the specific integration you need against the version of the upstream library you are pinned to. A pipeline framework that sits between your code and five other tools inherits the version constraints of all of them.

Finally, the project ships a lot of surface area. The top-level repository contains a helm directory, an infra directory, an alembic.ini for database migrations, a docker directory, and separate docker-compose and cloud build configurations. That is the footprint of a system with a server, a database schema and deployment tooling, and it is more to operate than a library you import.

ZenML against MLflow, and why the comparison is asked so often

The most common question people put to a search engine about this project is how it relates to MLflow, and the README answers it obliquely by listing MLflow as one of the tools ZenML integrates with. That is the actual difference in approach. MLflow is an experiment tracking and model registry layer that you call from inside your training code; ZenML is a pipeline and stack layer that wraps the code itself, and it can record into MLflow as one of its integrations. If your problem is "I cannot compare my runs", MLflow alone is the smaller answer. If your problem is "my training code is a script that nobody can reproduce end to end, and it needs to run on different infrastructure for different people", the pipeline layer is what you are missing, and running MLflow underneath it is a normal configuration rather than a competing choice.

The same logic applies to the other comparisons people search for. Kubeflow and Airflow both own scheduling and infrastructure in ways ZenML does not attempt to replace, while Metaflow shares more ground with ZenML in that it also wraps Python functions into flows, but it does not present the same client-server and stack abstraction. The honest framing is that ZenML is the layer above your training code and below your infrastructure, and the tools it is compared to mostly occupy one side of that position or the other.

Licence, maintenance and what an upgrade costs

The repository is licensed Apache-2.0, declared both in the LICENSE file and in pyproject.toml, which permits commercial use and modification under the usual conditions of that licence, including preservation of notices. There is also a CLA.md at the top level, which is relevant if you intend to contribute rather than consume. Nothing in the repository suggests a dual-licence trap for the open source package, and the README separately points to a commercial offering, ZenML Pro, with a sign-up link. Treat the boundary between the two as something to confirm against the documentation rather than assume.

On maintenance, the repository is not archived and the last push was on 2026-09-09, with releases 0.96.4 on 2026-09-04, 0.96.3 on 2026-08-07 and 0.96.2 on 2026-07-17. The cadence is roughly monthly on the patch line, with the version sitting at 0.96.x, which is pre-1.0. That matters for upgrade planning: a pre-1.0 project can change APIs between minor versions, and the changelog page linked from the README is the place to read before bumping. The server side adds a second upgrade axis, because alembic.ini at the top level indicates database migrations are part of the release process. Upgrading a client is one thing; upgrading a server with an existing database is another, and the repository does not document a downgrade path. Budget for that asymmetry.

One detail worth flagging for build hygiene: pyproject.toml configures uv with exclude-newer = "3 days", a rolling window that refuses packages published in the last three days as a supply chain measure. The file warns that when a version is filtered out, uv does not say so and resolution simply fails with "no compatible version found", which is a genuinely confusing failure mode if you hit it while debugging dependency conflicts.

Editorial conclusion

Adopt ZenML if you already have Python training or agent code and want run tracking, artifact lineage and a stack abstraction without rewriting it, and if you are willing to run a server component for anything beyond local work. Skip it if you need a scheduler with long-standing operational depth or a pure experiment tracker with no pipeline concept. Before committing, verify three things yourself: that the integrations you depend on exist at the version you need, that your Python version falls inside the >=3.10,<3.15 range declared in pyproject.toml, and that you are comfortable with the client-server split the README describes, because pip install zenml gives you the slimmer client and pip install "zenml[server]" gives you the server capabilities.

Frequently asked questions

How do I install ZenML?

The README gives three commands: pip install "zenml[server]" for the server-capable package, zenml init to initialise a repository, and zenml login to start a local server or connect to a remote one. The README notes that pip install zenml without the extra installs a slimmer client instead.

What is ZenML used for?

According to the README, it is built for ML or AI engineers working on traditional ML, LLM workflows or agents in a company setting, and it lets you write pipelines that run on any infrastructure backend. It then containerizes and tracks your code, records runs with metrics, logs and metadata, and integrates tools such as MLflow, Langgraph, Langfuse, Sagemaker and GCP Vertex.

Is ZenML open source and is it free?

The repository is licensed Apache-2.0, declared in both the LICENSE file and pyproject.toml, so the package is open source. The README also links to a separate commercial offering, ZenML Pro, so the open source package and the paid product are distinct things.

What is the difference between ZenML and MLflow?

The README lists MLflow as one of the tools ZenML integrates with, which places it underneath rather than against the pipeline layer. MLflow is where runs and models get recorded; ZenML is the layer that wraps your Python code into pipelines and decides which stack they execute on.

How do I use ZenML?

The README recommends starting with the quickstart example in the examples directory, which it says demonstrates pipelines, steps, artifacts, snapshots and deployments. The documentation links a "Your First AI Pipeline" page and a starter guide described as going from zero to production in 30 minutes.

What does ZenML do?

The README says it lets you write workflows that run on any infrastructure backend, then automatically containerizes and tracks your code, tracks individual runs with metrics, logs and metadata, and abstracts away infrastructure complexity. It also integrates existing tools such as MLflow, Langgraph, Langfuse, Sagemaker and GCP Vertex.

Official sources

  1. License: Apache-2.0
  2. Project website
  3. README
  4. Releases
  5. zenml-io/zenml on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/zenml-io-zenml.svg)](https://hysenlabs.com/projects/zenml-io-zenml)