Kubeflow Pipelines: End-to-End ML Workflow Orchestration on Kubernetes
Machine Learning Pipelines for Kubeflow
At a glance
- What is it?
- Kubeflow Pipelines is the Apache-2.0 workflow engine for ML on Kubernetes, built on Argo Workflows and a MySQL-backed API server. It fits teams already running Kubernetes, and it is heavy for anyone who is not.
- Who is it for?
- Adopt Kubeflow Pipelines if you already operate Kubernetes and need repeatable, reusable ML workflows with a Python SDK and a managed experiment UI. Do not adopt it for a single training script or a batch job that cron and a container image already handle, because you would be operating Argo Workflows, MySQL, and the KFP API server for nothing.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository received new commits within the last day.
- What is it written in?
- Mainly Go, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The gap Kubeflow Pipelines fills between a notebook and production
A training script in a notebook is not a pipeline. It has no defined inputs, no artifact passing, no retry policy, and no record of which data produced which model. Kubeflow Pipelines addresses that gap directly: its stated goals are end-to-end orchestration of ML workflows, easy experimentation across many trials, and re-use of components and pipelines so teams do not rebuild the same steps each time.
The intended audience is an ML platform team or an MLOps engineer who already runs Kubernetes. The README frames Kubeflow as an ML toolkit for making deployments of ML workflows on Kubernetes simple, portable, and scalable, and Kubeflow Pipelines as the workflow layer inside it. If your organization has no cluster and no one to operate one, the project's own installation paths (the Kubeflow Platform or a standalone deployment) both assume that infrastructure exists.
The re-use claim is the part worth taking seriously. Components are versioned units, and a pipeline composes them, so a data validation step written once can appear in ten pipelines without being copied. That is a different value proposition from a general-purpose scheduler, which orchestrates containers but has no opinion about artifacts or experiments.
How the KFP control plane and the v2 engine actually fit together
The README's acknowledgment section states that Kubeflow Pipelines uses Argo Workflows by default under the hood to orchestrate Kubernetes resources. That is the core architectural fact: KFP is a layer over Argo, not a replacement for it.
The Go module file confirms the dependency at the source level. go.mod requires github.com/argoproj/argo-workflows/v4 v4.1.2, alongside MySQL drivers (github.com/go-sql-driver/mysql v1.10.1), a PostgreSQL driver (github.com/jackc/pgx/v5 v5.10.0), S3 clients from the AWS SDK and minio-go, and grpc-gateway for the API surface. The dependency list maps to the runtime shape: an API server that stores pipeline and run metadata in a relational database, an object store for artifacts, and a controller that submits Argo workflows.
The executor matters for portability. The README notes that the Docker container runtime was deprecated on Kubernetes 1.20+, and that Kubeflow Pipelines switched to the Emissary Executor by default from version 1.8. Emissary is described as container-runtime agnostic, which means KFP can run on a cluster using any of the container runtimes Kubernetes supports. That choice removed a hard dependency on Docker, and it is the reason older installation guides that assume Docker-in-Docker no longer describe the default behavior.
Data flow follows the same pattern as Argo. A pipeline run becomes a workflow; each component becomes a pod; artifacts move through the object store rather than through the pod filesystem. The API server is the system of record for what ran, and the UI reads from it.
Installing Kubeflow Pipelines and running a first pipeline
The README gives two installation routes and no third. You can install Kubeflow Pipelines as part of the Kubeflow Platform, or deploy it as a standalone service. Both are documented on the Kubeflow site rather than in the README, so the first step is following the installation guide for whichever route you pick. There is no pip install that gives you the service; the Python package is the SDK that talks to an already-running deployment.
For local development against the repository itself, the justfile exposes named recipes that wrap existing Make targets. The README shows the entry point. Running just with no argument lists the available recipes:
just # list available recipes
just backend-test
just backend-imagesThe justfile notes that there is intentionally no generic build or test recipe, and that heavy or Docker-building flows are exposed only through explicitly named recipes such as backend-images. A recipe like kind-standalone creates a local standalone Kind cluster with KFP deployed, which is the shortest path from a checkout to a working control plane.
Client-side, pipelines are written with the Python SDK published as kfp on PyPI. The SDK reference documentation is at kubeflow-pipelines.readthedocs.io, and the pipeline API specification is documented separately. The README points to the Kubeflow Pipelines overview for a first pipeline, so the concrete component syntax belongs there rather than in this article. What you should expect after a successful standalone install is a UI backed by the API server, with your submitted run appearing as a workflow that Argo executes.
The Argo 3.x deprecation is the upgrade decision that matters
The compatibility matrix in the README lists Argo Workflows v3.7 and v4.1, and MySQL v8. Underneath that table sits a notice that changes the calculus for anyone running an older installation: Argo Workflows 3.x remains supported for KFP 2.x but is deprecated and will not be supported by Kubeflow Pipelines 3.0. Operators must upgrade their Argo Workflows controller to a supported 4.x release before upgrading to KFP 3.0.
This is not a routine version bump. Replacing the workflow controller under a running pipeline service is an infrastructure change with its own failure modes, and the notice explicitly directs operators to a tracking issue for the removal and migration work. If you are on Argo 3.x today, the migration is a prerequisite you should schedule rather than discover during a KFP upgrade window.
The second constraint is the database. MySQL v8 is the listed version, and go.mod shows drivers for both MySQL and PostgreSQL, but the matrix names only MySQL. Running the API server means running a stateful database, backing it up, and upgrading it. That is a real operational cost that a stateless scheduler would not impose. Teams that treat the KFP install as a single Helm chart and forget the database underneath it tend to find out during a restore.
A third limitation is scope: KFP orchestrates workflows. It is not a feature store, not a model registry, and not a serving layer. The README links to a blog post describing an end-to-end lifecycle, which implies those adjacent concerns live elsewhere in the Kubeflow ecosystem or outside it entirely.
Kubeflow Pipelines compared with plain Argo Workflows
The most direct alternative is Argo Workflows itself. KFP uses it as the execution engine, so the honest question is what the extra layer buys.
The difference is in what each system considers a first-class object. Argo models workflows, templates, and steps; artifacts are supported but the abstraction is generic. KFP models pipelines, components, experiments, and runs, with a Python SDK that generates the workflow definition. If your team writes Python and wants a component to be a function with typed inputs and outputs, the KFP SDK is the shorter path. If your team already writes Argo YAML and does not need experiment tracking or a component registry, adding KFP means adding an API server, a database, and a UI for capabilities you are not using.
The trade-off is control. Going through KFP means your workflow spec is generated, so debugging a malformed pod spec sometimes means reading the compiled Argo output rather than the file you wrote. The repository acknowledges this class of problem in its tooling: the Makefile has a regenerate-all target that rebuilds generated files and golden test files, and a check-diff target that verifies generated files are current. A project with that much generated-code machinery is telling you that the source of truth and the deployed artifact are not the same file.
A second alternative is a general CI system with container steps. It will schedule your training job and show you logs. It will not give you artifact lineage between components or an experiments view, which are the specific things KFP's goals name.
Maintenance, release cadence and the Apache-2.0 licence
The repository is not archived, and the last push was on 2026-09-10. Releases are frequent: 2.17.0 on 2026-07-09, 2.17.1 on 2026-08-27, and 2.17.2 on 2026-09-04. For an operator, that cadence means patch releases arrive often enough that pinning a version and reviewing upgrades on a schedule is more practical than tracking every tag.
Upgrade cost is dominated by the dependencies rather than the KFP code. The compatibility matrix ties each KFP line to specific Argo Workflows and MySQL versions, and the Argo 3.x deprecation notice means a KFP 3.0 upgrade is gated on an Argo controller upgrade. Budget for that as a separate project. The Makefile's check-go-version and update-go-version targets, driven by a GO_VERSION variable, show that the Go toolchain version is itself managed and checked in CI, so building from source requires matching the toolchain rather than using whatever Go is installed.
Licensing is Apache-2.0, which permits commercial use, modification, and redistribution with the usual attribution and notice requirements. The repository carries a LICENSE file and an AUTHORS file at the top level. This article is not legal advice; if you redistribute KFP as part of a product, have counsel review the notice obligations and the licences of the bundled dependencies, which include Argo Workflows and MySQL drivers.
Editorial conclusion
Adopt Kubeflow Pipelines if you already operate Kubernetes and need repeatable, reusable ML workflows with a Python SDK and a managed experiment UI. Do not adopt it for a single training script or a batch job that cron and a container image already handle, because you would be operating Argo Workflows, MySQL, and the KFP API server for nothing. Before committing, verify your Kubernetes version supports the Emissary executor, confirm your Argo Workflows controller is on a 4.x release if you plan to move to KFP 3.0, and check that your team can run and upgrade a MySQL 8 instance.
Frequently asked questions
How do you install Kubeflow Pipelines?
The README gives two routes: install it as part of the Kubeflow Platform, or deploy Kubeflow Pipelines as a standalone service. Both are documented on the Kubeflow site rather than in the README itself.
What does Kubeflow Pipelines use under the hood to orchestrate workflows?
The README's acknowledgments state that Kubeflow Pipelines uses Argo Workflows by default under the hood to orchestrate Kubernetes resources. The repository's go.mod requires the argo-workflows/v4 module.
Which Argo Workflows and MySQL versions does Kubeflow Pipelines support?
The README's dependency compatibility matrix lists Argo Workflows v3.7 and v4.1, and MySQL v8. Argo Workflows 3.x remains supported for KFP 2.x but is deprecated and will not be supported by Kubeflow Pipelines 3.0.
What is the Docker container runtime situation for Kubeflow Pipelines?
The Docker container runtime was deprecated on Kubernetes 1.20+, and Kubeflow Pipelines switched to the Emissary Executor by default from version 1.8. The README describes the Emissary executor as container runtime agnostic, so it runs on clusters using any container runtime Kubernetes supports.
Is there an optional command runner for working on the Kubeflow Pipelines repository?
Yes. The repository includes an optional just command runner at the repo root that provides short aliases for existing make targets. The README notes that all just recipes are thin wrappers around existing make targets and that there is intentionally no generic just build or just test recipe.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/kubeflow-pipelines)