Kubeflow Kale: compile a tagged Jupyter notebook into a KFP pipeline
Kubeflow’s superfood for Data Scientists
At a glance
- What is it?
- Kale is a JupyterLab extension and CLI that reads cell tags in an ordinary notebook, builds a dependency graph, and emits a Kubeflow Pipelines deployment. It fits teams already running KFP 2.16 or later who want their notebooks to stay notebooks.
- Who is it for?
- Adopt Kale if your notebooks already run end to end and you have a Kubeflow Pipelines 2.16 or later endpoint to point at, because the tag-and-compile route costs far less than rewriting the same code as KFP components. Do not adopt it if your pipeline is mostly non-Python services, if you need loops and conditionals expressed as first-class KFP control flow, or if you cannot run a Kubernetes cluster at all.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 1 day ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The gap Kale fills between a working notebook and a scheduled pipeline
A data scientist finishes a model in JupyterLab. The notebook runs top to bottom on one machine. The next request is to schedule it, run it nightly, or scale it across a cluster, and that request usually means rewriting the whole thing against the Kubeflow Pipelines SDK: splitting the code into components, declaring inputs and outputs, and debugging the YAML that comes out. Kale's answer is to leave the notebook alone and let cell tags carry the structure. The README puts it plainly: tag cells with labels such as imports, step, or skip, and Kale analyzes the code, detects dependencies between steps, and generates a Kubeflow Pipeline. The audience is the person who wrote the notebook, not a platform engineer. The README's own list of selling points includes "requires no direct knowledge of KFP SDK" and "no rewriting code into pipeline components". That is the whole pitch, and it is a narrow one on purpose: Kale is not a general workflow engine, it is a compiler from notebook to KFP.
How cell tags become a pipeline graph
The mechanism is tag-driven. Kale reads the notebook through nbformat and nbconvert, parses the Python with astor and pyflakes, and builds a graph with networkx. Six tag types are documented. imports marks libraries every step needs, functions marks helper functions available to all steps, pipeline-parameters marks variables a user can tune between runs, pipeline-metrics marks metrics that appear in the KFP UI, skip marks exploratory code to exclude, and step:step_name marks an actual pipeline step. The interesting part is what you do not write. The README states that Kale automatically detects which variables flow between steps, so inputs and outputs are inferred rather than declared. That inference is the product. It is also the part most likely to surprise you, because a variable that moves between two tagged cells through a name Kale cannot trace becomes a gap in the generated graph. The output is a KFP deployment produced from Jinja2 templates, with the kfp[kubernetes] package pinned at 2.16.0 or later as a hard dependency in pyproject.toml.
Installing Kubeflow Kale and compiling a first notebook
The README's install path is the PyPI route, and it assumes a Kubeflow Pipelines instance is already running and reachable. It pairs the extension with JupyterLab 4 or later and pulls the jupyter extra, which brings in jupyter_server. The project requires Python 3.11 or newer.
pip install "jupyterlab>=4.0.0" kubeflow-kale[jupyter]
jupyter labAfter JupyterLab starts, the README says to open any notebook from examples, click the Kale icon in the left sidebar, toggle Enable to see the notebook rendered as a pipeline graph, then click Compile and Run. If you prefer the terminal, the CLI is installed as the kale entry point and takes a notebook path plus a KFP host.
kale --nb examples/base/candies_sharing.ipynb --kfp_host http://localhost:8080 --run_pipelineThe --kfp_host value must point at a reachable KFP API server, and --run_pipeline submits rather than only compiling. There is also a source checkout path for contributors, using make dev and make jupyter, plus a Docker route where make docker-build and make docker-run serve JupyterLab at http://localhost:8889. To reach a real cluster from that container, the README shows serving a development wheel with make kfp-serve and forwarding the KFP API with kubectl port-forward -n kubeflow svc/ml-pipeline 8080:8888.
Where the tag model breaks down
The tag vocabulary is flat. There is a step tag and a skip tag, and nothing in the documented table expresses a loop, a conditional branch, or a retry policy as pipeline structure. If your notebook contains a for loop over model variants, Kale sees Python inside one step, not a fan-out of steps. That is a real ceiling for anyone coming from hand-written KFP, where control flow is part of the DSL. The dependency inference has a matching failure mode: it works on names it can trace through the AST, and the README does not document what happens when a value crosses a step boundary in a form the analyzer cannot follow, such as a value written to disk in one step and read by path in another. The README also does not document rollback of a submitted pipeline, or how to reconcile a notebook change with an already-running schedule. There is a FAQ.md and a ROADMAP.md in the repository root, which is where the project says known limitations live, and the README links to both rather than restating them. Finally, this is not a tool for non-Python pipelines. A workflow that mostly calls out to Java services, Spark jobs, or shell tooling has little to gain from an AST-based notebook compiler.
Kale against hand-written Kubeflow Pipelines components
The obvious alternative is the KFP SDK itself, which Kale depends on rather than replaces. The difference in approach is who owns the structure. With the SDK you write a Python function per component, decorate it, and wire the graph explicitly; the pipeline definition is code you maintain, and the notebook, if there is one, is a scratchpad that gets thrown away. Kale inverts that: the notebook is the artifact, and the pipeline is generated from it. The trade is control for proximity. Hand-written components let you express control flow, set retries and caching per task, and version the pipeline independently of any notebook. Kale gives you a pipeline whose shape is inferred from tags and variable flow, which means less code to write and less code to review, but also less to point at when the graph is wrong. A second alternative worth naming is Airflow, which people search for alongside Kubeflow in this space. Airflow schedules arbitrary tasks on its own workers and does not assume Kubernetes-native containers or a KFP artifact store; if your work is not containerized Python, Airflow asks less of you. If you are already on Kubeflow Pipelines, that comparison is settled and Kale is the shortcut inside it.
Licence, releases and the cost of staying current
Kale is Apache-2.0, the same licence as the wider Kubeflow project, and pyproject.toml carries the text "Apache License Version 2.0" with the OSI-approved classifier. For most internal use that means no relicensing obligation beyond retaining notices; if you intend to redistribute a modified Kale, read the LICENSE and the headers marked by .license-header.txt rather than treating this paragraph as advice. On maintenance, the repository is not archived and the most recent push recorded is 2026-09-04, with v2.2.0 released on 2026-08-05 and v2.1.1 before it on 2026-06-04. Upgrade cost is dominated by one line: the kfp[kubernetes]>=2.16.0 dependency, which the README ties to Kubeflow Pipelines 2.16.0 or later. A cluster older than that floor is a blocker, not a warning. The README also notes a v2.0 support statement for KFP 2.16.0 and a comment that v2.0 was "coming soon to PyPI" in the quick-install block, while the Installation section documents the PyPI command directly, so the two paths in the README do not describe the same state of the release.
Editorial conclusion
Adopt Kale if your notebooks already run end to end and you have a Kubeflow Pipelines 2.16 or later endpoint to point at, because the tag-and-compile route costs far less than rewriting the same code as KFP components. Do not adopt it if your pipeline is mostly non-Python services, if you need loops and conditionals expressed as first-class KFP control flow, or if you cannot run a Kubernetes cluster at all. Before trusting it, verify three things yourself: that the notebook parses under the tags you chose, that the generated pipeline behaves the same as the notebook when run, and that the KFP version on your cluster matches the kfp[kubernetes]>=2.16.0 floor in pyproject.toml.
Frequently asked questions
What is Kubeflow Kale and what does it do?
Kale is a JupyterLab extension and CLI that converts a tagged notebook into a Kubeflow Pipelines deployment. It reads cell tags such as imports, step, and skip, detects dependencies between steps, and generates the pipeline without you writing KFP component code.
How do I install Kubeflow Kale?
The README gives a PyPI install that pairs JupyterLab 4 or later with the jupyter extra, then launches JupyterLab. It requires Python 3.11 or newer and a reachable Kubeflow Pipelines instance.
Which Kubeflow Pipelines version does Kubeflow Kale need?
The README lists Kubeflow Pipelines v2.16.0 or later as a requirement, and pyproject.toml pins kfp[kubernetes]>=2.16.0 as a dependency. A cluster below that version is not supported.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/kubeflow-kale)