MLRun: an MLOps orchestration layer that turns notebook functions into served pipelines
MLRun is an open source MLOps platform for quickly building and managing continuous ML applications across their lifecycle. MLRun integrates into your development and CI/CD environment and automates the delivery of production data, ML pipelines, and online applications.
At a glance
- What is it?
- MLRun packages data ingestion, feature storage, batch jobs, pipelines and model serving into one project abstraction, with a Python client that runs the same function locally or on Kubernetes.
- Who is it for?
- Adopt MLRun if your team already writes Python functions and needs one abstraction that covers data ingestion, feature definitions, batch jobs, pipelines and serving, and if you are willing to run or connect to a server rather than a library alone. Skip it if you only need experiment tracking, or if you cannot operate the Kubernetes and Nuclio side that the serving path assumes.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 5 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 25, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What MLRun is for, and who ends up using it
MLRun sits between a notebook and a production cluster. The README describes it as an open source AI orchestration platform for building and managing continuous AI applications across their lifecycle, and lists the job as automating delivery of production data, ML pipelines and online applications. That phrasing points at a specific audience: teams that already have models working somewhere, and are now stuck wiring the same preprocessing code into training jobs, a serving container and a monitoring job.
The unit of organization is the project. According to the README, assets and metadata (data, functions, jobs, artifacts, models, secrets) are grouped into projects, and a project can be imported or exported as a whole, or mapped to a git repository or an IDE project. That is the design decision that separates MLRun from a plain training library: it assumes your work has a boundary you can version and move between environments.
It is a poor fit for someone who only wants to log metrics from a scikit-learn script. There is no version of MLRun that is a single import with no server behind it. The client is a Python package, but the interesting parts (running functions, storing artifacts, serving models) assume an MLRun service somewhere.
Functions, runs and artifacts: the mechanism behind an MLRun job
The core abstraction is the function. A function wraps a piece of Python code with a runtime, and MLRun executes it as a run. The repository ships examples/function.py and examples/handler.py, which is the shape the documentation expects: you write a handler, decorate or register it, and MLRun handles packaging, parameters and logging.
Runs produce artifacts. Artifacts are the outputs a later step consumes, and they are addressed by project and key rather than by file path on whichever machine happened to execute the job. The README states that MLRun provides abstract interfaces to various offline and online data sources through its datastore, plus data lineage and versioning. That abstraction is what lets the same function read from a local file in development and from an object store in production without a code change.
On top of that sits the Feature Store, which the README says automates collection, transformation, storage, catalog, serving and monitoring of features across the lifecycle, with reuse and sharing as the goal. The practical effect is that a feature defined once can be retrieved for training as a batch and for inference as a real-time lookup. That dual path is the part teams underestimate: it constrains how you write the transformation, because it has to be expressible in both modes.
Pipelines are assembled from runs. The README links an Automated ML Pipeline tutorial and describes CI/CD integration, and the client ships pipeline adapters (the repository has a pipeline-adapters/ directory and requirements.txt pins mlrun-pipelines-kfp-common and mlrun-pipelines-kfp-v1-8). So the orchestration layer is Kubeflow Pipelines underneath, wrapped in MLRun's project and function model.
Installing the MLRun client and running a first function
The package is on PyPI as mlrun, and setup.py names it mlrun with the description "Tracking and config of machine learning runs". Install it into a Python environment:
pip install mlrunA run needs a project. The README points at the Quick start tutorial for this, and the examples directory contains examples/new-project.ipynb and examples/load-project.ipynb, which are the two entry points the repository itself provides. The repository's development setup uses uv with python-version 3.11 and a required uv version of 0.10.9 or newer, per pyproject.toml, if you intend to build from source rather than install the wheel.
Two practical notes before you go further. The client is not pinned to a server by pip, so the version you install and the version of the MLRun service you point at are separate decisions. And the README sends readers to the Set up your client environment guide for configuring the connection, which is where the endpoint and credentials belong rather than in your notebook.
Where MLRun gets in your way
The heaviest constraint is the serving path. The README states that MLRun serving productizes a trained model as a serverless function using real-time auto-scaling Nuclio serverless functions. Nuclio is not optional decoration here; it is the runtime. If your organization has standardized on a different serving stack, the deployment half of MLRun's story does not apply to you, and you are left using it as a job and pipeline runner.
Dependency pinning is the second constraint, and it is visible in requirements.txt rather than in prose. The file caps pandas below 2.2, numpy below 1.27.0, pyarrow below 24, and fsspec between 2025.5.1 and 2025.7.0, with a comment that fsspec should be identical to gcs and s3fs. A project that needs a newer pandas has to resolve that conflict before it can use MLRun at all. The comment about pandas 2.2 requiring sqlalchemy 2 explains why the cap exists, but it does not make the cap go away.
The release cadence is worth reading carefully. The most recent release listed is v1.12.0-rc30 from 2026-08-23, alongside v1.13.0-rc6 and v1.13.0-rc5 from the days before. These are release candidates, not final versions, and the default branch is development. The last push to the repository was on 2026-08-23, so the codebase is moving, but anyone deploying from these tags is deploying an RC.
Finally, the README is a map, not a manual. It links out to docs.mlrun.org for nearly every task and does not document rollback, migration between server versions, or what happens to stored artifacts when you upgrade. Those answers live in the documentation site, not in the repository README.
MLRun compared with MLflow and Kubeflow
The comparison people search for is MLRun versus MLflow, and the difference is scope rather than quality. MLflow is an experiment tracker and model registry: you log runs, compare them, and register a model. It does not give you a feature store with an online serving path, and it does not run your function on a cluster. MLRun covers tracking as one part of a larger object, and its project model assumes functions, artifacts and pipelines are all first-class.
The comparison with Kubeflow is closer and more confusing, because MLRun uses Kubeflow Pipelines underneath. The pinned mlrun-pipelines-kfp-common and mlrun-pipelines-kfp-v1-8 packages make that explicit. The difference is the layer above: Kubeflow Pipelines orchestrates containers, while MLRun gives you a Python function abstraction, a feature store, and a serving path over the same orchestration. If you already run Kubeflow and are happy writing pipeline YAML and container images by hand, MLRun adds a Python layer you may not want. If you find yourself rebuilding the same glue between pipeline steps, that glue is what MLRun is trying to replace.
ZenML occupies similar ground from a different angle, and the search data shows people comparing them. Without running either, the honest statement is that both wrap pipeline steps in Python abstractions; the repository here shows MLRun leaning on Kubeflow Pipelines and Nuclio specifically, which ties the deployment story to those projects.
Licence and the cost of keeping up
MLRun is Apache-2.0, and setup.py declares the same license. Apache-2.0 permits commercial use, modification and redistribution, and includes an explicit patent grant. It does not require you to publish your changes. This is a permissive licence, and it is the same one used by a large part of the surrounding ecosystem, which matters if you are assembling a stack: there is no copyleft obligation to reason about when MLRun links into your code.
That is a description of the licence text and not legal advice; if you are redistributing MLRun inside a product, have counsel read the NOTICE and attribution requirements, since the repository carries a .licenserc.json and the lint configuration enables the flake8-copyright rule with author "Iguazio".
The upgrade cost is the real ongoing expense. The client and server versions move independently, the current tags are release candidates, and the dependency pins in requirements.txt are tight enough that a major upgrade of pandas or pyarrow in your own code will collide with MLRun's caps. Budget for reading the release notes on each upgrade rather than assuming the client is backward compatible with an older server. The README does not document a compatibility matrix, so that information has to come from the documentation site or from testing.
What to check before you commit to MLRun
Start with the ecosystem page. The README explicitly points readers to docs.mlrun.org/en/stable/ecosystem.html for the list of supported data stores, development tools, services and platforms, and that list is the fastest way to find out whether your existing storage and CI systems are first-class or unsupported. If your data lives somewhere that page does not mention, assume you will be writing the integration.
Second, confirm which half of MLRun you actually need. The MLOps half (projects, functions, jobs, pipelines, feature store) works without committing to the gen AI serving path. The gen AI half leans on vector databases, LLM evaluation and fine-tuning, plus the Nuclio-based serving graph, and the README links demos for a call center and a banking agent as the reference implementations. Those two halves have different prerequisites.
Third, look at the examples before the docs. The repository ships runnable notebooks for the common cases, including mlrun_basics.ipynb, mlrun_jobs.ipynb, mlrun_dask.ipynb, mlrun_sparkk8s.ipynb and xgb_serving.ipynb. Reading one of those end to end will tell you more about whether the abstraction fits your code than the architecture page will, because you can see exactly how much boilerplate a function needs.
Editorial conclusion
Adopt MLRun if your team already writes Python functions and needs one abstraction that covers data ingestion, feature definitions, batch jobs, pipelines and serving, and if you are willing to run or connect to a server rather than a library alone. Skip it if you only need experiment tracking, or if you cannot operate the Kubernetes and Nuclio side that the serving path assumes. Before committing, verify three things: that the server version you deploy matches your client version, that the data stores you use appear in the supported ecosystem list, and that your serving target can run Nuclio functions.
Frequently asked questions
What is MLRun used for?
MLRun is an open source AI orchestration platform for building and managing continuous AI applications across their lifecycle. It organizes data, functions, jobs, artifacts and models into projects, and automates delivery of production data, ML pipelines and online applications.
How does MLRun fit into an MLOps CI/CD pipeline?
The README states that assets and metadata are organized into projects that can be imported or exported as a whole and mapped to git repositories or IDE projects, which enables versioning, collaboration and CI/CD. It links a dedicated CI/CD Integration page in the documentation for the details.
How are machine learning pipelines built in MLRun?
Pipelines are assembled from runs of MLRun functions. The client depends on the mlrun-pipelines-kfp-common and mlrun-pipelines-kfp-v1-8 packages, and the README points to an Automated ML Pipeline tutorial for the workflow.
What is an MLRun pipeline in the context of LLMs?
The README describes an application pipeline that accepts events or data, contextualizes it with a state, prepares the required model features, infers results using one or more models, and drives actions. It links serving gen AI models and a realtime serving graph in the documentation.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/mlrun-mlrun)