Model or dataset
mlflow/mlflow avatar
mlflow/mlflow

MLflow: an AI engineering platform you can run locally

The open source AI engineering platform for agents, LLMs, and ML models. MLflow enables teams of all sizes to debug, evaluate, monitor, and optimize production-quality AI applications while controlling costs and managing access to models and data.

28,183 stars6,389 forksPythonApache-2.0

At a glance

What is it?
MLflow is an Apache-2.0 platform for tracing, evaluating and governing LLM applications, with a separate model-training toolchain underneath. It installs from PyPI, starts with uvx mlflow server, and stores traces in a local backend on port 5000.
Who is it for?
Adopt MLflow if you want OpenTelemetry-based tracing, prompt versioning and evaluation in one place and you are willing to run a tracking server and own its storage. Do not adopt it as a drop-in replacement for a hosted observability product if nobody on the team can operate a backend database.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository received new commits within the last day.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What MLflow is for, and who ends up running it

MLflow covers two jobs that share one server and one UI. The first is the machine learning lifecycle: experiment tracking, model evaluation, a model registry and deployment to batch or real-time scoring. The second is LLM and agent work: tracing, evaluation, prompt management, prompt optimization and an AI Gateway that routes requests to model providers through an OpenAI-compatible interface.

The README frames the whole thing as an AI engineering platform for agents, LLMs and ML models, and describes teams using it to debug, evaluate, monitor and optimize applications while controlling costs and managing access to models and data. The intended audience is broad by design, which is also the first trade-off a reader should notice. A two-person team that only wants to see what an agent did on the last request gets the same install as a platform group that wants a registry, a gateway and a deployment story.

That breadth is visible in the repository layout. There is a single mlflow package, but also charts/, docker/, docker-compose/, examples/ with directories per framework, and libs/. The README's own quickstart is deliberately narrow: start a server, turn on autologging, run a model call, look at the UI. Nothing in that path requires you to understand the registry or the gateway.

How tracing and the tracking server actually fit together

The mechanism described in the README is a client-server split. Your application process imports mlflow, points at a tracking URI, and the library ships spans and runs to a server over HTTP. The server persists them and serves the UI. The README's example sets the tracking URI to http://localhost:5000, which is the default port the local server listens on.

Tracing is built on OpenTelemetry, and the README states support for any LLM provider and agent framework on that basis. It also states native integration with MCP. Practically, that means the library is not inventing a proprietary wire format for spans; it is an OTel-based collector with an MLflow-specific storage and UI layer on top. If you already emit OTel traces, that is the part worth checking against your existing pipeline before you commit.

Autologging is the other half. The README shows mlflow.openai.autolog() as a single call that instruments the OpenAI client, so calls made afterwards produce traces without manual span creation. The README describes one-line automatic instrumentation across frameworks, and the examples/ directory carries per-framework folders such as langchain/, crewai/, agno/, anthropic/ and gemini/, which is where the concrete per-framework wiring lives.

One consequence of this design: the server is on the critical path for ingestion. If the tracking server is unreachable, your application still runs, but the traces do not land, and the README does not document a client-side buffering or retry contract for that case.

Installing MLflow and getting a first trace

The README offers two paths. The fastest is a single CLI command that it says installs the MLflow skills and launches a coding agent to add tracing to an app:

bash
uvx mlflow@latest agent setup

If you would rather wire it up yourself, the README gives three steps. First, start the server. The uvx form runs MLflow without a separate install step:

bash
uvx mlflow server

Second, in your application, set the tracking URI and switch on OpenAI autologging:

python
import mlflow

mlflow.set_tracking_uri("http://localhost:5000")
mlflow.openai.autolog()

Third, run code that calls the provider. The README's example uses the OpenAI client's responses endpoint:

python
from openai import OpenAI

client = OpenAI()
client.responses.create(
    model="gpt-5.4-mini",
    input="Hello!",
)

After that, the README says to explore traces and metrics in the MLflow UI at http://localhost:5000. Expect one trace per instrumented call, with the model name and input attached. For installation on a specific platform, the README itself does not give per-OS instructions; it points at the documentation quickstart, and the package metadata in pyproject.toml sets requires-python to >=3.10, which is the constraint to check first on an older interpreter.

Where the model-training side stops and the LLM side begins

The README splits the feature set explicitly. Under LLMs and Agents it lists observability, evaluation, prompts and optimization, and the AI Gateway. Under Model Training it lists experiment tracking, model evaluation, model registry and deployment, with targets including Docker, Kubernetes, Azure ML and AWS SageMaker.

These are not the same evaluation. The model-training evaluation is described as automated evaluation tools integrated with experiment tracking. The LLM evaluation is described as 50+ built-in metrics and LLM judges, or your own, with the goal of catching regressions before production. If you arrive expecting one evaluation concept, you will find two, reachable from different parts of the docs.

The AI Gateway is the piece that changes deployment shape most. It is described as a unified gateway for LLM providers with routing, rate limits, fallbacks, credential management, guardrails and traffic splitting for A/B testing, exposed through an OpenAI-compatible interface. Adding it means your application talks to MLflow instead of directly to the provider, which is a meaningful architectural decision, not a library import. The README links a gateway quickstart but does not describe the failure behaviour when the gateway itself is down.

Limitations and cases where MLflow is the wrong tool

The README is a marketing-shaped document, and several operational questions are simply not answered there. It does not document rollback, backup, retention or how to migrate a tracking store between backends. It does not state the storage requirements for traces at production volume. It does not describe what happens to in-flight traces when the server restarts. Those gaps matter more than any feature list, because a tracing system that loses data quietly is worse than no tracing system.

Dependency policy is a second constraint, and it is stated plainly in pyproject.toml: version ranges express compatibility, not security, and the project does not raise version floors to exclude dependency versions with known CVEs, leaving patching to the application. That is a defensible position, but it means MLflow does not solve vulnerability management for you. If your process requires a library to pin around CVEs, this is a mismatch.

Third, MLflow is Python-first. The README claims support for TypeScript/JavaScript, Java and any other programming language, and OTel gives that claim a real foundation, but the quickstart, autologging and examples are Python. A JVM or Node team will be working from the OpenTelemetry path rather than the documented happy path.

Finally, MLflow is the wrong tool if you want a managed service with someone else on call. Running mlflow server means owning a process, a backend store and an artifact location. The README's localhost example hides that; production does not.

MLflow compared with Langfuse

Langfuse is the comparison people search for, and the difference is one of scope rather than features. MLflow is a platform with a model registry, a deployment story and a training-side toolchain bolted to the same server that holds your LLM traces. Langfuse is an LLM observability product.

That matters at the boundary. If your problem is purely tracing and evaluating LLM calls, a narrower tool has less surface to operate and fewer concepts to learn. If your problem includes versioning a model, promoting it through a registry, and shipping it to a scoring endpoint, MLflow already has that path in the same UI, and stitching two systems together costs integration work.

The second difference is the OpenTelemetry stance. The README states MLflow is built on OpenTelemetry and integrates natively with it, so traces can in principle flow from any OTel-instrumented service. That is a portability argument in MLflow's favour: your instrumentation is not tied to MLflow's SDK alone. Verify it against your own collector setup before treating it as settled, because the README does not spell out the exporter configuration.

A third, smaller difference: the AI Gateway. MLflow ships provider routing, fallbacks and credential management as part of the platform. If you already run a gateway, that component is redundant for you, and you would be adopting MLflow for tracing and evaluation only.

Maintenance, licensing and what an upgrade costs

The repository is not archived. The last push to the default branch was on 2026-04-06, and the most recent release listed is v3.15.2 on 2026-08-26, preceded by v3.15.1 on 2026-08-03. There is also a model-catalog/latest tag dated 2026-04-06. Release tags and default-branch activity are separate signals here, and the gap between them is worth noting if you track master rather than releases.

The licence is Apache-2.0, with LICENSE.txt at the repository root and an Apache Software License classifier in pyproject.toml. Apache-2.0 permits commercial use and modification and includes a patent grant. It also requires that you preserve notices and state changes. That is a summary of the licence text, not legal advice; if you redistribute MLflow inside a product, read LICENSE.txt and your own counsel's guidance.

Upgrade cost is dominated by the server, not the client. The development pyproject.toml carries a version of 3.16.1.dev0 while the latest listed release is 3.15.2, so the default branch is ahead of the released line. The release file, pyproject.release.toml, replaces the development metadata at release time, which means installing from source and installing from PyPI are not the same artifact. Teams that pin to a release and teams that track master will have different upgrade stories, and the README does not describe a schema migration policy for the tracking store.

Editorial conclusion

Adopt MLflow if you want OpenTelemetry-based tracing, prompt versioning and evaluation in one place and you are willing to run a tracking server and own its storage. Do not adopt it as a drop-in replacement for a hosted observability product if nobody on the team can operate a backend database. Before committing, verify what the README does not state: which backend store and artifact location you will use in production, whether the AI Gateway fits your provider mix, and whether your Python version satisfies the requires-python constraint of >=3.10.

Frequently asked questions

What is MLflow used for?

The README describes it as an AI engineering platform for agents, LLMs and ML models, used to debug, evaluate, monitor and optimize AI applications. It also covers the ML lifecycle: experiment tracking, model evaluation, a model registry and deployment.

Is MLflow a part of Databricks?

The README presents MLflow as an open source project with its own website, docs and demo, and the repository is developed in the public mlflow/mlflow GitHub repository. Databricks appears in the package keywords and as a dependency (databricks-sdk), but the README does not describe MLflow as a Databricks product.

Is MLflow free or paid?

The repository is licensed under Apache-2.0, so the software itself is free to use and modify under that licence. The README does not describe a paid tier of MLflow.

Is MLflow a MLOps tool?

It covers MLOps ground: the README lists experiment tracking, model evaluation, a model registry and deployment to Docker, Kubernetes, Azure ML and AWS SageMaker. It also covers LLM and agent observability, evaluation and prompt management, which sit outside classic MLOps.

How to install MLflow?

The README's quickstart starts the server with uvx mlflow server, which runs it without a separate install step. The package is also published on PyPI as mlflow, and pyproject.toml sets requires-python to >=3.10.

How to use MLflow locally?

Start the server with uvx mlflow server, then in your application set the tracking URI to http://localhost:5000 and call mlflow.openai.autolog(). The README says traces and metrics then appear in the MLflow UI at that same address.

Official sources

  1. Official documentation
  2. Official README
  3. Project repository
  4. Release notes
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/mlflow-mlflow.svg)](https://hysenlabs.com/projects/mlflow-mlflow)