MLServer: A V2-Compliant Python Inference Server for Multi-Model Serving
An inference server for your machine learning models, including support for multiple frameworks, multi-model serving and more
At a glance
- What is it?
- MLServer is an open-source Python inference server from Seldon that exposes both REST and gRPC endpoints compliant with the KFServing V2 Dataplane spec. It supports ten framework runtimes out of the box, parallel inference workers, adaptive batching, and Kafka integration, and serves as the core inference server within KServe and Seldon Core.
- Who is it for?
- MLServer fits teams that need a V2-compliant inference server deployable locally or in Kubernetes via Seldon Core or KServe, with multi-model serving and adaptive batching out of the box. It supports Python 3.9 through 3.12; Python 3.13 is explicitly not supported.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 1 day ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.
Editorial analysis
Multi-Model Serving With a Standard Inference Protocol
Most inference servers tie a running process to a single model, forcing separate deployments when a team needs to serve multiple models on the same machine. MLServer lets multiple models run in the same process, each independently loadable and unloadable from a model repository without restarting the server.
All models are exposed through a single REST and gRPC interface that follows the KFServing V2 Dataplane specification, which has been standardized and adopted by several model serving frameworks. V2 compliance means MLServer can slot into Kubernetes orchestration layers that expect V2 endpoints, including Seldon Core and KServe, where MLServer is the core Python inference server.
The project is from Seldon Technologies and is Apache-2.0 licensed. The most recent release is 1.7.1 from 2025-06-06. The last push was on 2026-09-26.
Architecture: FastAPI, gRPC, and the Inference Runtime Layer
The core server is built on FastAPI and uvicorn for the REST layer, grpcio for gRPC, and pydantic 2.x for request and response validation. The pyproject.toml pins fastapi below 0.125.0 and requires uvicorn 0.38.0 or later.
Between the server layer and the ML framework lies the inference runtime: a set of adapter classes that define how a specific model type should load, run inference, and return results. Runtimes are separate installable packages (for example, mlserver-sklearn, mlserver-mlflow) rather than bundled dependencies. This lets teams install only the runtimes they need.
The Dockerfile builds from python:3.10-slim for wheel building and then switches to a ubi9-minimal base for the final image. Key environment variables set in the Dockerfile include MLSERVER_MODELS_DIR=/mnt/models and MLSERVER_ENV_TARBALL=/mnt/models/environment.tar.gz, which set the default model storage paths inside the container. The Docker image is published as seldonio/mlserver.
MLServer also integrates OpenTelemetry (opentelemetry-sdk, instrumentation for FastAPI and gRPC) and Prometheus metrics via starlette-exporter and py-grpc-prometheus.
Installing MLServer and Serving a First Model
Install the core package:
pip install mlserverTo serve a scikit-learn model, also install the sklearn runtime:
pip install mlserver-sklearnThe README points to the full list of examples in docs/examples/index.md for next steps after installation. During development, run the full test suite from the repository root:
make testTo run tests for a single file:
tox -e py3 -- tests/batch_processing/test_rest.pyPython 3.9 through 3.12 are supported. The pyproject.toml requires Python between 3.9 and 3.13 exclusive, meaning Python 3.13 is not supported. Python 3.7 and 3.8 are unsupported.
Inference Runtimes: Frameworks Covered and License Differences
MLServer ships ten pre-packaged runtimes. The framework table from the README lists Scikit-Learn, XGBoost, Spark MLlib, LightGBM, CatBoost, Tempo, MLflow, Alibi-Detect, Alibi-Explain, and HuggingFace as supported.
Each runtime is a separately installed package, named with the mlserver- prefix followed by the framework name. The README includes a license note that is significant for production deployments: Alibi-Detect and Alibi-Explain are both licensed under the Business Source License 1.1, not Apache-2.0. The rest of MLServer is Apache-2.0. Any deployment that uses the Alibi runtimes must comply with the BSL 1.1 terms for those packages.
Custom runtimes are documented at ./docs/runtimes/custom.md in the repository. Writing a custom runtime is the path for frameworks not on the built-in list.
The HuggingFace runtime (mlserver-huggingface) enables serving HuggingFace transformers through the same V2 interface, making it possible to put a language or vision model behind a standard inference endpoint without writing custom serving code.
Adaptive Batching, Parallel Workers, and Kafka Integration
MLServer supports adaptive batching, which groups inference requests together on the fly before passing them to the model. This is useful when the model's throughput scales better with batch inputs than with individual calls. Configuration is documented at mlserver.readthedocs.io/en/latest/user-guide/adaptive-batching.html.
For vertical scaling, MLServer can run inference in parallel across a pool of workers. Each worker handles one concurrent inference request, and the number of workers can be tuned to match available CPUs or GPU slots. The documentation for parallel inference is at mlserver.readthedocs.io/en/latest/user-guide/parallel-inference.html.
The aiokafka dependency in pyproject.toml indicates support for Kafka-based inference request routing. This enables asynchronous inference pipelines where requests arrive via Kafka topics rather than direct HTTP or gRPC calls, which fits streaming or event-driven architectures.
The tritonclient dependency (version 2.5 to below 2.61) points to existing integration with NVIDIA Triton Inference Server for use cases that involve Triton as a backend.
Deploying on Kubernetes With Seldon Core and KServe
MLServer is designed to deploy on Kubernetes through two specific platforms. KServe (formerly KFServing) uses MLServer as one of its Python inference server options for serving scikit-learn, XGBoost, and other Python-based models. Seldon Core also uses MLServer as its V2-compliant inference backend.
In both cases, the V2 Dataplane compliance is what enables the integration. KServe and Seldon Core expect inference servers to expose V2-compatible endpoints; MLServer's REST and gRPC interfaces meet that contract without extra configuration.
The Dockerfile target is designed for this Kubernetes deployment path: the ubi9-minimal base and the environment variable defaults (MLSERVER_MODELS_DIR=/mnt/models) align with the volume mount conventions that Kubernetes deployments typically use for model artifacts.
For local development without Kubernetes, the mlserver CLI is sufficient. The README's example section covers running MLServer directly against a local testdata directory.
MLServer vs Triton Inference Server
NVIDIA Triton Inference Server is a well-known, high-performance inference serving platform that supports TensorFlow, PyTorch, ONNX, TensorRT, and other frameworks on both CPU and GPU. It is optimized for throughput and latency on NVIDIA hardware and ships with dynamic batching and model ensemble support.
MLServer and Triton address the same problem from different positions. Triton is built for maximum inference throughput on GPU-heavy deployments with NVIDIA backends. MLServer is built for the Python ML ecosystem (scikit-learn, XGBoost, MLflow, HuggingFace) and emphasizes V2 compliance, multi-model serving in a single Python process, and Kubernetes integration through Seldon Core and KServe. The two are not mutually exclusive: MLServer's tritonclient dependency indicates it can work alongside Triton in architectures that need both.
For teams whose workloads are primarily Python-based classical ML models or HuggingFace models, MLServer's runtimes are a lower-friction entry point. For teams serving large neural networks on GPU clusters where throughput is the primary constraint, Triton's hardware optimization is the stronger case.
Editorial conclusion
MLServer fits teams that need a V2-compliant inference server deployable locally or in Kubernetes via Seldon Core or KServe, with multi-model serving and adaptive batching out of the box. It supports Python 3.9 through 3.12; Python 3.13 is explicitly not supported. Before choosing any optional runtime, check its individual license: Alibi-Detect and Alibi-Explain are under the Business Source License 1.1, not Apache-2.0. Run pip install mlserver and then the framework-specific package to start a first serving session.
Frequently asked questions
What is MLServer compared to KServe?
MLServer is the Python inference server component; KServe is the Kubernetes-based platform for deploying and managing ML models. KServe uses MLServer as one of its inference server runtimes for Python framework models. MLServer can also be run standalone without KServe.
What is MLServer compared to MLflow?
MLflow is an experiment tracking and model registry platform that also includes a model serving component. MLServer is a dedicated inference server with a separate MLflow runtime (mlserver-mlflow) that serves models stored in the MLflow model registry through the V2 protocol. They are complementary: MLflow manages the model lifecycle, MLServer handles the inference endpoint.
What are the alternatives to MLServer?
NVIDIA Triton Inference Server is the most commonly compared alternative, optimized for GPU workloads with TensorFlow, PyTorch, and ONNX. BentoML is another Python-native inference server with a different API model. MLServer is distinguished by its V2 Dataplane compliance and tight integration with Seldon Core and KServe.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/seldonio-mlserver)