Self-hosted service
SeldonIO/seldon-core avatar
SeldonIO/seldon-core

Seldon Core 2: Kubernetes-Native MLOps for Deploying and Managing AI Systems

An MLOps framework to package, deploy, monitor and manage thousands of production machine learning models

4,782 stars866 forksGoNOASSERTION

At a glance

What is it?
Seldon Core 2 is an open-source MLOps framework that runs on Kubernetes and handles deployment, scaling, and experiment routing for machine learning models and LLM pipelines. It was last updated on 2026-03-23.
Who is it for?
Seldon Core 2 suits platform engineering teams that need a standardized way to deploy and manage many models on Kubernetes, especially when A/B testing, multi-model serving, and Kafka-based data pipelines are requirements. It is not the right choice for teams without Kubernetes expertise, for simple single-model deployments that do not need pipeline orchestration, or for anyone who needs a fully open-source license without the Business Source License restrictions.
Can I use it commercially?
Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
Is it still maintained?
Activity is slowing. The repository last received commits 6 months ago.
What is it written in?
Mainly Go, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What Seldon Core 2 solves and who it targets

Machine learning model deployment on Kubernetes involves more than wrapping a model in a container. Teams need to handle traffic routing between model versions, scale inference servers under load, share expensive GPU resources across multiple models, and run controlled experiments that compare two versions of the same model without exposing all users to the new version. Seldon Core 2 is an MLOps framework that addresses these problems together rather than leaving each to a custom implementation.

The README describes it as serving both MLOps use cases (deploying and managing traditional ML models) and LLMOps use cases (deploying large language model pipelines). It targets data science and platform engineering teams operating at scale on Kubernetes, on-premises or in any cloud. The project homepage is at seldon.io, and a commercial version with additional support is available through Seldon's website.

The README is clear that this is a Kubernetes-native tool. Teams without a Kubernetes cluster and the operational expertise to manage one will find the prerequisites significant. Seldon Core 2 is not a standalone inference server like Triton or TorchServe that can run on a single machine.

Core architecture: operator, scheduler, and Kafka

Seldon Core 2 runs as a Kubernetes operator. The operator manages Custom Resource Definitions (CRDs) that represent models, pipelines, servers, and experiments. When a user applies a model CRD to the cluster, the operator provisions the required inference server, mounts the model, and registers it for routing.

The README lists the key components. The scheduler coordinates model placement across inference servers. Pipelines connect models and custom components together using Kafka for real-time data streaming between pipeline steps. This means a pipeline can fan data out to multiple models, merge results, and pass outputs to downstream components with the delivery semantics Kafka provides.

The Makefile in the repository shows the Kubernetes deployment steps used during development:

bash
kubectl create ns seldon-mesh
kubectl create -f k8s/yaml/crds.yaml
kubectl create -f k8s/yaml/components.yaml -n seldon-mesh
kubectl create -f k8s/yaml/servers.yaml -n seldon-mesh

This installs the CRDs, the core components, and the inference servers into the `seldon-mesh` namespace. The tear-down target in the Makefile deletes these resources in reverse order and waits for clean removal before deleting the CRDs.

Multi-model serving and overcommit to reduce infrastructure cost

A common infrastructure problem in ML serving is that most models sit idle most of the time. If every model occupies a dedicated inference server, the GPU or CPU allocation remains reserved even when the model receives no requests. Seldon Core 2's multi-model serving feature consolidates multiple models onto shared inference servers, reducing the number of pods required.

The overcommit feature extends this further. It allows teams to deploy more models than the available memory can hold simultaneously. When a model that is not currently loaded receives a request, the scheduler evicts a less-recently-used model to free memory and loads the requested model. This is useful for large catalogs of infrequently used models where maintaining all of them in memory simultaneously is not economically justified.

The README frames both features as infrastructure cost savings. The trade-off is latency: a request to an evicted model will be slower on first load. Teams with strict latency SLOs for every model should evaluate whether the load time is acceptable or keep critical models pinned in memory.

Experiment routing: A/B tests and shadow deployments

Seldon Core 2 includes an experiments subsystem for routing traffic between candidate models or pipelines. The README describes support for A/B tests and shadow deployments. In an A/B test, a fraction of live traffic goes to a new model version, and the rest continues to the existing one. Shadow deployments send a copy of the same traffic to both models but only return the response from the current production model, leaving the candidate's outputs unused but observable.

This allows teams to validate a new model against live traffic without changing what users see. The experiment configuration is a CRD applied to the cluster, and the scheduler handles the traffic split. The README references an experiments section in the documentation at docs.seldon.ai and an example notebook at `samples/experiment_versions.ipynb` in the repository.

The README does not describe the statistical analysis of experiment results. Teams will need to bring their own metrics collection and significance testing, or use Seldon's commercial products for that layer.

Seldon Core 2 versus KServe and Ray Serve

KServe is another Kubernetes-native model serving framework, originally part of the Kubeflow project. Both KServe and Seldon Core 2 define custom resources for models and servers on Kubernetes and support standardized inference protocols. The key difference is scope: Seldon Core 2 places explicit emphasis on multi-model serving, overcommit, and Kafka-based pipeline composition, which are features more suited to large-scale deployments with many models. KServe focuses more tightly on individual model lifecycle management and integrates closely with the broader Kubeflow ecosystem.

Ray Serve takes a different approach: it runs on Ray, a distributed Python framework, and is designed for building serving applications in Python with granular control over deployment topology. Ray Serve is more flexible for teams already using Ray for distributed training or data processing, but it requires Ray infrastructure rather than Kubernetes native resources.

For teams whose primary concern is running hundreds of models on shared infrastructure with experiment routing and Kafka pipelines, Seldon Core 2 is the more purpose-built option. For teams that want a lightweight model server integrated with Kubeflow, KServe is worth comparing.

Business Source License and maintenance status

Seldon Core 2 is distributed under the Business Source License 1.1 (BUSL-1.1). This is not an open-source license by the Open Source Definition: it restricts production commercial use until a specified Change Date, after which the code converts to a fully open license. The README states that any contribution to the project will also be licensed under BUSL-1.1. Before deploying Seldon Core 2 in a commercial production environment, legal review of the full LICENSE file is necessary to confirm that the intended use is permitted.

The interfaces directory contains files dual-licensed under GPL-2.0-or-later (as indicated in their headers). The README does not document a Change Date for the BUSL-1.1 conversion.

The last push to the repository was on 2026-03-23, which is more than six months before this review. The most recent release in the GitHub releases list is v2.10.2, published on 2025-12-19. The v1 branch continues to receive separate maintenance as v1.19.0, published on 2026-01-23. Teams evaluating Seldon Core 2 should check whether the maintainers have resumed activity since March 2026, and review the CHANGELOG for any breaking changes between their intended version and the current release.

Repository layout and getting started

The repository is organized into several top-level directories. The `operator/` directory contains the Kubernetes operator code. The `scheduler/` directory holds the scheduling component. The `components/` directory includes custom pipeline components. The `samples/` directory provides Jupyter notebooks and YAML manifests for hands-on examples, including notebooks for HuggingFace model deployment (`samples/huggingface.ipynb`), explainers (`samples/explainer-examples.ipynb`), and inference testing (`samples/inference.ipynb`).

For local development, the Makefile provides a `deploy-local` target that delegates to the scheduler's local start commands:

bash
make deploy-local

This avoids requiring a full Kubernetes cluster for basic development work. The `init-go-modules` target sets up the Go workspace, which spans multiple modules in the repository:

bash
go work init
go work use operator
go work use scheduler

The `docs/` directory contains Sphinx-based documentation that can be built and served locally. The project's public documentation is at docs.seldon.ai, which covers installation, user guides for servers, models, pipelines, and experiments, and performance tuning.

Editorial conclusion

Seldon Core 2 suits platform engineering teams that need a standardized way to deploy and manage many models on Kubernetes, especially when A/B testing, multi-model serving, and Kafka-based data pipelines are requirements. It is not the right choice for teams without Kubernetes expertise, for simple single-model deployments that do not need pipeline orchestration, or for anyone who needs a fully open-source license without the Business Source License restrictions. Before adopting it, verify that the Business Source License terms permit your production use case, confirm that your Kubernetes cluster meets the prerequisites in the installation documentation at docs.seldon.ai, and check the CHANGELOG for breaking changes since the last release.

Frequently asked questions

What is Seldon Core?

Seldon Core is a Kubernetes-native MLOps framework for deploying, scaling, and managing machine learning models and pipelines. Version 2 adds multi-model serving, overcommit, Kafka-based pipelines, and experiment routing for A/B tests and shadow deployments.

Is Seldon Core open source?

Seldon Core 2 is available on GitHub but is licensed under the Business Source License 1.1, which restricts production commercial use until a specified Change Date. The README notes that the interfaces directory contains files licensed under GPL-2.0-or-later. Confirm the LICENSE terms before deploying commercially.

How does Seldon Core compare to KServe?

Both are Kubernetes-native model serving frameworks that define custom resources for models and inference servers. Seldon Core 2 emphasizes multi-model serving, overcommit, and Kafka pipeline composition for large-scale deployments. KServe integrates more closely with the Kubeflow ecosystem and focuses on individual model lifecycle management.

What are the alternatives to Seldon Core?

The README does not list alternatives, but commonly compared tools include KServe for Kubernetes-native model serving, Ray Serve for Python-based distributed serving on Ray clusters, and MLflow for model registry and basic serving. The choice depends on infrastructure preferences and pipeline complexity.

What is Seldon Core used for with MLflow?

The README does not describe direct MLflow integration, but Seldon Core 2 is designed to deploy models in standardized formats. Teams using MLflow for experiment tracking and model versioning can package their MLflow models and deploy them through Seldon Core's server configurations on Kubernetes.

Official sources

  1. Issues
  2. Project website
  3. README
  4. Releases
  5. SeldonIO/seldon-core on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/seldonio-seldon-core.svg)](https://hysenlabs.com/projects/seldonio-seldon-core)