OME turns a model file into a custom resource and a runtime decision
Open Model Engine (OME) — Kubernetes operator for LLM serving, GPU scheduling, and model lifecycle management. Works with SGLang, vLLM, TensorRT-LLM, and Triton
At a glance
- What is it?
- A Kubernetes operator for large language model serving, licensed Apache-2.0, that parses model files to extract architecture, parameter count and capabilities, scores that against available runtimes, and schedules the result onto GPUs with a bin-packing scheduler installed alongside the default one. The API is v1beta1 and the charts ship from a moirai-internal registry.
- Who is it for?
- Adopt OME if you run more than one inference runtime in one cluster and want the choice of runtime, the placement of the workload and the rollout to be declarative objects you can inspect, or if you want to reuse the SGLang and vLLM integration rather than writing your own. The weighted runtime selection and the prefill and decode disaggregation support are the parts that would cost you the most to build.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository received new commits within the last day.
- What is it written in?
- Mainly Go, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 2, 2026, and from our analysis. They are not legal advice.
Editorial analysis
Model files become custom resources, and parsing reads them to decide
The design decision that shapes the whole operator is that a model is a first-class custom resource rather than a container image someone tagged. Models are parsed directly from the model files, and the parser extracts three things: the architecture, the parameter count, and the capabilities. Those three facts are then the input to every downstream decision, which is why the formats it accepts matter. SafeTensors, PyTorch, TensorRT and ONNX are all named, and storage is distributed with automated repair, double encryption and namespace scoping. There is also a pre-configured catalog of more than two hundred models in the repository's supported models reference, covering the Llama, Qwen, DeepSeek, Gemma and Phi families. The practical consequence is a two-tier experience. A model in that catalog is a manifest you write; a model outside it needs its architecture recognised by the parser before the operator can do anything with it, so the failure mode for an unknown architecture arrives early rather than at serving time. Deployment patterns sit on the same abstraction: prefill and decode disaggregation, multi-node inference and traditional Kubernetes deployments are all supported, with canary and blue-green rollout strategies and scaling controls alongside them. That combination, declarative model plus declarative rollout, is the reason to run an operator rather than a Helm chart per model.
Runtime choice is a weighted score, not a setting you pick
Intelligent runtime selection is described as automatic matching of models to runtime configurations through weighted scoring, and the inputs to that score are enumerated: architecture, format, quantization, parameter size, and framework compatibility. Five inputs, weighted, is a design you can reason about in a way a hard-coded mapping does not allow. It also means the operator can be wrong in a way that is hard to see, because a score produces a choice without a stated reason, so the CLI feature below exists partly to fill that gap. Two runtimes are called out as first-class integrations. SGLang is described in detail, with cache-aware load balancing, multi-node deployment, prefill-decode disaggregated serving and multi-LoRA adapter serving. vLLM is named for high-throughput inference. The project description also names TensorRT-LLM and Triton among the engines it works with, so the integration surface is wider than the two the feature list details, and the exact support level for each is the kind of thing to confirm rather than assume. Automated benchmarking is built in rather than bolted on, through a BenchmarkJob custom resource with configurable traffic patterns, concurrent load testing and stored results, which is what makes a runtime comparison reproducible instead of anecdotal.
The GPU scheduler is installed as a second scheduler and workloads opt in by name
Resource optimization is handled by a specialized GPU bin-packing scheduler with dynamic re-optimization, and the install notes make its mechanics unusually clear. It is not a replacement for the default Kubernetes scheduler. It is installed as a second scheduler, and workloads opt in by setting their scheduler name in the spec. It also requires the scheduler-plugins PodGroup custom resource to be present, which is an external dependency you have to satisfy before the chart is useful. Bin packing is the right default for expensive accelerators, where fragmentation costs more than imbalance, and a scheduler that re-optimizes dynamically is addressing the case where a cluster's shape changes after placement. The cost is that a second scheduler with a required CRD is more moving parts between you and a running inference service. That is why it ships as an optional add-on rather than part of the core install, and it is worth remembering that the workloads already scheduled when you add it are not moved by installing it.
Alfred recommends GPU migrations and does not perform them by default
Fleet capacity and GPU operations are the newest part of the project and are labelled alpha, which is the word to weigh before planning around them. Two components ship there. The quota manager assembles an AcceleratorQuota tree and renders it into Kueue, or projects per-cluster shares onto members from a management cluster, which is the bridge between a fleet policy and the gang scheduling that Kueue already provides. Alfred is described as the GPU cluster caretaker: it observes the physical GPU layer and recommends corrective migrations. The detail that matters is the qualifier attached to that, which is recommend-only by default. So the operator can tell you that a workload belongs on a different node or a different card, and the action of moving it is left to a human or to something you configure. That is a deliberate choice about blast radius, and it is also why Alfred is listed as an optional add-on you install separately rather than something the core install turns on. The rest of the platform is exposed the same way, through Kubernetes components rather than through new ones: Kueue handles gang scheduling for multi-pod workloads, LeaderWorkerSet provides resilient multi-node deployments, KEDA provides autoscaling on custom metrics, the Gateway API handles traffic routing, and the Gateway API Inference Extension covers standardized inference endpoints.
kubectl ome is a kubectl plugin that explains its own decisions
The command-line interface is distributed as a kubectl plugin, so it arrives as kubectl ome rather than as a separate binary you have to keep on your path. Its feature list is the most interesting part of the project, because it is written around explaining rather than operating. It inspects models, runtimes, inference services and logical instances. It explains runtime and accelerator selection, which is the answer to the question a weighted scoring function raises. It reports rollout, autoscaling, placement, quota and traffic evidence. It streams component logs. And it submits guarded actions, named as rollout pause, resume, promote and rollback, traffic drain, and migration requests. The word doing the work there is guarded, since it describes an interface that stages an action for review rather than firing it. Accelerator management is the matching piece on the cluster side: AcceleratorClass resources define GPU capabilities, discovery patterns and cost information, and selection policies are named BestFit, Cheapest and MostCapable, which is a cost-aware scheduler built on top of the standard one. The graphical side is a separate repository, ome-console, described as a modern web interface for models, serving runtimes and inference services with real-time updates and HuggingFace model search integration. Keeping it outside this repository means the operator can be adopted without a console, which is the right split for a component that reconciles in a cluster.
The charts ship from a moirai-internal registry, not the ome-projects one
The install instructions are worth reading twice, because the registry host does not match the project name. The recommended path installs two Helm charts from an OCI registry, first the custom resource definitions and then the resources themselves:
# Install OME CRDs
helm upgrade --install ome-crd oci://ghcr.io/moirai-internal/charts/ome-crd --namespace ome --create-namespace
# Install OME resources
helm upgrade --install ome oci://ghcr.io/moirai-internal/charts/ome-resources --namespace omeThe alternative path clones the repository and installs the same two charts from the local charts directory, which is the one to use if you are modifying the operator. Both land in a namespace called ome, and the CRD chart is the one that creates the namespace on a first install. There is a documented procedure for moving an existing manifest installation onto these charts, which tells you manifest installs were the earlier path. Two optional charts sit on top: a serving chart that deploys pre-configured ClusterBaseModels, ClusterServingRuntimes and InferenceServices, and the three alpha add-ons. Every command points at the moirai-internal registry path while the repository lives under ome-projects, so anyone behind a mirror or an air-gapped registry needs to rewrite those paths before the first install.
Go 1.26, Kubernetes 1.28, and a v1beta1 API with real test coverage
The engineering constraints are stated precisely enough to plan an upgrade against. Installation requires Kubernetes 1.28 or newer, and the Go module is sigs.k8s.io/ome declaring Go 1.26.0, with the Kubernetes libraries at 0.36.2 across api, apimachinery, client-go, apiserver, cli-runtime and others. The dependency list is the one you would expect from an operator that also runs a fleet control plane: KEDA for autoscaling, Kueue through its Go client, LeaderWorkerSet, Prometheus for metrics, Open Policy Agent's cert-controller, a Common Expression Language library for validation, Istio's client, Knative libraries, cobra and viper for the CLI, and ginkgo with gomega for tests. The production readiness list claims a v1beta1 API, documentation, unit and integration test coverage, production deployments with large-scale workloads, monitoring through standard metrics and Kubernetes events, role-based access control with model encryption, and a high availability mode backed by redundant model storage. The tree backs that up with a scheduler directory, charts, cmd, internal and pkg, config, tests, hack, dockerfiles, a devcontainer, a golangci configuration and a pre-commit setup. The Makefile is where the release process shows itself: the image tag defaults to whatever git describe reports with the dirty flag, that string and the commit hash are compiled into a version package through linker flags, the Go version is read out of go.mod so there is one source of truth, the build prefers nerdctl over docker when it is on the path, and two options exist for generating the CRDs and for enabling a self-signed certificate authority. Releases are on a steady monthly cadence, 1.2.0 in July 2026, then 1.2.1 four days later and 1.2.2 in August, with the default branch last pushed on 1 October 2026.
Editorial conclusion
Adopt OME if you run more than one inference runtime in one cluster and want the choice of runtime, the placement of the workload and the rollout to be declarative objects you can inspect, or if you want to reuse the SGLang and vLLM integration rather than writing your own. The weighted runtime selection and the prefill and decode disaggregation support are the parts that would cost you the most to build. Do not adopt it for a single model on a single node, since the operator, three optional add-ons and a second scheduler are a lot of machinery for one deployment. Four things to check first. Your Kubernetes version, since 1.28 or newer is required. Which optional pieces you actually want, because the scheduler, the quota manager and Alfred are each separately installable and each is marked alpha. Whether your model is in the pre-configured catalog of more than two hundred entries or needs a new entry. And whether the registry you can reach is the internal one the install commands point at.
Frequently asked questions
What does OME stand for?
Open Model Engine. It is a Kubernetes operator for enterprise-grade management and serving of large language models, automating model management, runtime selection, resource utilization and deployment patterns. The project description names SGLang, vLLM, TensorRT-LLM and Triton as the engines it works with.
How does OME choose an inference runtime for a model?
Automatically, through weighted scoring over architecture, format, quantization, parameter size and framework compatibility. The parser extracts the architecture, parameter count and capabilities directly from the model file, and a catalog of more than two hundred pre-configured models covers the Llama, Qwen, DeepSeek, Gemma and Phi families.
What are the alpha add-ons in OME?
Three optional charts published alongside the core ones. The GPU bin-packing scheduler is installed as a second scheduler that workloads opt into via spec.schedulerName and that requires the scheduler-plugins PodGroup CRD; the quota manager assembles the AcceleratorQuota tree and renders it into Kueue; and Alfred observes the physical GPU layer and recommends corrective migrations, recommend-only by default.
How do I install OME?
Kubernetes 1.28 or newer is required. The recommended path installs the CRD chart and the resources chart from an OCI registry with two helm upgrade --install commands into a namespace called ome, and the source path clones the repository and installs the same charts from the local charts directory. An optional serving chart deploys pre-configured ClusterBaseModels, ClusterServingRuntimes and InferenceServices.
What does kubectl ome do?
It is the official OME CLI distributed as a kubectl plugin. It inspects models, runtimes, inference services and logical instances, explains runtime and accelerator selection, reports rollout, autoscaling, placement, quota and traffic evidence, streams component logs, and submits guarded actions such as rollout pause, resume, promote and rollback, traffic drain and migration requests.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/ome-projects-ome)