Koordinator: QoS-Based Scheduling for Mixed Kubernetes Workloads
A QoS-based scheduling system brings optimal layout and status to workloads such as microservices, web services, big data jobs, AI jobs, etc.
At a glance
- What is it?
- Koordinator is an Apache-2.0 scheduling system that layers QoS classes, co-location and descheduling on top of vanilla Kubernetes. The design is coherent, but the documentation lives on koordinator.sh rather than in the repository, and that shapes how you should evaluate it.
- Who is it for?
- Adopt Koordinator if you already run Kubernetes at a scale where node fragmentation and noisy-neighbour interference cost you real capacity, and you are willing to run a second scheduler alongside the default one. Do not adopt it if you cannot operate a custom kube-scheduler deployment or if your workloads are uniformly latency-sensitive, because the QoS classes only pay off when there is a mix.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 6 days ago.
- What is it written in?
- Mainly Go, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 25, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What Koordinator Actually Schedules
Kubernetes treats every pod as roughly equal once requests and limits are set. That assumption breaks when you put a latency-sensitive web service and a Spark job on the same node. The web service needs predictable tail latency; the Spark job wants whatever CPU is left over. The default scheduler has no vocabulary for that distinction, so operators either over-provision or accept interference.
Koordinator's answer is a QoS layer. The README describes the project as "a QoS based scheduling system for hybrid orchestration workloads on Kubernetes," with a stated goal of improving runtime efficiency and reliability for both latency-sensitive workloads and batch jobs, simplifying resource tuning and raising pod density. The audience is therefore platform teams running heterogeneous clusters, not single-workload shops. If every pod you run is a stateless HTTP service with identical resource profiles, the QoS machinery has nothing to arbitrate.
The Component Split: koord-scheduler, koord-manager, koordlet
The repository layout and the Makefile together reveal the architecture more clearly than the README does. The Makefile defines image targets for five separate binaries: koordlet, koord-manager, koord-scheduler, koord-descheduler and koord-device-daemon. That is a multi-component system, not a scheduler plugin you drop into a config file.
koord-scheduler is the replacement scheduler that makes the placement decisions. koord-manager is the controller side, holding the CRDs and reconciling state. koordlet runs per node as an agent, which is where the runtime side of QoS enforcement happens; the go.mod dependencies on go-nvml, rdmamap, containerd/nri and netlink point at device topology, RDMA, NRI hooks and network configuration respectively. koord-descheduler is a separate image, which means eviction and rebalancing are opt-in rather than bundled into the scheduling path.
The practical consequence is that installing Koordinator changes how pods get placed cluster-wide, and it adds a node agent. Both are operational commitments. The README itself does not describe this split; you have to read the Makefile and the directory tree to see it.
Installing Koordinator and Running a First Co-located Workload
The README does not contain installation commands. It points to the project website: "You can view the full documentation from the Koordinator website" at koordinator.sh/docs, and directs readers to the installation page for installing or upgrading to the latest version. There is no helm repo URL, no kubectl apply line and no manifest path given in the README, so any command I printed here would be invented. Do not trust a tutorial that supplies one without checking the installation page first.
What the repository does give you is the build path, which is what you would use if you want to build images yourself rather than pull them. The Makefile parameterises the registry and namespace, with defaults of ghcr.io and koordinator-sh, and tags images by branch and short commit:
REG ?= ghcr.io
REG_NS ?= koordinator-sh
KOORD_SCHEDULER_IMG ?= "${REG}/${REG_NS}/koord-scheduler:${GIT_BRANCH}-${GIT_COMMIT_ID}"Those are make variables, not shell exports, so they belong in a make invocation or a local override, not in your shell profile.
For a first real use, the repository ships worked examples rather than prose. The examples directory contains nginx, spark-jobs, spark-operator-chart and runtime-hook-server subdirectories. The README points at a best-practices page titled colocation-of-spark-jobs, which is the documented path for running co-located workloads. Start there: the Spark example is the one the project itself chose to publish, and it exercises the batch-plus-service case the QoS classes were designed for. Expect to need a working Spark operator before the example means anything, since spark-operator-chart is a separate entry in that directory.
Where Koordinator Is the Wrong Tool
The clearest limitation is visible in the README's own framing. Every benefit it lists, improved resource utilization, reduced interference, flexible scheduling policies, is conditional on having a mixed workload profile. A cluster of uniform services gains nothing from QoS classes and takes on a second scheduler plus a node agent for no return.
A second constraint is the documentation split. The README is an introduction and a pointer; it documents no installation procedure, no CRD schemas, no configuration keys and no rollback path. Everything operational lives on koordinator.sh. That is a normal choice for a CNCF-adjacent project, but it means the repository alone is not sufficient to evaluate or operate the system, and the version of the docs you read may not match the tag you deploy.
Third, the component count is a real cost. Five images in the Makefile means five things to version, upgrade and monitor. The descheduler being separate is good design, since you can leave it off, but the scheduler and the node agent are not optional if you want the QoS behaviour. On managed Kubernetes offerings where you cannot replace the scheduler, Koordinator's core value is unavailable.
Finally, the last push to the repository was on 2026-04-16, the same day as the v1.8.0 release. The release cadence shown is v1.6.1 in August 2025, v1.7.0 in October 2025 and v1.8.0 in April 2026. The project is not archived, but the gap between releases is widening, and anyone planning a long-term dependency should weigh that.
Koordinator vs Volcano: Different Layers, Not Just Different Code
The comparison people search for is Koordinator against Volcano, and the difference is architectural rather than feature-level. Volcano is a batch scheduling system: it grew out of the need to run jobs, with queue management, gang scheduling and job-level fairness as first-class concepts. Its unit of concern is the job.
Koordinator's unit of concern is the node and the pod's resource class. Its stated goal is improving runtime efficiency for latency-sensitive workloads and batch jobs together, which means the service side is not an afterthought. The presence of koordlet as a per-node agent, with dependencies on NRI, NVML and netlink, is the tell: Koordinator enforces behaviour at runtime on the node, not only at admission time. Volcano's scheduling decisions are largely made in the scheduler.
That difference determines which one fits. If your problem is that a training job's pods need to be scheduled together or not at all, that is gang scheduling, and Koordinator is not the tool described here. If your problem is that a latency-sensitive service is being starved by a batch job sharing its node, that is co-location interference, and it is the case Koordinator was built for. The two are not mutually exclusive in principle, but running both schedulers over the same cluster is a configuration problem the README does not address.
Licence, Upgrade Path and What Maintenance Actually Costs
Koordinator is licensed under the Apache License, Version 2.0, per the LICENSE file and the README's licence section. That is a permissive licence with an explicit patent grant and no copyleft obligation on your own code. It does not answer the questions that matter for a fork or a redistribution, such as what happens to the project name and the NOTICE file, so treat the licence text as the thing to read rather than any summary, including this one.
Upgrade cost is driven by the five-image split. Because the Makefile tags images with branch and commit rather than a semantic version, a self-built deployment has no version identifier that maps cleanly to a release tag; you get v1.8.0 only if you build from that tag or pull the published image. Upgrading means coordinating koord-scheduler, koord-manager and koordlet together, and the README documents no rollback procedure. The CHANGELOG.md at the repository root is where the release notes live, and it is the file to read before any upgrade.
Security reporting goes to [email protected], with a SECURITY.md and an embargo-policy.md in the repository. There is also an ADOPTERS.md and an AGENTS.md, the latter suggesting the project has thought about automated contribution workflows. None of that tells you how quickly a fix ships.
Editorial conclusion
Adopt Koordinator if you already run Kubernetes at a scale where node fragmentation and noisy-neighbour interference cost you real capacity, and you are willing to run a second scheduler alongside the default one. Do not adopt it if you cannot operate a custom kube-scheduler deployment or if your workloads are uniformly latency-sensitive, because the QoS classes only pay off when there is a mix. Before committing, verify three things yourself: that your Kubernetes version is covered by the installation page at koordinator.sh/docs/installation, that the koord-scheduler image for v1.8.0 exists in the registry you pull from, and that your nodes expose the resource-isolation features koordlet needs, since the README does not enumerate them.
Frequently asked questions
How does Koordinator compare with Volcano?
The README describes Koordinator as improving efficiency for latency-sensitive workloads and batch jobs together, with a per-node agent (koordlet) that has dependencies on NRI, NVML and netlink. Volcano is not discussed in the repository, so the comparison cannot be settled from what the project documents.
How does Koordinator differ from a manager?
The repository does not describe a manager role in that sense. It does ship a koord-manager image target in the Makefile, which is the controller component that holds CRDs and reconciles state, separate from koord-scheduler and koordlet.
Is there a Koordinator alternative?
The repository does not name alternatives. The nearest thing it documents is the default Kubernetes scheduler, which Koordinator's koord-scheduler replaces rather than extends, judging by the separate image target in the Makefile.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/koordinator-sh-koordinator)