# KAI Scheduler: gang scheduling, queues and fairness for GPUs on Kubernetes

> A scheduler plug-in built for the part of the AI lifecycle where a training job needs eight GPUs on one node or none at all, with Dominant Resource Fairness across queue hierarchies.

**kai-scheduler/KAI-Scheduler** — KAI Scheduler is an open source Kubernetes Native scheduler for AI workloads at large scale

- Repository: https://github.com/kai-scheduler/KAI-Scheduler
- Stars: 1,527 · Forks: 272
- Language: Go
- License: Apache-2.0
- Published: 2026-10-07 · Updated: 2026-10-07 · Language: en
- Canonical page: https://hysenlabs.com/projects/kai-scheduler-kai-scheduler

## Built for the gang-scheduling problem

The project describes itself as a Kubernetes scheduler that optimises GPU resource allocation for AI and machine learning workloads, designed for clusters with thousands of nodes and high workload throughput. It is not a scheduler in the sense of replacing kube-scheduler entirely; the README states it can run alongside other schedulers installed on the cluster, which is the deployment model most people will want.

The feature list starts with batch scheduling: all pods in a group are scheduled simultaneously or not at all. That single capability is what separates a GPU scheduler from a general one. A distributed training job with eight workers is useless if seven of them land on one node and the eighth is pending for an hour, and the default scheduler has no concept of that requirement.

The rest of the list is a governance vocabulary. Bin packing versus spread scheduling lets you choose between minimising fragmentation and maximising resiliency. Workload priority exists, and separately so does preemptibility, with a dedicated design document explaining why they were split into two independent policies. Hierarchical queues apply quotas, limits, priorities and fairness across multiple levels. Resource distribution lets you tune quotas, over-quota weights, limits and priorities per queue, and fairness policies use Dominant Resource Fairness with reclamation across queues.

Two entries are unusual enough to be worth calling out. Min-guaranteed-runtime establishes a window during which a running workload cannot be preempted or reclaimed even if it is marked preemptible, which is how you protect a checkpointing or evaluation phase. Time-based fairshare accounts for historical usage with decay, so a team that burned a cluster last month is not penalised forever.

## Topology awareness and elastic workloads

Version 0.10.0, in October 2025, is described as the major feature release, and its three headline items are topology-aware scheduling, hierarchical PodGroups, and time-based fairshare.

Topology-aware scheduling comes from the disaggregated serving pattern, where prefill and decode run as separate components that need to be placed on nodes with the right interconnect rather than merely on any node with a GPU. The KubeCon NA 2025 lightning talk covers exactly this, under the title Mind the Topology. The `go.mod` pulls in `k8stopologyawareschedwg/noderesourcetopology-api` and `ray-project/kube`, which is the wiring rather than the theory.

Elastic workloads scale between defined minimum and maximum pod or SubGroup thresholds, and hierarchical PodGroups let you express a gang requirement for a group of groups rather than one flat set of pods. The `examples/` directory shows how these get used in practice, with directories for batch, external PodGroup, preemption delay, quickstart, Ray, staleness grace period, time-based fairshare and topology aliases.

Dynamic Resource Allocation support is also listed, for vendor-specific hardware through Kubernetes ResourceClaims, with GPUs from NVIDIA named as the case. Release v0.17.1 contains a related bug fix: a GPU shared by multiple pods through one DRA ResourceClaim was being counted once per pod, producing negative idle GPU counts, and is now counted once per node.

## Ray, Argo Workflows and the Kubeflow operators

The `go.mod` file is a better picture of the project's real integration surface than the feature list is. Alongside the topology API it depends on `ray-project/kube`, `kubeflow/training-operator`, `kubeflow/mpi-operator`, `argoproj/argo-workflows/v3`, `NVIDIA/go-nvml`, `openshift/api` and the Prometheus operator client libraries.

Two of these are announced rather than inferred. KAI Scheduler is natively integrated for Ray workloads on Kubernetes, documented on Ray's own site, which makes sense given hierarchical gang scheduling is precisely what a multi-component agentic pipeline needs. And Grove from AI Dynamo integrates KAI's topology-aware and hierarchical gang scheduling to orchestrate disaggregated serving workloads, with a write-up on NVIDIA's developer blog.

The Kubeflow training and MPI operators in the dependency list mean KAI is positioned to sit behind those workload controllers rather than replace them. The v0.17.2 note about releasing pods no longer blocking required inter-pod anti-affinity is the sort of interaction bug that only appears once two schedulers' constraints meet.

The project also targets OpenShift, given the `openshift/api` dependency, and carries a kubestellar ACMM badge in the README for conformance.

## A repository with a governance surface, not just code

The tree shows what a project under CNCF-style scrutiny looks like. There is `GOVERNANCE.md`, `MAINTAINERS.md`, `OWNERS`, `CLA.md`, `code_of_conduct.md`, `SECURITY.md` and `SUPPORT.md` at the root, alongside `CONTRIBUTING.md` and `CHANGELOG.md`. The README badge points at the OpenSSF Best Practices project entry.

The development tooling is also more structured than most Go projects of this size. There is a `CLAUDE.md` and an `AGENTS.md` for automated coding agents, a `.agents/` directory, a `.coderabbit.yaml`, a `.changie.yaml` with a matching `.changes/` directory for unreleased changelog entries, `.golangci.yaml` for linting, a `Makefile`, and `hack/` for build scripts. `roadmap.md` is present, and `ADOPTERS.md` suggests a downstream adoption list.

Coverage is tracked through a workflow called `update-coverage-badge.yaml`, with the badge pointing at a `coverage-badge` branch. The Dockerfile is worth a look: it builds a debug stage from `golang:1.26.3`, installs Delve at version 1.27.2, drops in a prebuilt binary named by architecture, sets group and other-write permissions recursively, and switches to the numeric user 65532. That is a container hardened for a shared or restricted runtime rather than a development convenience image.

## Version 0.17 is fixing correctness issues in gangs and reclaim

The two most recent releases are a useful window into what is still being worked on, because the fixes are all about scheduling correctness rather than new features.

v0.17.2, published 2026-09-16, does four things. It respects hierarchical gang floors during allocation and during stale-gang eviction. It isolates same-named PodGroups across Kubernetes namespaces, which previously meant two teams using the same PodGroup name could collide. It stops released pods from blocking required inter-pod anti-affinity, so reclaim and preemption can actually place those pods. And it no longer grants the GPU-sharing node score to pods that request no GPU at all.

v0.17.1, published 2026-09-02, renamed the Helm value `global.fips` to `global.fipsMode` with on and off options, propagated Kubernetes client QPS and burst settings to operator-managed controllers, allowed `minMember` and `minSubGroup` of zero for workloads with no gang requirement, and fixed the DRA double-counting described earlier. It also fixed stale gang eviction no longer evicting the remaining pods of a gang whose pods completed successfully.

That last one is the kind of bug that only appears in production: a job that finished correctly still got its surviving pods evicted because the completion signal raced with the eviction check.

## What the documentation has to cover

Nearly every feature in the README is a link into `docs/`, which tells you where the real content is. There are directories for batch scheduling, priority, queues, fairness, elastic workloads and topology, plus `docs/developer/designs/` holding the longer proposals: priority-preemptibility-separation, hierarchical-podgroup, time-based-fairshare, and min-runtime. The time-based fairshare proposal is even linked to its batch working group discussion and a recording.

The README's news section is mostly conference material: a KubeCon EU 2026 talk on GPU reservations for balancing utilisation and fairness across teams, the KubeCon NA 2025 topology talk, the Grove integration post, and an April 2025 introduction recorded at the batch working group meeting.

For someone deciding whether to adopt this, the honest reading is that the design documents are where the semantics live. Whether a workload can be preempted, what guarantees it gets at minimum, and how a queue's share decays over time are all questions answered by those proposals rather than by the README. Given the project's stated purpose of running large AI clusters with many consumers, that documentation surface is the thing to read before installing anything.

## Conclusion

KAI Scheduler is worth evaluating when your GPU cluster has more than one team in it, because queue hierarchy, priority separated from preemptibility, and Dominant Resource Fairness are the problems that appear at that point and that the default scheduler does not address. It runs alongside other schedulers rather than replacing them, which lowers the risk of trying it. The repository is Apache 2.0 licensed with an NVIDIA copyright header on the build, carries an OpenSSF Best Practices badge and a `GOVERNANCE.md`, and reached v0.17.2 on 2026-09-16. With 186 open issues against 1,527 stars, expect an active project with real bug traffic: the v0.17.2 notes alone fix four scheduling correctness issues.

## FAQ

### What does Kubernetes Scheduler do?

It watches pods that have no node assigned and picks a node for each one based on predicates such as resource fit and affinity, then binds them. The default scheduler treats each pod independently, which is why workload-level scheduling features need an extension.

### How does Kubernetes decide which node to schedule a pod on?

Through a set of filter and score plugins. Filters eliminate nodes that cannot run the pod, and scores rank the survivors on things like least requested or balanced allocation. A GPU scheduler extends this with gang requirements, topology awareness and fairness across queues.

### What problem does gang scheduling solve for AI workloads?

Distributed training and disaggregated serving need pods to start together or not at all. Without it, a scheduler can place seven of eight workers and leave the eighth pending, which wastes the partial allocation. KAI Scheduler's batch scheduling makes the whole group scheduled or none of it.

### Does KAI Scheduler replace kube-scheduler?

No. The README states it can run alongside other schedulers installed on the cluster, so it is added as a scheduling capability rather than swapped in for the default one.

### How does KAI Scheduler divide GPUs between teams?

Through hierarchical queues with quotas, limits, priorities and weights, combined with Dominant Resource Fairness and resource reclamation. There is also time-based fairshare, which accounts for historical usage with time decay so heavy use in the past does not penalise a team indefinitely.

## Sources

- [Issues](https://github.com/kai-scheduler/KAI-Scheduler/issues)
- [kai-scheduler/KAI-Scheduler on GitHub](https://github.com/kai-scheduler/KAI-Scheduler)
- [License: Apache-2.0](https://github.com/kai-scheduler/KAI-Scheduler/blob/main/LICENSE)
- [README](https://github.com/kai-scheduler/KAI-Scheduler/blob/main/README.md)
- [Releases](https://github.com/kai-scheduler/KAI-Scheduler/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/kai-scheduler-kai-scheduler
