Self-hosted service
kubernetes-sigs/kueue avatar
kubernetes-sigs/kueue

Kueue puts a queue in front of your Kubernetes batch jobs

Kubernetes-native Job Queueing

2,998 stars813 forksGoApache-2.0

At a glance

What is it?
Kueue is a job-level manager that admits work only when quota allows it, so a training run or a Ray job waits its turn instead of fighting the scheduler for room.
Who is it for?
Kueue is at its best when a cluster already has expensive batch work and a reason to be fair about who gets the accelerators. The project is honest about its shape: it does not schedule pods, it decides whether a pod is allowed to exist, and everything it manages has to plug into a Workload object and a ClusterQueue.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 15 days ago.
What is it written in?
Mainly Go, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 23, 2026, and from our analysis. They are not legal advice.

Editorial analysis

Admission, not scheduling

The README opens by stating the scope in one sentence and then immediately rules out the obvious alternative. Kueue is a set of APIs and controllers for job queueing. It is a job-level manager that decides when a job should be admitted to start, meaning pods can be created, and when it should stop, meaning active pods should be deleted. Admission checks and eviction are the verbs, not scheduling.

That distinction has real consequences for anyone arriving from Slurm, Volcano or a GPU operator. Those systems sit between you and the scheduler and shape placement. Kueue does not compute where a pod goes. It decides whether the job is allowed to exist at all. Once admitted, ordinary Kubernetes scheduling applies, and the pods you get are scheduled the way any other pods are.

The feature list follows from that positioning. Job management supports priority-based queueing with two named strategies, StrictFIFO and BestEffortFIFO. Advanced resource management adds resource flavor fungibility, Fair Sharing, cohorts, and preemption with a variety of policies between tenants. The project also handles partial admission, which lets a job run with reduced parallelism based on the quota actually available, and dynamic reclaim, which releases quota as pods complete rather than at the end of the job. Whether partial admission is what you want is a policy decision your cluster has to make out loud, because it trades job completion time for utilization.

What is built in for batch workloads

The integrations section is the most concrete part of the README, and it is worth reading as a compatibility list rather than a feature list. There is built-in support for BatchJob, Kubeflow training jobs, RayJob, RayCluster, JobSet, and plain Pods or Pod Groups.

For a mixed cluster that matters, because the last item is the escape hatch. If your workload does not fit any of the named integrations, a plain Pod or Pod Group with Kueue-managed labels can still be admitted and counted. That is the difference between a queue you adopt incrementally and one you adopt by migrating every job type at once.

The serving side is handled separately. Kueue describes mixing training and inference as simultaneous management of batch workloads along with serving workloads such as Deployments or StatefulSets, and the documentation links to a dedicated task page for running a Deployment under Kueue rather than assuming a Job shape. Treating quota as a cluster-wide currency that both a model server and a fine-tuning run draw from is the whole point of the project, and it is the reason the resource model has a concept as general as flavor fungibility.

On the observability side there are built-in Prometheus metrics and an on-demand visibility endpoint for looking at pending workloads. That second feature is the one operators tend to need first: when something is not running, the useful question is which queue it is waiting in and what is ahead of it.

AdmissionChecks and the autoscaler handshake

AdmissionChecks are a mechanism for internal or external components to influence whether a workload can be admitted. Read that carefully, because it is the extension point that makes Kueue usable in a cluster that is also growing.

The concrete example in the README is autoscaling. Kueue integrates with cluster-autoscaler's provisioningRequest through admissionChecks. The pattern is: a job asks for more than exists, Kueue withholds admission, the provisioning request asks for nodes, the nodes arrive, and admission proceeds. Without this, a job that needs capacity in a cluster too small to hold it simply waits forever, which is indistinguishable, from the outside, from a broken queue.

The same mechanism is general enough for anything else that needs to veto an admission. The cost is that you now have two controllers making scheduling-adjacent decisions, and the README does not claim otherwise. The feature list also mentions MultiKueue, which lets a cluster search other clusters for capacity and off-load the main one, and topology-aware scheduling to optimize pod-to-pod throughput against data-center topology. Topology awareness is a good example of scope discipline: Kueue does not place pods, but it can ask that placement respect the network when it grants quota.

Two release lines shipping patches minutes apart

The release data is the most interesting thing in the repository after the feature list, and it is not what you would guess. On 2026-09-17, v0.19.5 was published at 14:55:26 UTC. On the same afternoon, v0.18.9 went out at 14:35:05 UTC, twenty minutes earlier. Both lines are being patched, which means an operator has a real choice to make rather than a single obvious upgrade target.

Both notes carry the same shape of instruction: review the `.0` release notes for each minor version you cross, and review the patch notes only within the line you are already on. Kueue publishes upgrade work as a document rather than a changelog entry, and the notes for both versions open with a line telling the reader they must read them. Take that literally if you are upgrading across a minor boundary.

One of those two releases contains a genuine API contract change for anyone extending Kueue rather than only running it. v0.19.5 fixes a bug where workloads were finalized as orphaned while their owning Job was being deleted, so Kueue now waits until the owner Job is gone. The same note says that if you implement the ComposableJob interface in a custom integration, your Load method must change to return `(*jobframework.LoadResult, error)` instead of `(bool, error)`. That is a compile error for custom integrations, delivered in a patch release, and it is exactly the kind of thing the line-scoped reading instruction is designed to surface.

What v0.20.0-rc.0 refuses to start over

The release candidate published on 2026-09-10 is where the project's caution becomes visible. It fixes a bug in which DRA device-class mapping, or a resource transformation using the reserved resource name `pods`, was silently discarded or left the workload permanently pending. The upgrade note does not describe a workaround, it describes a precondition: remove or rename those entries before upgrading, or the controller manager will fail to start.

The second entry is a quota bypass, which is a different category of problem. Raising `spec.leaderWorkerTemplate.size` on an already-admitted, Kueue-managed LeaderWorkerSet could run more pods per group than the reserved quota covered. The fix makes that field immutable while Kueue manages the resource, behind a new `LWSImmutableGroupSize` feature gate in Beta.

There is a naming trap in the first entry worth spelling out: renaming a mapping name or an `outputs` key also requires updating the matching ClusterQueue `nominalQuota` entries in the same change. Two objects have to move together, and the release note says so once. If you are on the release candidate, read that line twice.

The API surface has otherwise been stable enough for the README to claim v1beta2 while respecting the Kubernetes Deprecation Policy, which is a stronger statement than most projects in this space make.

Repository layout, build, and what is vendored in

The tree is what you would expect from a Kubernetes SIG project and slightly more opinionated in places. There are `apis/`, `client-go/`, `cmd/`, `config/`, `charts/`, `hack/`, `internal/`, `pkg/`, `test/`, and `vendor/`, plus `keps/` for enhancement proposals, a `CHANGELOG/` directory, `release-timelines/`, and a `site/` directory with a Netlify configuration for the docs. Two files sit at the root that are worth knowing about: `AGENTS.md` and `CLAUDE.md`, plus `.krew.yaml`, which is how Kueue installs itself as a kubectl plugin.

The build is a two-stage Dockerfile that fetches dependencies before copying sources, with a retry wrapper around `go mod download`:

dockerfile
# defaulted ARGs need to be declared first
ARG BUILDER_IMAGE=public.ecr.aws/docker/library/golang:1.27
ARG BASE_IMAGE=gcr.io/distroless/static:nonroot

The dependency list in go.mod is where Kueue's integration claims become literal. Alongside the usual controller-runtime and Kubernetes libraries at v0.37.0, there is `github.com/kubeflow/training-operator v1.9.4`, `github.com/ray-project/kuberay/ray-operator v1.7.0`, `github.com/kubeflow/spark-operator/v2`, `github.com/kubeflow/mpi-operator`, `sigs.k8s.io/jobset v0.12.0`, and `sigs.k8s.io/lws v0.10.0`. Each of those is a project with its own release cadence, and Kueue is pinned to a specific version of each.

The Makefile separates platform matrices for controller images, the CLI, the visualization frontend and backend, and Ray, which notes that Ray only provides PyPI wheels for amd64 and arm64. Image staging goes to us-central1, and every image repository is a variable, so the registry is a build argument rather than a constant.

Editorial conclusion

Kueue is at its best when a cluster already has expensive batch work and a reason to be fair about who gets the accelerators. The project is honest about its shape: it does not schedule pods, it decides whether a pod is allowed to exist, and everything it manages has to plug into a Workload object and a ClusterQueue. The two release lines shipping patches twenty minutes apart on the same afternoon are the operational detail worth planning around, and the upgrade notes for v0.20.0-rc.0 name the resource names that will refuse to start the controller at all. Read the overview on kueue.sigs.k8s.io, then check the quota names in your ClusterQueue against the reserved-resource rules before choosing a version.

Frequently asked questions

How does kueue work?

Kueue sits between you and the scheduler. It tracks batch workloads against ClusterQueues that carry quota, and it admits a job only when the quota covers it, at which point pods are created and normal scheduling takes over. Work that does not fit waits, and work that is admitted can later be stopped, with partial admission and dynamic reclaim letting it run smaller or release quota as pods finish.

Does Kueue replace the Kubernetes scheduler?

No, and the README is explicit that Kueue decides when a job is admitted to start and when it should stop, not where pods land. Once admitted, your pods are scheduled by the standard Kubernetes scheduler. Kueue adds topology-aware scheduling and a quota layer above it, but it does not replace placement.

Which workload types work with Kueue out of the box?

BatchJob, Kubeflow training jobs, RayJob, RayCluster, JobSet, and plain Pods or Pod Groups are all supported, with documentation pages for running Deployments and StatefulSets under Kueue as well. The plain Pod path is the escape hatch for workload shapes that have no dedicated integration.

What happens if Kueue holds back a job waiting for nodes?

The autoscaling integration answers this. Kueue connects to cluster-autoscaler's provisioningRequest through admissionChecks, so a job that needs more capacity than the cluster has can trigger node provisioning and then be admitted once the nodes arrive, instead of sitting pending indefinitely.

Which Kueue version should a new install pick?

Two lines were patched on 2026-09-17: v0.19.5 and, twenty minutes earlier, v0.18.9, and v0.20.0-rc.0 shipped on 2026-09-10. The choice depends on which minor line you already run, and the upgrade notes instruct you to read the `.0` notes for every minor boundary you cross. The v0.20.0-rc.0 notes also require renaming any device-class mapping or resource transformation that uses the reserved name `pods`, or the controller will not start.

Official sources

  1. kubernetes-sigs/kueue on GitHub
  2. License: Apache-2.0
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/kubernetes-sigs-kueue.svg)](https://hysenlabs.com/projects/kubernetes-sigs-kueue)