kubernetes-sigs/descheduler: evicting pods after the scheduler has moved on
Descheduler for Kubernetes
At a glance
- What is it?
- The descheduler does not schedule anything. It reads a policy, finds pods sitting on nodes the cluster has outgrown, and evicts them so kube-scheduler can place them again. Here is what the policy keys actually control, how to install it with Helm or the Kubernetes manifests, and where it stops being the right tool.
- Who is it for?
- Adopt the descheduler if you run a long-lived cluster where node utilization, taints and affinity requirements drift after the initial placement, and you accept that eviction is the only lever it has. Skip it if you need placement guarantees at creation time, or if you are on a platform that already bundles its own descheduler build.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 4 days ago.
- What is it written in?
- Mainly Go, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The gap between a scheduling decision and the cluster it was made in
kube-scheduler binds pending pods to nodes using predicates and priorities evaluated against the cluster state at the moment a pod appears. The README frames the whole project as a consequence of that timing: clusters are dynamic, and the decision that was correct when the pod was created may not hold later. Nodes become under or over utilized, taints and labels are added or removed, pod and node affinity requirements stop being satisfied, nodes fail and their pods land elsewhere, and new nodes join.
The descheduler is the component that acts on that drift. Its scope is deliberately narrow. According to the README, it finds pods that can be moved and evicts them, and it does not schedule replacements: the default scheduler does that. That single sentence defines the tool's cost model. Every action the descheduler takes is a pod deletion, and the recovery depends entirely on kube-scheduler making a better choice the second time.
The audience is whoever owns cluster efficiency rather than a single workload. Platform teams running multi-tenant clusters, operators who see a handful of nodes pinned at high utilization while others idle, and anyone who has changed node labels or taints and watched existing pods ignore the change. Application developers rarely need it, because the problem it addresses is a property of the cluster, not of one deployment.
Strategy plugins, the default evictor, and where the policy keys apply
The policy is the interface. The README describes it as configurable, with default strategy plugins that can be enabled or disabled, plus a common eviction configuration at the top level and configuration from the Evictor plugin (Default Evictor when nothing else is named). Both the top-level configuration and the Evictor plugin configuration apply to all evictions, which is worth reading twice: a limit you set at the top level is not scoped to one strategy.
The top-level keys are what keep the tool from being reckless. maxNoOfPodsToEvictPerNode caps evictions from each node summed across all strategies. maxNoOfPodsToEvictPerNamespace does the same per namespace. maxNoOfPodsToEvictTotal caps evictions per rescheduling cycle, again summed across strategies. All three default to nil, which means no cap unless you write one. nodeSelector limits which nodes are processed, and the README is explicit that it is only used when nodeFit is true and only by the PreEvictionFilter extension point, so setting it alone does not restrict the run the way a reader might assume.
The metricsCollector block is marked deprecated in the README table, with a metricsCollector.enabled boolean that defaults to false and turns on Kubernetes Metrics Server collection for actual resource utilization. Because it is deprecated, treat it as a migration surface rather than a feature to build on. The repository ships example policies under examples/, including policy.yaml, high-node-utilization.yml, low-node-utilization.yml, node-affinity.yml, topology-spread-constraint.yaml, pod-life-time.yml, pod-life-time-transition.yml, too-many-restarts.yml and failed-pods.yaml. Those filenames map to the strategies far more clearly than any prose summary, and they are the fastest way to see what a working policy looks like.
Installing the descheduler as a Job, CronJob or Deployment
The README's quick start states that the descheduler runs as a Job, a CronJob or a Deployment inside the cluster, and that it has the advantage of running repeatedly without user intervention. The pod runs as a critical pod in the kube-system namespace so it is not evicted by itself or by the kubelet.
Every variant starts with the same two base objects, RBAC and the policy ConfigMap, then one workload manifest. For a one-off Job:
kubectl create -f kubernetes/base/rbac.yaml
kubectl create -f kubernetes/base/configmap.yaml
kubectl create -f kubernetes/job/job.yamlSwap the last line for kubernetes/cronjob/cronjob.yaml to run on a schedule, or kubernetes/deployment/deployment.yaml to keep it running. The ConfigMap is where your policy lives, so edit it before creating the workload if you want caps in place from the first cycle.
Helm is the shorter path and is what most readers searching for the descheduler helm chart want. The README notes that an official chart has existed since release v0.18.0, points to the chart README in charts/descheduler/README.md for detailed instructions, and lists the chart on artifact hub. Kustomize is also supported; the README gives this example for the Job form, pinned to a release branch:
kustomize build 'github.com/kubernetes-sigs/descheduler/kubernetes/job?ref=release-1.34' | kubectl apply -f -The same pattern is documented for kubernetes/cronjob and kubernetes/deployment. Note the ref: the README warns that master is in development and its information may not apply to previous versions, and it publishes a table mapping releases such as v0.36.x to release-1.36 and v0.35.x to release-1.35. Pin to the branch matching the image you actually run.
Eviction is the only verb, and that is the limitation
The descheduler cannot move a pod. It can only ask for it to be deleted, and from that point the outcome is out of its hands. If the replacement pod cannot be scheduled, or the scheduler picks the same node again, the eviction bought nothing except a restart. The README states plainly that the current implementation does not schedule replacement of evicted pods and relies on the default scheduler for that, so any planning conversation about the descheduler is really a conversation about whether kube-scheduler will make a different decision once the pod is gone.
That makes the tool a poor fit for workloads with expensive cold starts or long initialization, since the cost of a wrong eviction is paid by the application rather than the cluster. It is also the wrong tool for enforcing placement at creation time: if a pod must never land on a particular node class, that belongs in the scheduler's constraints, not in a periodic cleanup that runs after the fact. And because the three eviction caps default to nil, an unconfigured policy has no built-in brake. The README documents the keys but does not document rollback, so there is no described procedure for undoing a rescheduling cycle that went badly; the practical recovery is the workload's own controller recreating pods, which is not the same thing.
One more boundary is versioning. The README's documentation table exists because behaviour differs across releases, and it directs users of a published image such as registry.k8s.io/descheduler/descheduler:v0.36.0 to that version's release branch. Reading master while running an older image is a documented way to get confused.
How it differs from Karpenter and from scheduler-side approaches
The comparison readers search for most often is descheduler versus Karpenter, and the split is clean. Karpenter is a provisioner: it decides what nodes should exist and creates or removes them in response to pending pods and workload requirements. The descheduler never touches the node set. It assumes the nodes are there and asks whether the pods are on the right ones. A cluster can use both, and the two act on opposite ends of the same problem: Karpenter changes supply, the descheduler redistributes demand across what already exists.
The other real alternative is not a different tool but a different layer. Kubernetes scheduling constraints, affinity, anti-affinity, topology spread constraints and taints, express placement rules that kube-scheduler enforces when a pod is created. They are declarative and they do not delete anything. What they cannot do is retroactively fix pods that were placed before the rule existed, which is exactly the case the descheduler was written for. The repository's own examples/topology-spread-constraint.yaml is a useful illustration of the overlap: a topology spread constraint shapes where new pods go, while the strategy that reads it rebalances pods that are already running.
There are also platform builds. The related searches include descheduler openshift and descheduler operator, which reflects that OpenShift ships its own descheduler with an operator-driven configuration model rather than the ConfigMap and manifest workflow described in this README. If you are on OpenShift, check what the platform already installs before adding the upstream manifests on top of it.
Maintenance, release cadence and the licence
The repository is not archived, and its last push was on 2026-09-07. Recent releases are v0.36.0 and descheduler-helm-chart-0.36.0, both dated 2026-05-20, with descheduler-helm-chart-0.35.1 on 2026-03-09. The chart and the binary version in lockstep, so the chart release is the practical upgrade signal for Helm users.
The upgrade cost is the documentation table. The README maintains per-release branches from release-1.30 through release-1.36 and states that master is in development and may not work for previous versions. That means an upgrade is not just a new image tag: you re-read the README on the target release branch and check whether your policy keys still behave as documented. The metricsCollector deprecation is the kind of change that shows up this way, and the README does not describe a migration path beyond the deprecation marker itself.
Building from source is a Go project. go.mod declares module sigs.k8s.io/descheduler and go 1.26.0, and the Dockerfile uses golang:1.26.0 as the build stage and a scratch final image running as USER 1000 with the binary at /bin/descheduler. The Makefile builds for amd64, arm and arm64 and defaults REGISTRY to gcr.io/k8s-staging-descheduler, which is a staging registry, not a production one. The licence is Apache-2.0, the same licence as Kubernetes itself, which is permissive and includes an explicit patent grant; the Dockerfile and Makefile both carry the Apache header. This is not legal advice, and if you redistribute a modified build you should read the licence text yourself.
Editorial conclusion
Adopt the descheduler if you run a long-lived cluster where node utilization, taints and affinity requirements drift after the initial placement, and you accept that eviction is the only lever it has. Skip it if you need placement guarantees at creation time, or if you are on a platform that already bundles its own descheduler build. Before enabling anything, read the README for your release branch rather than master, and set maxNoOfPodsToEvictPerNode, maxNoOfPodsToEvictPerNamespace and maxNoOfPodsToEvictTotal to small values so the first rescheduling cycle cannot empty a node.
Frequently asked questions
Does the descheduler schedule the replacement pods after eviction?
No. The README states that in the current implementation the descheduler does not schedule replacement of evicted pods and relies on the default scheduler for that. The descheduler only evicts; kube-scheduler decides where the pod goes next.
What is the difference between karpenter and descheduler?
They operate on opposite sides of the cluster. Karpenter decides which nodes should exist and provisions or removes them, while the descheduler leaves the node set alone and evicts already-running pods so kube-scheduler can place them again.
What are the alternatives to the kubernetes descheduler?
The alternative at the scheduling layer is expressing placement as scheduler constraints such as affinity, anti-affinity, topology spread constraints and taints, which are enforced at pod creation and never delete anything. Those cannot retroactively fix pods placed before the rule existed, which is the case the descheduler addresses.
How do I install the descheduler with Helm?
The README states that an official Helm chart has existed since release v0.18.0, points to charts/descheduler/README.md for detailed instructions, and notes the chart is listed on artifact hub. The chart version tracks the binary, for example descheduler-helm-chart-0.36.0 alongside v0.36.0.
Which policy keys limit how many pods the descheduler evicts?
maxNoOfPodsToEvictPerNode, maxNoOfPodsToEvictPerNamespace and maxNoOfPodsToEvictTotal, each summed across all strategies. All three default to nil in the README table, so no cap applies unless you set one.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/kubernetes-sigs-descheduler)