Self-hosted service
volcano-sh/volcano avatar
volcano-sh/volcano

Volcano: a Kubernetes batch scheduler for AI, big data and HPC jobs

A Cloud Native Batch System (Project under CNCF)

5,956 stars1,539 forksGoApache-2.0

At a glance

What is it?
Volcano replaces the default kube-scheduler with a batch-aware scheduler built on the kube-batch codebase, adding gang scheduling, queues and job lifecycle plugins. It is a CNCF incubating project written in Go under Apache-2.0, and it is the wrong tool for plain long-running web services.
Who is it for?
Adopt Volcano when your workloads are multi-pod jobs that must start together, such as distributed training, Spark or MPI runs, and you need queue-level fairness on top of Kubernetes. Do not adopt it for a cluster of ordinary stateless services; the default kube-scheduler already covers that, and the extra scheduler and controllers add operational surface for no gain.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 14 days ago.
What is it written in?
Mainly Go, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 17, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The gap Volcano fills in a stock Kubernetes cluster

The default kube-scheduler places one pod at a time. That is fine for a web service, where each replica is independent. It is a problem for a distributed training job: if the scheduler admits 7 of 8 workers and the eighth cannot fit, the seven running pods hold GPUs while doing nothing. Volcano calls this the gang problem and makes it the centre of its design. The README describes the project as a Kubernetes-native batch scheduling system that extends and enhances the standard kube-scheduler, aimed at AI, ML/DL, bioinformatics and genomics, and other big data applications. The audience is therefore platform teams running multi-pod, resource-hungry jobs on shared clusters, not teams deploying request-response services. The README also lists the frameworks the project has integrations for, including Spark, Flink, Ray, TensorFlow, PyTorch, Argo, MindSpore, PaddlePaddle, Kubeflow, MPI, Horovod, MXNet and KubeGene. That list is the practical boundary of who this is for: if your framework is not on it, you are writing your own integration.

How Volcano schedules: queues, PodGroups and plugins

Volcano does not patch kube-scheduler. It runs its own scheduler and controllers alongside the cluster. The README notes that the scheduler is built on kube-batch, citing issues #241 and #288 for the history. A batch job is expressed as a PodGroup, which groups the pods that must be scheduled as a unit and carries the minimum member count. The scheduler collects pending pods into those groups, then runs them through an ordered chain of plugins that decide whether the group can be placed, which nodes it should use, and in what order. Queues sit above PodGroups and provide the sharing and priority layer: a queue has a weight and a capability, and the scheduler allocates resources between queues rather than treating every pod as an equal citizen. This is the mechanism that makes fair sharing between teams possible on one cluster. The repository layout reflects the split: pkg/ holds the scheduler and controllers, cmd/ holds the binaries, config/ and installer/ hold the deployment material, and defs/ holds the CRD definitions. The go.mod file shows the project tracking Kubernetes closely, vendoring k8s.io/kubernetes v1.36.1 and controller-runtime v0.24.1, so the scheduler is coupled to a specific Kubernetes minor line.

Installing Volcano and submitting a first job

The README points to https://volcano.sh/ as the project home and the repository carries an installer/ directory plus a config/ directory for deployment. The docs directory holds the user guides, and the example directory holds runnable manifests. The README does not spell out the full install command, so treat the installer directory and the website as the authoritative source rather than copying a command from a blog post. What the repository does give you is a set of example jobs to submit once the scheduler is running. The example/job.yaml file is the smallest starting point, and the README's ecosystem links point at framework-specific guides such as docs/user-guide/how_to_use_pytorch_plugin.md. The repository's example directory also contains example/jobflow/ for job dependencies, example/cronjob/ for scheduled runs, example/hierarchical-jobs/, example/sharding/ and example/task-start-dependency/, each of which is a manifest you can apply directly. The Makefile shows how the project builds its own images: it defines IMAGE_PREFIX=volcanosh and OUTPUT_DIR ?= _output, so a source build writes its binaries under _output/bin. The same Makefile exposes CRD_OPTIONS and CRD_VERSION ?= v1 for CRD generation, which is the part you touch if you build the project from source rather than installing a release.

Where Volcano is the wrong tool

Volcano is a second scheduler, not a drop-in replacement for every workload. If your cluster runs stateless services with independent replicas, gang scheduling buys you nothing, and you have added a scheduler, a set of controllers and their CRDs to the control plane. The upgrade path is the sharper constraint. Because Volcano vendors k8s.io/kubernetes and client-go at a specific version (v1.36.1 in the current go.mod), a Volcano release is tied to a Kubernetes minor line, and upgrading Kubernetes ahead of Volcano is a coordination problem rather than a routine patch. The README does not document a rollback procedure for the scheduler, so plan for the case where you need to move pods back to kube-scheduler and confirm that path yourself before production. There is also an integration cost that the ecosystem list hides: a framework appears on the list because someone wrote a plugin or a guide, and the depth of that integration varies. The repository's example directory is the honest signal here. If your framework has no example and no guide under docs/, you are on your own.

Volcano against the Kubernetes scheduler and against Kueue

The obvious alternative is doing nothing: keep kube-scheduler and accept partial placement. That is a real option for jobs that tolerate stragglers, and it costs nothing to operate. The other alternative worth naming is Kueue, which also addresses batch admission on Kubernetes but from a different direction. Kueue is a job-queueing controller layer that works with the existing kube-scheduler rather than replacing it; Volcano ships its own scheduler with its own plugin chain and its own PodGroup abstraction. In practice the difference shows up in two places. First, scheduling decisions: Volcano's plugins decide node placement and ordering, so features like topology-aware placement live inside Volcano, while a Kueue-style setup leaves placement to kube-scheduler and its own plugins. Second, the unit of admission: Volcano's PodGroup is a Volcano CRD, so a job framework must be taught to create it, which is why the ecosystem integrations exist. Neither approach is strictly better. If you want to keep the stock scheduler and add queueing, the Kueue direction is less invasive. If you want scheduling policy itself under your control, Volcano is the more direct route.

Maintenance cost, release cadence and licence

The repository is not archived, and the last push was on 2026-09-10, so the project is being worked on. Recent releases include v1.15.2 and v1.14.5, both published on 2026-08-29, and v1.13.4 on 2026-08-31. That pattern matters more than any single version number: the project maintains more than one release line at a time, so a team on an older minor still receives patches. The cost of running Volcano is not the install, it is the coupling. Every Kubernetes upgrade has to be checked against the Kubernetes libraries the release was built with, and the go.mod file is where you find that number. The project is Apache-2.0, which permits commercial use and modification, and the repository carries a LICENSE file plus a licenses/ directory for third-party notices. That is a permissive arrangement, but it is not legal advice, and the FOSSA and OpenSSF Scorecard badges in the README point at tooling you can consult for dependency and supply-chain detail rather than relying on the badge itself.

Editorial conclusion

Adopt Volcano when your workloads are multi-pod jobs that must start together, such as distributed training, Spark or MPI runs, and you need queue-level fairness on top of Kubernetes. Do not adopt it for a cluster of ordinary stateless services; the default kube-scheduler already covers that, and the extra scheduler and controllers add operational surface for no gain. Before committing, verify three things in your own cluster: that your job framework has a documented Volcano integration (the README lists Spark, Flink, KubeRay, PyTorch and TensorFlow among others), that your Kubernetes version matches the k8s.io/kubernetes v1.36.1 vendored in go.mod, and that you can install the chart or manifests and confirm the volcano-scheduler pod becomes ready.

Frequently asked questions

How do I use Volcano with Kubernetes?

Volcano runs as its own scheduler and controllers alongside the cluster, and jobs are submitted as manifests that create a PodGroup. The repository's example directory contains ready manifests such as example/job.yaml and example/jobflow/, and the docs directory holds framework-specific guides. The README points to https://volcano.sh/ as the project home for installation material.

What is Volcano in the Kubernetes ecosystem?

It is a Kubernetes-native batch scheduling system that extends the standard kube-scheduler, and the README describes it as targeting AI, ML/DL, bioinformatics and other big data workloads. It is an incubating project of the Cloud Native Computing Foundation and its scheduler is built on kube-batch.

What are the alternatives to Volcano for batch scheduling on Kubernetes?

Keeping the stock kube-scheduler is one option if your jobs tolerate partial placement. Kueue is another: it adds queueing on top of the existing kube-scheduler, while Volcano ships its own scheduler with its own plugin chain and PodGroup abstraction. The README also lists integrations with Spark, Flink, KubeRay, PyTorch and TensorFlow.

Which version of Kubernetes does Volcano require?

The current go.mod vendors k8s.io/kubernetes v1.36.1 and client-go v0.36.1, so a Volcano release is built against a specific Kubernetes minor line. Check the go.mod of the release you plan to run before upgrading your cluster.

Official sources

  1. License: Apache-2.0
  2. Project website
  3. README
  4. Releases
  5. volcano-sh/volcano on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/volcano-sh-volcano.svg)](https://hysenlabs.com/projects/volcano-sh-volcano)