Open-source project
SchedMD/slurm avatar
SchedMD/slurm

Slurm Workload Manager: What It Does and When It Is the Wrong Tool

Slurm: A Highly Scalable Workload Manager

4,396 stars891 forksCNOASSERTION

At a glance

What is it?
Slurm is a Linux cluster resource manager and job scheduler from SchedMD. This covers its three core functions, the documented build path, and the cases where a container orchestrator fits better.
Who is it for?
Adopt Slurm if you run batch and parallel work on dedicated Linux compute nodes and need allocation, a queue, and a job launch framework in one system. Do not adopt it if your workloads are long-running services that need rolling updates and horizontal autoscaling, which is what Kubernetes is built around.
Can I use it commercially?
Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
Is it still maintained?
Yes. The repository received new commits within the last day.
What is it written in?
Mainly C, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The Problem Slurm Solves on a Shared Linux Cluster

Put more than a handful of users on the same set of Linux machines and the coordination problem appears immediately. Two people start memory-heavy jobs on the same node and one of them dies. Someone leaves a process running for a week and blocks a resource that three other people need. There is no record of who asked for what, and no way to queue work that cannot start yet. Slurm exists to answer those questions. The README states three key functions: allocating exclusive and/or non-exclusive access to compute nodes for a duration, providing a framework for starting, executing, and monitoring work on the allocated nodes, and arbitrating conflicting requests by managing a queue of pending work. Those three functions map to the commands users touch most. sbatch submits a batch script and returns immediately. salloc requests an interactive allocation. srun launches a job step inside an allocation. The target audience is not a single developer on a laptop. It is a research group, an HPC center, or a company with a fixed pool of compute nodes and more demand than capacity. Slurm has been tested only under Linux, per the README, so a mixed Windows cluster is out of scope.

How the Slurm Architecture Splits Control and Execution

The repository layout tells you a lot about the design. Under src/ the source is split into self-describing subdirectories, and src/slurmctld is the controller daemon. That is the central process that holds the queue, tracks node state, and makes scheduling decisions. It is a single logical point of coordination, which is why production deployments usually run a backup controller alongside the primary. Jobs do not execute on the controller. The controller hands an allocation to slurmd daemons running on the compute nodes, and those daemons start and supervise the work locally. The public interface for programs that want to talk to the scheduler is a C API, with headers such as slurm.h and slurm_errno.h installed from the slurm/ directory. That matters for integration: a site that wants its own portal or workflow engine compiles against the API rather than scraping command output. Configuration lives in etc/, which holds a sample config file and supporting scripts, and the manual pages for commands and configuration files ship under doc/. The testsuite/ directory contains tests written for Check, Expect and Pytest, so the project tests itself at more than one level. The trade-off in this design is explicit: a centralized controller gives you a global view of the cluster and a single queue, and it also means controller availability is the thing you have to engineer around.

Building Slurm from the Source Distribution

The README does not inline install steps. It points to https://slurm.schedmd.com/quickstart_admin.html and states that extensive documentation is available from the home page. The repository ships the build machinery itself: configure and configure.ac at the top level, Makefile.am, auxdir/ for autotools scripts, and an INSTALL file. Because the README gives no command sequence, the honest thing to do is name what is present rather than invent a transcript. The top-level directory listing includes configure, Makefile.am, Makefile.in, aclocal.m4, config.h.in, and auxdir/, which is the standard autotools set, and the README tells you to see INSTALL for the build. The INSTALL file and the quickstart admin page are therefore the two documents to read before running anything, because the required options depend on your site. Two other files matter at this stage. slurm.spec is an RPM spec file, and debian/ holds packaging metadata, so a site that prefers a package build has a starting point in the repository rather than in the README. The README also notes that contribs/ holds tools outside Slurm proper and that building them requires a separate make contrib/install-contrib step, which is the one build invocation the README does spell out. The README does not document rollback or uninstall, so plan the install prefix before you build.

Where Slurm Is the Wrong Choice

Slurm allocates nodes for a duration and runs work inside that allocation. It is not a service platform. If your workload is a web API that needs to stay up, scale out on request rate, and roll forward without dropping connections, Slurm's model fights you: the scheduler wants to know when a job ends, and a process that is supposed to run forever never gives that signal cleanly. The same mismatch appears with container-native pipelines that expect a reconciler to restart failed pods automatically. Slurm will requeue a job if you configure it to, but that is a policy decision you make, not a default behavior you inherit. There is also an operational cost that the README is honest about in a different way: the project calls itself fault-tolerant, and fault tolerance here means the system is designed to survive component failure, not that it removes the need to run and monitor a controller. A single-node workstation user should not install Slurm at all. The overhead of a controller, daemons, and a configuration file exceeds anything a laptop workload needs. Finally, the README says Slurm has been tested only under Linux, so treat any non-Linux deployment as unsupported territory.

Slurm vs Kubernetes: Different Units of Work

The comparison people search for is Slurm versus Kubernetes, and the difference is what each one schedules. Slurm schedules a job onto a set of nodes for a bounded time, then releases them. Kubernetes schedules a long-lived container and keeps it running, replacing it when it dies. A training run that occupies eight nodes for six hours is a Slurm job. A model-serving endpoint that must answer requests continuously is a Kubernetes deployment. The second difference is the queue. Slurm has a first-class notion of pending work with priorities and time limits, and users expect to wait. Kubernetes has no queue in that sense; a pod that cannot be scheduled is an error condition to be resolved, not a normal state to sit in. The third is the interface. Slurm's integration surface is a C API plus a command set, and its configuration is a flat file. Kubernetes exposes an HTTP API and declarative YAML objects. Neither is a superset of the other. Sites that need both often run them side by side, with Slurm owning the batch nodes and Kubernetes owning the services, rather than trying to force one system to do both jobs.

Release Lines, Licence, and the Cost of Staying Current

The repository shows three maintained release lines at the same time: v26.05.4, v25.11.8, and v25.05.9, all tagged on 2026-09-02. That pattern is the upgrade story in miniature. SchedMD maintains parallel branches, so a site can stay on an older line and still receive point releases, and the newest line gets features first. The cost is that you must choose a line deliberately and read CHANGELOG.md and the CHANGELOG/ directory before moving, because the changelog is where behavior changes are recorded. The last push to the repository was on 2026-09-23, so the master branch moves independently of the release tags. Building from master is a different commitment than building from a release tag, and the README's own instruction to read INSTALL and the quickstart admin page applies either way. On licensing, the README states the software is distributed under the GNU General Public License and points to COPYING, DISCLAIMER, and LICENSE.OpenSSL for details. The repository metadata does not resolve to a standard SPDX identifier, so read those files directly rather than relying on a label. The OpenSSL file exists because of a separate linking exception; that is a detail your legal review should see, not something to infer from a licence badge.

Working with the Slurm Community and Source Tree

The README is direct about where bugs go: the official issue tracker is at https://support.schedmd.com/, not the GitHub issue queue. Contributions and patches are welcome, and the README points to CONTRIBUTING.md for the process. The repository also carries CODEOWNERS, CODE_OF_CONDUCT.md, SECURITY.md, and a .pre-commit-config.yaml, so there is a defined path for a patch from formatting through review. For anyone extending Slurm rather than just using it, three directories matter. src/api is where the client-facing library code lives. contribs/ holds tools outside Slurm proper, and the README notes that building them requires a separate make contrib/install-contrib step. The slurm/ directory is the installed include set, which is what you compile against when writing a program that submits or queries jobs. If your team plans to build a portal on top of Slurm, budget time for the C API and for reading the manual pages under doc/, because that is where the command and configuration reference actually lives.

Editorial conclusion

Adopt Slurm if you run batch and parallel work on dedicated Linux compute nodes and need allocation, a queue, and a job launch framework in one system. Do not adopt it if your workloads are long-running services that need rolling updates and horizontal autoscaling, which is what Kubernetes is built around. Before committing, read the quickstart admin page at https://slurm.schedmd.com/quickstart_admin.html and the INSTALL file in the repository, then decide whether you will track one of the maintained release lines or the master branch, because that choice determines your upgrade cadence.

Frequently asked questions

What is Slurm used for?

Slurm is a cluster resource management and job scheduling system. The README describes three functions: allocating access to compute nodes for a duration, starting and monitoring work on those nodes, and arbitrating conflicting requests through a queue.

Is Slurm named after Futurama?

The README and repository files give no origin for the name, so this cannot be confirmed from the available material.

What is Slurm vs Kubernetes?

Slurm allocates nodes for a bounded duration and runs work inside that allocation, with a queue of pending jobs. Kubernetes keeps long-lived containers running and replaces them when they fail. The unit of scheduled work differs, not just the implementation.

Is Slurm owned by Nvidia?

The repository lists SchedMD as the owner and points to https://support.schedmd.com/ as the official issue tracker. No ownership by Nvidia appears in the README or repository metadata.

How do I install Slurm?

The README does not inline install steps; it directs readers to https://slurm.schedmd.com/quickstart_admin.html and to the INSTALL file in the repository. The distribution ships configure, Makefile.am, auxdir/, slurm.spec and debian/, so both source and package builds have a starting point in the tree.

How do I use salloc in Slurm?

The README does not document salloc usage. It states that Slurm provides a framework for starting, executing, and monitoring work on allocated nodes, and the manual pages under doc/ are where the command reference ships.

Official sources

  1. Issues
  2. Project website
  3. README
  4. Releases
  5. SchedMD/slurm on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/schedmd-slurm.svg)](https://hysenlabs.com/projects/schedmd-slurm)