GPUStack: a GPU cluster manager for vLLM and SGLang serving, plus on-demand SSH GPU instances
Project brief: A GPU cluster manager for high-performance AI model serving (vLLM, SGLang) and on-demand SSH-accessible GPU instances.
At a glance
- What is it?
- GPUStack is an Apache-2.0 Python control plane that turns a set of GPU machines into a serving cluster. It is aimed at teams who want vLLM or SGLang without hand-writing the launch flags, and it is a poor fit if you only ever run one model on one box.
- Who is it for?
- Adopt GPUStack if you have several GPU machines, a mix of NVIDIA, AMD, Ascend or other accelerators, and no appetite for maintaining per-model vLLM launch scripts. Skip it if a single GPU serving a single model is your whole workload, because the server, worker agents and database add moving parts you will never use.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository received new commits within the last day.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The problem GPUStack solves: engine flags, not GPUs
Running vLLM or SGLang on one machine is a command line. Running it across a heterogeneous fleet is a configuration problem that never stops: which GPU gets which model, how many GPUs a tensor-parallel deployment needs, which engine version pairs with which driver, and what happens when a worker dies at 3am. GPUStack positions itself as the layer that answers those questions. The README describes it as an open-source GPU cluster manager for AI model serving and GPU instance provisioning, and the architecture section says a single GPUStack server can manage multiple GPU clusters across on-premises and cloud environments, with a scheduler that allocates GPUs and selects inference engines.
The intended audience is stated fairly plainly: development teams, IT organizations and service providers delivering Model-as-a-Service at scale. That is a different buyer from the person who wants a chat UI on a laptop. GPUStack also provisions SSH-accessible GPU instances on demand, which targets a second audience: teams that need a raw GPU box for fine-tuning or interactive work, not an inference endpoint. Both audiences share the same inventory of machines, which is the actual argument for one tool instead of two.
How the scheduler, workers and engines fit together
The architecture is a server plus workers. The server holds no GPU requirement of its own; the README states it can run on a CPU-only machine, and Docker Desktop on Windows and macOS is supported for it. Workers are the GPU side, and the README is explicit that only Linux is supported for worker nodes, with WSL2 suggested for Windows and macOS not supported at all.
Workers register with the server using a token issued from the UI. The registration command carries --server-url, --token and --advertise-address, and mounts the Docker socket so the worker can start containers on its own host. That is the data path for provisioning: the server decides, the worker executes locally.
Above that sits the engine layer. GPUStack does not implement inference itself. The README describes a pluggable engine architecture that configures vLLM, SGLang, TensorRT-LLM or your own engine. This is the design decision worth noting: the project's value is in selection and parameterization, so its quality depends on how well it tracks upstream engines. The README claims this enables deploying new models on the day they are released, and that pre-tuned modes exist for low latency or high throughput. It also lists extended KV cache systems (LMCache, HiCache) and speculative decoding (EAGLE3, MTP, N-grams). Those are upstream engine capabilities that GPUStack exposes through configuration rather than reinventing.
Operations tooling is bundled rather than bolted on: the README names automated failure recovery, load balancing, monitoring, authentication and access control, and mentions Grafana and Prometheus dashboards for system health and metrics.
Installing GPUStack with Docker and deploying a first model
The README's quick start assumes one node with at least one NVIDIA GPU, plus the NVIDIA driver, Docker and the NVIDIA Container Toolkit installed. The server itself is started with a single container. Note the published port: the command maps 80:80, so the UI answers on plain HTTP at the host's address.
sudo docker run -d --name gpustack \
--restart unless-stopped \
-p 80:80 \
--volume gpustack-data:/var/lib/gpustack \
gpustack/gpustackIf Docker Hub is slow or unreachable, the README offers a Quay mirror. The extra flag tells GPUStack to pull its own images from that registry too, which matters on networks that cannot reach Docker Hub.
sudo docker run -d --name gpustack \
--restart unless-stopped \
-p 80:80 \
--volume gpustack-data:/var/lib/gpustack \
quay.io/gpustack/gpustack \
--system-default-container-registry quay.ioStartup is asynchronous. The README directs you to follow the logs and then read the generated admin password out of the data volume, which is the credential you log in with alongside the username admin.
sudo docker logs -f gpustack
sudo docker exec gpustack cat /var/lib/gpustack/initial_admin_passwordAdding capacity happens in the UI: Clusters, then Add Cluster, then Docker as the provider. GPUStack generates a worker command for you. The README shows the shape of it, including --privileged, host networking and the Docker socket mount. Treat the generated token and address as secrets and placeholders to replace, not values to copy literally.
sudo docker run -d --name gpustack-worker \
--restart=unless-stopped \
--privileged \
--network=host \
--volume /var/run/docker.sock:/var/run/docker.sock \
--volume gpustack-data:/var/lib/gpustack \
--runtime nvidia \
gpustack/gpustack \
--server-url http://your_gpustack_server_url \
--token your_worker_token \
--advertise-address 192.168.1.2The worker then appears on the Workers page. From there the README's model flow is: open the Catalog page, pick a model (its example is Qwen3.5-0.8B), wait for the deployment compatibility checks, and save. Compatibility checking before deployment is the part that saves time, because it is where a mismatch between model size, GPU memory and engine support surfaces as a warning rather than a crash loop.
Where GPUStack gets in the way
The worker requirement is the sharpest limitation. GPUStack worker nodes are Linux only. If your GPU machines run Windows, the README points at WSL2 and warns against Docker Desktop for that role. If they run macOS, there is no worker path at all. A cluster manager that cannot manage half your fleet is a partial adoption, and partial adoption usually means two systems.
The worker container runs --privileged with the Docker socket mounted and host networking. That is a wide grant of authority to a process that also accepts instructions from a central server. The README does not describe a reduced-privilege mode, so this is a design boundary to accept or reject, not a flag to tune.
The README also does not document rollback. There is no described procedure for reverting a model deployment, downgrading the server, or recovering a worker that registered with a stale token. Alembic appears in the dependency list and alembic.ini is at the repository root, so schema migrations exist, which makes version skew between server and workers a real question the documentation leaves open. For a component that sits in front of production inference traffic, that gap matters more than any missing feature.
Finally, GPUStack is the wrong tool when the answer is one model on one GPU. A single vLLM process with a systemd unit has fewer failure modes, no database, no scheduler and no privileged agent. GPUStack earns its complexity at the point where you have more than one machine to fill.
GPUStack against vLLM and Ollama
The comparison people search for most is GPUStack versus vLLM, and the honest answer is that they are not substitutes. vLLM is the inference engine. GPUStack launches and configures vLLM, and the README lists it first among the engines it manages. Choosing GPUStack does not remove vLLM from your stack; it removes the shell scripts that start vLLM. If you already have a working vLLM deployment pipeline and it fits on one node, GPUStack adds a control plane you do not need. Where the difference bites is multi-node: tensor parallelism across machines, engine selection per model, and a scheduler deciding placement are things vLLM alone leaves to you.
Against Ollama the split is about workload shape. Ollama is built for local, single-user inference with a simple model pull and run flow. GPUStack targets serving at scale with authentication, access control, token metering and API request rate accounting, and it speaks industry-standard APIs for LLM, voice, image and video models. Those two designs serve different jobs. If your requirement is a developer running a model on a workstation, Ollama is the shorter path. If your requirement is many users hitting shared GPUs with quotas and auditability, Ollama is not the shape of the problem.
The Kubernetes question is separate again. The repository ships a charts/ directory and the dependency list includes both kubernetes and kubernetes-asyncio client libraries, so Kubernetes is a first-class environment for GPUStack rather than an afterthought. The README's multi-cluster claim includes Kubernetes clusters alongside on-premises servers and cloud providers. That means GPUStack is not an alternative to Kubernetes; it can be the layer that decides what runs on top of it.
Licence, versioning and what maintenance costs you
GPUStack is Apache-2.0, which permits commercial use, modification and redistribution provided the licence and notices are preserved. Nothing in the repository suggests a dual-licence or open-core split, but the README does not address licensing at all, so a legal review of your specific distribution model is your call, not something this article can settle.
The maintenance signal is straightforward. The last push to main was on 2026-07-31, and the most recent release listed is v2.2.3 on the same date, preceded by v2.2.3rc1 and v2.2.2 on 2026-07-24. That is a project shipping on a roughly weekly cadence, with release candidates published alongside finals. It is not archived.
The upgrade cost is where the design shows its seams. GPUStack wraps upstream engines, so an engine upgrade and a GPUStack upgrade are separate events that have to stay compatible, and a model that deploys today may fail compatibility checks after either changes. The dependency list is long and pinned with upper bounds in places (fastapi below 0.137.0, transformers excluding 4.57.0, kubernetes capped below 34.0.0), which suggests the maintainers hit breakage and responded by constraining versions. The pyproject.toml declares requires-python >=3.10,<3.13, so a Python 3.13 environment is out of range for running from source. There is also a pinned gpustack-runner==0.1.28 dependency, meaning the runner is versioned separately from the server and can drift. Plan upgrades as a coordinated server-and-worker operation, not a rolling restart of one component.
Editorial conclusion
Adopt GPUStack if you have several GPU machines, a mix of NVIDIA, AMD, Ascend or other accelerators, and no appetite for maintaining per-model vLLM launch scripts. Skip it if a single GPU serving a single model is your whole workload, because the server, worker agents and database add moving parts you will never use. Before committing, verify three things on your own hardware: that your accelerator appears in the supported list, that the worker node runs Linux (Windows via WSL2 only, macOS unsupported), and that the container registry and model source reachable from your network are the ones the install command points at.
Frequently asked questions
What is GPUStack?
It is an open-source GPU cluster manager for AI model serving and GPU instance provisioning. It configures and orchestrates inference engines such as vLLM, SGLang and TensorRT-LLM, and can launch SSH-accessible GPU instances on demand.
How does GPUStack compare with vLLM?
They operate at different levels. vLLM is an inference engine, and GPUStack is the manager that configures and orchestrates it. The README lists vLLM first among the pluggable engines GPUStack supports, so adopting GPUStack does not replace vLLM.
How does GPUStack compare with Ollama?
The README positions GPUStack for Model-as-a-Service at scale, with built-in user authentication, access control, real-time GPU monitoring and token usage metering. Ollama is not mentioned anywhere in the repository material, so a direct feature comparison is not something the documentation supports.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/gpustack-gpustack)