Self-hosted service
dstackai/dstack avatar
dstackai/dstack

dstack: a control plane for GPU workloads across clouds, Kubernetes and bare metal

Vendor-agnostic orchestration for training, inference and agentic workloads across NVIDIA, AMD, TPU, and Tenstorrent on clouds, Kubernetes, and bare metal.

2,252 stars262 forksPythonMPL-2.0

At a glance

What is it?
dstack provisions accelerators and runs dev environments, jobs and inference services from YAML definitions. It is a strong fit for teams spread over several GPU providers, and a poor fit for anyone who wants a fully managed platform.
Who is it for?
Adopt dstack if you already run workloads on more than one GPU provider, or you expect to, and you are comfortable operating a server and writing YAML. Do not adopt it if you want a hosted platform with no server to run, or if you only ever use one cloud's native scheduler.
Can I use it commercially?
Yes, with conditions. MPL-2.0 is a weak copyleft licence: you can use it inside commercial and closed-source software, but if you distribute changes to its own files, you must publish those changes under the same licence.
Is it still maintained?
Yes. The repository last received commits 12 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 17, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The problem dstack targets: one workload definition, many GPU sources

The README describes dstack as "a unified control plane for GPU provisioning and orchestration that works with any GPU cloud, Kubernetes, or on-prem clusters." That sentence is the whole pitch, and it names the actual pain. A team that trains on one provider, runs inference on another, and keeps a bare-metal box for experiments ends up maintaining three sets of provisioning scripts, three notions of a job, and three ways to expose a port. dstack replaces those with one server and one YAML format.

The intended audience is narrower than the description suggests. You need workloads that want GPUs, you need more than one place to run them, and you need to be willing to run a server yourself. A single-developer project on one cloud gains little from an abstraction layer that has to be installed and configured. The README's own note that backend configuration is unnecessary for on-prem SSH fleets hints at the second audience: labs with physical machines and no cloud account at all.

Server, backends and YAML: how the pieces fit together

The architecture has two processes and a set of configuration files. A dstack server holds state and talks to your configured backends. A CLI, or the programmatic API, or an AI agent using the published skills, submits work to that server. The server handles provisioning, job queuing, auto-scaling, networking, volumes, run failures, out-of-capacity errors and port-forwarding, according to the README.

What you submit is a configuration file. The README lists six kinds: fleets, which provision and manage clusters; dev environments, which launch environments accessible from an IDE or an agent; tasks, for training and batch jobs across a single node or a cluster; services, which deploy model inference as endpoints; presets, described as agent-driven inference optimization and marked experimental; and volumes, for instance and network volumes that persist data. These live as YAML files inside your repository, which is the design decision that matters most. Workload definitions are versioned alongside the code they run, so a change to the cluster shape shows up in a pull request.

Accelerator support is stated as NVIDIA, AMD, Google TPU and Tenstorrent out of the box. The recent release notes add more texture: a Slurm backend in 0.20.27, multiple Kubernetes clusters in 0.20.21, NVIDIA Dynamo integration in 0.20.20, Kubernetes volumes in 0.20.17, PD disaggregation in 0.20.10 and replica groups in 0.20.7. That cadence suggests the project is being extended along the axes its users ask for, though the README does not say which of these are stable.

Installing dstack and submitting a first configuration

The README installs the server with uv, and the server requires Git and OpenSSH. It runs on Linux, macOS and Windows via WSL 2. Starting it prints an admin token and a local address.

bash
uv tool install "dstack[all]" -U
dstack server

The documented output shows the admin token and the line "The server is running at http://127.0.0.1:3000/". Keep that token; the CLI needs it. If you install the CLI separately rather than with the server, the README gives this second command and then a project registration step that writes to ~/.dstack/config.yml.

bash
uv tool install dstack -U
dstack project add \
    --name main \
    --url http://127.0.0.1:3000 \
    --token bbae0f28-d3dd-4820-bf61-8f4bb40815da

The README states that configuration is updated at ~/.dstack/config.yml after this runs. The token in the example is the one the server printed on startup, not a fixed value.

With the server reachable, you write a configuration file and apply it. The README does not reproduce a full YAML example inline, so the field names for a given workload type have to come from the concepts pages for fleets, tasks, services and the rest. What the README does say is that you apply configurations with the dstack apply CLI command, through the programmatic API, or via the agent skills, which are installed with a single npx command.

bash
npx skills add dstackai/dstack

That installs skills that let Claude, Codex and Cursor create and manage fleets and submit workloads on your behalf, per the README.

Where dstack gets in the way

The most concrete limitation is the one the project states about itself. pyproject.toml carries the classifier "Development Status :: 4 - Beta". Whatever the release cadence looks like, the maintainers are not claiming production maturity, and you should read the version numbers in that light. The 0.x series with frequent point releases means configuration schemas can move.

Operating a server is a real cost, not a footnote. The README points to a server deployment guide for configuration options but does not document rollback, backup or recovery of server state. If your orchestration layer is a single server you run yourself, its availability becomes your problem in a way that a cloud provider's native scheduler never is.

There is also a category error to avoid. dstack orchestrates workloads; it does not train models, serve them well by default, or replace your experiment tracker. The presets feature is explicitly labelled experimental, so treating agent-driven inference optimization as a production capability would be reading ahead of the documentation. And if your entire estate sits on one Kubernetes cluster that your platform team already manages well, adding a control plane above it buys abstraction you may not want to pay for.

dstack and SkyPilot: two answers to the same question

The comparison people search for is dstack versus SkyPilot, and the difference is mostly about where the control plane lives. SkyPilot's model is a CLI that submits jobs to whichever cloud or Kubernetes cluster you point it at, with the cluster lifecycle managed from the client side. dstack's model is a long-running server that holds the state, with the CLI as a client of that server. That server is what makes fleets, services with endpoints, volumes and multi-cluster Kubernetes support coherent, and it is also what you have to operate.

The practical consequence: if you want a tool you install and use without any persistent component, dstack's architecture is heavier than what you are looking for. If you want shared state across a team, with dev environments, long-running inference services and provisioned clusters all visible in one place, the server is the feature rather than the tax. The README's mention of a programmatic API and agent skills depends on that server existing.

A second alternative is simply the native scheduler of whichever cloud you already use. It will be better integrated with that provider's networking, storage and identity, and it will not help you at all the moment you add a second provider or an on-prem box.

Licence, maintenance and the cost of upgrades

dstack is licensed under the Mozilla Public License 2.0, and pyproject.toml classifies it as "License :: OSI Approved :: Mozilla Public License 2.0 (MPL 2.0)". MPL-2.0 is file-level copyleft: modifications to MPL-covered files must be made available under the same licence, while larger works that combine dstack with other code can be licensed differently. Running the server and applying your own YAML configurations does not trigger distribution obligations. This is a description of the licence text, not legal advice; if you plan to ship modified dstack files, have counsel read LICENSE.md.

The repository is not archived and the last push was on 2026-09-10, so the project is being worked on. Release 0.21.5 landed on 2026-09-03, a week before that push. The upgrade cost follows from the 0.x versioning: a control plane that provisions real GPU capacity is not something you upgrade casually, because a schema change in your YAML can fail at apply time rather than at build time. Pinning the version in your install command and reading the release notes between pins is the cheap insurance here. The README does not document a migration process for configuration files across versions.

Editorial conclusion

Adopt dstack if you already run workloads on more than one GPU provider, or you expect to, and you are comfortable operating a server and writing YAML. Do not adopt it if you want a hosted platform with no server to run, or if you only ever use one cloud's native scheduler. Before committing, verify that your specific accelerators and backends are covered in the backends documentation, check that your Python version is at least 3.10 as pyproject.toml requires, and confirm the MPL-2.0 file-level copyleft terms against how you plan to distribute modifications.

Frequently asked questions

What is dstack?

It is an open-source control plane for GPU provisioning and orchestration that works with GPU clouds, Kubernetes and on-prem clusters. You install a server, configure backends, and apply YAML definitions for fleets, dev environments, tasks, services, presets and volumes.

What is a dstack alternative?

SkyPilot is the closest alternative, and the difference is architectural: SkyPilot submits jobs from a CLI, while dstack keeps state in a server you run. The native scheduler of a single cloud is also an alternative if you never add a second provider.

How do I install dstack?

The README installs the server with uv tool install "dstack[all]" -U and then runs dstack server. The CLI can be installed separately with uv tool install dstack -U and pointed at the server with dstack project add.

Which accelerators does dstack support?

The README states that NVIDIA, AMD, Google TPU and Tenstorrent accelerators are supported out of the box. Recent releases also added a Slurm backend and multiple Kubernetes clusters.

What licence does dstack use?

Mozilla Public License 2.0, which pyproject.toml lists under the MPL 2.0 classifier. It is file-level copyleft, so modifications to MPL-covered files must stay under the same licence.

Official sources

  1. dstackai/dstack on GitHub
  2. License: MPL-2.0
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/dstackai-dstack.svg)](https://hysenlabs.com/projects/dstackai-dstack)