SkyPilot: a unified control plane for GPUs across clouds and Kubernetes
The AI Compute Platform for frontier teams. SkyPilot turns fragmented AI compute into one AI supercomputer, so frontier AI teams build custom intelligence faster.
At a glance
- What is it?
- SkyPilot is an Apache-2.0 Python system that runs AI workloads on any cloud, Kubernetes cluster or Slurm cluster through one interface. It is strongest when you already own fragmented compute and need scheduling, failover and autostop on top of it.
- Who is it for?
- Adopt SkyPilot if you already hold GPUs across more than one cloud or cluster and want one job interface plus autostop and failover instead of a per-provider script. Do not adopt it if you run a single fixed cluster with no queueing or failover problem, or if you need a managed control plane that someone else operates.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 13 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 17, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The fragmentation problem SkyPilot targets
Most AI teams end up with compute in several places. A reserved block of GPUs in one cloud, a Kubernetes cluster someone else administers, a Slurm cluster inherited from an HPC group, plus spot capacity that is cheap until it disappears. Each of those has its own way to submit work, its own credential model, and its own failure behaviour. The team writes a different launch script per provider, and the scripts drift.
SkyPilot's pitch is that you describe the workload once and let the system place it. The README frames the split audience explicitly: AI teams get "a simple interface to run jobs on any infra", while infra teams get "a unified control plane to manage any AI compute". The project describes itself as BYOC, meaning everything is launched inside your own cloud accounts, VPCs and clusters rather than on SkyPilot-hosted machines.
That distinction matters more than the feature list. A hosted training platform takes your data and your credentials into its own tenancy. SkyPilot is a client you install and run; the compute stays where it already is. For teams with compliance constraints on where data may sit, the BYOC model is the whole reason to look at it.
How SkyPilot places a job across clouds and clusters
The unit of work is a task, described in YAML. A task names its resource requirements (accelerator type and count, whether spot instances are acceptable) and the commands to run. The README's summary is "Environment and job as code", which is the accurate description: the YAML is portable across the supported infrastructure list, and the same file can target Kubernetes, Slurm, AWS, GCP, Azure, OCI, CoreWeave, Nebius, Lambda Cloud, RunPod and others.
The scheduler is where the design shows. SkyPilot reads the candidate infrastructure for a task and picks a placement, with the README describing an "intelligent scheduler" that "automatically schedule[s] on the most available infra". When a chosen region or provider cannot deliver, the documented behaviour is flexible provisioning with smart failover: the job moves rather than failing outright. This is the mechanism that makes spot instances usable, because a preempted spot node is a placement failure the system can absorb.
The second mechanism is autostop, which the README lists under maximizing fleet utilization as "automatic cleanup of idle resources". A cluster that finishes its work and sits idle is the most common way GPU budgets leak. Autostop is a policy on the cluster, not something each job has to remember to do.
On Kubernetes specifically, SkyPilot positions itself as an AI-friendly layer over an existing cluster. The README claims gang scheduling, multi-node jobs, queueing, multi-cluster support, SSH into pods, code sync and IDE connection. Gang scheduling is the one that matters for distributed training: either all workers start or none do, which prevents a half-allocated job from holding GPUs while it waits.
Installing SkyPilot and launching a first job
The README gives the install command in the uv form, with pip, nightly and from-source listed as alternatives in the installation docs. The bracket selects which cloud providers you want wired in, so you only pull the dependencies you need.
uv pip install "skypilot[kubernetes,aws,gcp,azure,oci,nebius,lambda,runpod,fluidstack,paperspace,cudo,ibm,scp,seeweb,shadeform,verda]"After installing, the documented next step is the Quickstart, which the README describes as launching your first cluster in about two minutes. Before that, you need credentials for at least one provider in the normal place for that provider. SkyPilot does not replace your cloud CLI authentication; it reads it.
The interactive surface is the sky command line. The README does not print a check command in the excerpt above, so confirm the exact invocation against the installation docs rather than guessing a flag.
Task files live in the examples directory of the repository, which is the practical starting point. The examples tree covers distributed PyTorch, Ray training, cron-style jobs, custom images, disk sizing and a long list of provider-specific setups such as AWS EFA and CoreWeave InfiniBand. Copying one of those YAML files and editing the resources block is faster than writing a task from scratch, and it keeps you inside syntax the project actually tests.
For teams that want an agent to drive the tool, the README documents a separate path. Install the SkyPilot Skill and tell your agent to fetch and follow the INSTALL.md file in the agent directory of the repository:
Fetch and follow https://github.com/skypilot-org/skypilot/blob/HEAD/agent/INSTALL.md to install the skypilot skillThat is a distinct install route from the pip package, not a replacement for it.
Where SkyPilot is the wrong tool
The clearest limitation is that SkyPilot is not a managed service. The README points to a team deployment option through an API server, but the default posture is that you run the control plane yourself, on your own credentials, against your own accounts. If what you want is a vendor who pages someone at 3am when the scheduler is down, this is the wrong category of product.
The second constraint is that flexibility comes from breadth of provider support, and breadth is also a maintenance surface. The supported infrastructure list is long, and each entry implies credential handling, instance-type mapping and quota behaviour that SkyPilot has to track as providers change. The release notes for v0.13.0 mention "governance & robustness on API server", which suggests the API server path has been an area of hardening rather than a settled component.
Third, if your compute is one cluster that you fully control and you have no queueing, failover or idle-cleanup problem, the abstraction is overhead. You would be adding a YAML schema and a scheduler in front of a machine you already know how to use. The README's own framing is about unifying "multiple clusters, clouds, and hardware", and that is the condition under which the tool earns its place.
Finally, note the boundary the README does not document. It describes autostop as cleanup of idle resources, but it does not describe a rollback path for a cluster that was stopped when you did not want it stopped. Treat autostop as a policy to test on a throwaway cluster before you attach it to a shared one.
SkyPilot compared with Ray and Slurm
Ray and SkyPilot are often mentioned together, and the difference is in what each one owns. Ray is a distributed computing framework: you write your program against its APIs, and it handles task and actor distribution inside a cluster you have already provisioned. SkyPilot sits below that concern. It decides which infrastructure the work lands on and manages the lifecycle of that placement. A team running Ray on SkyPilot is using SkyPilot to solve the provisioning and failover problem and Ray to solve the in-cluster distribution problem. They are stacked, not substituted, and the repository reflects that: the examples tree includes distributed Ray training alongside distributed PyTorch.
Slurm is the closer comparison in interface terms, and the README makes the analogy itself, describing SkyPilot as offering "Slurm-like ease of use, cloud-native robustness" on Kubernetes. The real difference is the substrate. Slurm assumes a fixed set of machines that you own and administer, and its scheduling is confined to that set. SkyPilot assumes the set can change: clusters can be created and torn down, capacity can come from a cloud API, and a placement that fails can move somewhere else. If your machines are fixed and will stay fixed, Slurm's model is simpler and has decades of operational practice behind it. If your capacity is elastic or spread across providers, Slurm's assumption is the thing that hurts.
The repository also ships an examples/airflow directory, which is worth noting for teams whose orchestration already lives in Airflow. That is a different integration point from the scheduler comparison above: Airflow sequences work, SkyPilot places it.
Licence, releases and what upgrades cost
SkyPilot is Apache-2.0, which permits commercial use, modification and redistribution, and includes a patent grant. The repository ships the LICENSE file at the top level. This is a permissive licence with no copyleft obligation on your own code, and no separate commercial tier is described in the README. For the legal questions that follow from that (notably what happens when you redistribute a modified version), talk to your own counsel rather than treating a licence identifier as advice.
The release cadence visible in the repository is steady: v0.13.0 in July 2026, with a release candidate before it, and v0.13.1rc1 shortly after. The last push to the default branch was on 2026-09-10, so the codebase is moving. Release candidates appearing before final tags suggests the project tests changes in the open before tagging them, which is a reasonable signal for teams that pin versions.
Upgrade cost is dominated by the same surface that gives the tool its value. A new release can change how a provider is handled, and your task YAMLs may reference accelerator names or instance characteristics that providers rename. The practical approach is to pin the version in your environment and read the release notes before moving, because the notes are where the project records changes like the v0.13.0 Hugging Face storage and lifecycle hook additions. Nothing in the README describes an automatic migration path for task files.
Getting the first cluster right
The gap between reading about SkyPilot and depending on it is the first real cluster. The README's own path is one minute to install and two minutes to a first cluster, and that ordering is deliberate: the tool is meant to be evaluated by running something small.
Start with a single provider you already pay for and a task file copied from the examples directory, then change only the resources block. That isolates the variable. If the job lands, the credentials and the accelerator naming are correct, and everything after that is scheduling policy. If it does not land, you have one provider's configuration to debug instead of several.
Once that works, the interesting test is a deliberate failure. Request capacity that is unlikely to be available and see whether the job moves or errors. The README documents flexible provisioning and smart failover, and that behaviour is the reason to accept the extra layer in the first place. If failover does not fire in your setup, the value proposition is much thinner than the feature list suggests.
Editorial conclusion
Adopt SkyPilot if you already hold GPUs across more than one cloud or cluster and want one job interface plus autostop and failover instead of a per-provider script. Do not adopt it if you run a single fixed cluster with no queueing or failover problem, or if you need a managed control plane that someone else operates. Before committing, verify three things in your own account: that the credentials for each cloud are accepted, that the accelerators you name in a task YAML are actually available for that cloud, and that autostop behaves the way you expect on a throwaway cluster, because the README does not document rollback for a stopped cluster.
Frequently asked questions
Is SkyPilot free?
The project is licensed under Apache-2.0 and the README describes no paid tier or commercial edition. You still pay your cloud providers for the compute SkyPilot launches, since it runs everything inside your own accounts.
Is SkyPilot open source?
Yes. The repository is skypilot-org/skypilot, the primary language is Python, and the licence is Apache-2.0, with the LICENSE file at the top level of the repository.
How does SkyPilot work?
You describe a workload in a task YAML with its resource requirements, and SkyPilot selects a placement across your clouds, Kubernetes clusters or Slurm clusters, with documented smart failover when a placement cannot be satisfied. It also provides autostop for idle resource cleanup and binpacking on shared clusters.
How to install SkyPilot?
The README gives the uv form, uv pip install "skypilot[kubernetes,aws,gcp,azure,oci,nebius,lambda,runpod,fluidstack,paperspace,cudo,ibm,scp,seeweb,shadeform,verda]", choosing the providers you need. The installation docs also list pip, nightly and from-source options.
What does SkyPilot do?
It runs, manages and scales AI workloads on infrastructure you already own, presenting cloud accounts, Kubernetes clusters and Slurm clusters behind one interface. The README also lists autostop, binpacking and an intelligent scheduler as its utilization features.
How does SkyPilot compare with Slurm?
The README describes SkyPilot as offering Slurm-like ease of use with cloud-native robustness on Kubernetes. The difference is the substrate: Slurm schedules a fixed set of machines you administer, while SkyPilot can create and tear down clusters and move a job to different infrastructure.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/skypilot-org-skypilot)