Self-hosted service
loft-sh/vcluster avatar
loft-sh/vcluster

vCluster: tenant clusters with their own API server on shared infrastructure

vCluster creates tenant clusters: fully isolated environments delivered as managed Kubernetes, or as the foundation for Slurm, Ray, Run:ai and inference clusters. Each gets its own API server, CRDs and RBAC, and runs on an existing cluster or standalone on bare metal. CNCF Certified Kubernetes.

11,323 stars608 forksGoApache-2.0

At a glance

What is it?
vCluster is the Apache-2.0 Go engine behind Loft's tenant cluster product. It gives each team a certified Kubernetes control plane of its own, and the cost is that you now operate two clusters instead of one.
Who is it for?
Adopt vCluster when you need per-tenant CRDs, RBAC and API isolation on infrastructure you already run, and when your team can operate a second control plane per tenant. Do not adopt it if you only need namespace-level quotas and network policy: Capsule's tenant model is a fraction of the operational surface.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 1 day ago.
What is it written in?
Mainly Go, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The problem vCluster solves: tenants that need their own API server, not another namespace

Kubernetes namespaces are a weak tenancy boundary. They share one API server, one set of CRDs, one etcd, and one cluster-scoped RBAC surface. The moment a tenant wants to install a custom resource definition, run a cluster-scoped operator, or hold cluster-admin over their own resources, namespaces stop working. The usual answer is one full cluster per tenant, which means one control plane per tenant: more etcd, more certificates, more upgrade jobs.

vCluster takes a third path. Each tenant gets a cluster with its own API server, its own CRDs and its own RBAC, but that control plane runs on infrastructure you already have or standalone on bare metal. The README puts it as "Every cluster becomes a product you can ship." The intended audience is platform engineering teams serving internal teams or external customers, and the AI infrastructure crowd in particular: the repository topics include multi-tenancy, platform-engineering and cloud-native, and the project advertises itself as the foundation for Slurm, Ray, Run:ai and inference clusters.

One claim in the README deserves scrutiny rather than repetition: "no in-cluster agent pods, and no lateral path between environments." Whether there is a lateral path depends on your network policy and on which architecture mode you pick, and the comparison table further down makes that explicit. Treat the isolation sentence as a design intent, not a guarantee you inherit for free.

Inside the tenant cluster: a syncer between two API servers

The mechanism is a control plane plus a translation layer. A tenant cluster runs its own API server, backed by its own storage, and the workloads it schedules land on nodes that belong to a host cluster (or, in the standalone mode, on hardware you point it at). Because the tenant API server is upstream Kubernetes, the README states that kubectl, Helm, Argo, Crossplane, operators and CRDs work against it unmodified. That is the whole point of the design: no client-side fork, no special SDK.

What makes this non-trivial is state. A pod created inside the tenant cluster has to exist somewhere real, and the host cluster has to be able to schedule it. The project ships a syncer that reconciles objects between the virtual and host API servers, which is why the repository carries a large pkg/ tree and a vendor/ directory next to cmd/vcluster and cmd/vclusterctl. The two binaries map to the two halves: cmd/vcluster is the control plane that runs inside the tenant, cmd/vclusterctl is the CLI you drive it with.

The README also describes a control plane that stays invisible to tenants, and a demo where a tenant installs a CRD as cluster-admin while the same command in another tenant returns nothing. That is the concrete behaviour to test in your own environment, because it is exactly what namespaces cannot give you.

Install and create your first tenant cluster

The README's quickstart assumes a running Kubernetes cluster with kubectl configured. The CLI installs through Homebrew, and the package name is loft-sh/tap/vcluster:

bash
brew install loft-sh/tap/vcluster

With the CLI on your PATH, creating a tenant cluster takes one command. The README uses --namespace team-x, and the namespace is where the vCluster control plane objects live in your host cluster:

bash
vcluster create my-vcluster --namespace team-x

The CLI connects you to the new tenant, so the next command runs against the tenant API server and not the host. The README's example asks for namespaces, and you should see the tenant's own namespace list, not the host cluster's:

bash
kubectl get namespaces

If you have no Kubernetes cluster at all, the README points at vind, described as "vCluster in Docker". The create command takes a driver flag and runs the whole thing in Docker containers:

bash
vcluster create my-vcluster --driver docker
kubectl get namespaces

The README lists the features vind carries over from the full product: UI, sleep and resume, LoadBalancer, image cache, and external nodes joining over VPN. It also links a browser playground on Killercoda for readers who want to look before installing anything.

Four architectures, and the isolation each one actually buys

The README's architecture table is the most useful page of the material, because it separates two decisions that people usually conflate: where the control plane runs, and where tenant workloads land. Four modes are listed. Shared Nodes requires a control plane cluster and gives no node, CNI or CSI isolation; it is aimed at dev/test and density. Dedicated Nodes still requires a control plane cluster but adds node isolation, and the README positions it for production tenants. Private Nodes adds CNI and CSI isolation on top and is marked bare-metal ready, aimed at compliance and GPU workloads. Standalone needs no control plane cluster at all and targets AI factories and edge.

The honest reading is that "vCluster isolates tenants" is only true at the API layer in the first two modes. If your tenants need their own CNI or CSI, you are in Private Nodes or Standalone, and both of those carry more operational weight than the quickstart suggests. The README does not document rollback for a mode change, and it does not describe what happens to running tenant workloads during an upgrade of the host cluster.

The cluster-type table adds a second boundary. Kubernetes and nested clusters come from this repository. Inference, Ray, Run:ai, Slurm (marked Beta) and the coming agent sandbox clusters are delivered through vCluster Platform, which the README describes as a separate product with a free tier capped at 64 CPUs and 32 GPUs. The repository you are reading is, in the README's own words, "the open-source engine that every one of them is built on."

Where vCluster is the wrong tool

If your tenancy requirement is quota enforcement plus network policy between teams on one shared API server, vCluster is overkill. You will be running an extra API server, an extra datastore and an extra upgrade path per tenant, and the syncer becomes one more component that can fall out of sync. Capsule is the natural comparison here: it keeps a single API server and expresses tenancy through namespaces grouped under a tenant object, with policy applied by the existing controllers. The difference is architectural, not cosmetic. Capsule has no second control plane to operate; vCluster has one per tenant, and that is what buys per-tenant CRDs and cluster-scoped RBAC.

kind is a different kind of wrong tool. It is a local test harness that runs Kubernetes nodes in containers, and the README positions vind explicitly as "like kind, but with the full vCluster feature set", listing UI, sleep and resume, LoadBalancer, image cache and VPN-joined external nodes. If you want a throwaway cluster on a laptop, kind is simpler. If you want a long-lived tenant cluster with an API surface a customer can be handed, that is the gap vCluster is aiming at.

The third case is scale of operators. Each tenant cluster is a control plane you must patch, back up and monitor. A platform with fifty tenants and no automation around vCluster lifecycle will spend more time on control plane maintenance than on the workloads it was built to serve.

Licence, maintenance and what an upgrade actually costs you

The repository is Apache-2.0, which permits commercial use, modification and redistribution, and it carries no copyleft obligation on your own code. Note the split the README draws: this repository is the open-source engine, while the Platform, which delivers the Slurm, Ray, Run:ai, inference and agent sandbox cluster types, is a separate product with its own free tier. If your plan depends on one of those cluster types, the Apache-2.0 licence on this repository does not cover the component you need. That is a factual boundary, not a legal opinion; take licence questions to your own counsel.

Maintenance is easy to state from the facts: the repository is not archived, and the last push was on 2026-09-19, two days before this article's reference point. The most recent release is v0.37.1, dated 2026-09-14, preceded by two release candidates on the same day. The version number is the thing to plan around. A 0.x line means the upgrade path is not frozen, and the project's own release cadence, three artifacts inside one day, tells you patches arrive quickly. Budget for reading release notes between minor versions rather than assuming a drop-in upgrade.

The upgrade cost that the README does not address is the one that matters most in production: what happens to tenant workloads when the host cluster's nodes are drained, and how a tenant control plane is rolled forward independently. The material documents neither, so treat both as questions for a proof of concept on your own infrastructure before you commit a tenant to it.

Editorial conclusion

Adopt vCluster when you need per-tenant CRDs, RBAC and API isolation on infrastructure you already run, and when your team can operate a second control plane per tenant. Do not adopt it if you only need namespace-level quotas and network policy: Capsule's tenant model is a fraction of the operational surface. Before committing, verify the architecture mode you actually need against the comparison table, because shared nodes and private nodes differ on CNI and CSI isolation, and confirm which cluster types require vCluster Platform rather than this repository.

Frequently asked questions

What is a vCluster?

It is a tenant cluster: a Kubernetes environment with its own API server, CRDs and RBAC that runs on an existing cluster or standalone on bare metal. Because the API server is upstream Kubernetes, kubectl, Helm, Argo, Crossplane, operators and CRDs work against it unmodified.

What is the difference between vCluster and vCluster Platform?

This repository is the open-source engine. The Platform is the product that delivers the inference, Ray, Run:ai, Slurm and agent sandbox cluster types on top of it, and it has its own free tier capped at 64 CPUs and 32 GPUs.

how to install vcluster

The README's quickstart installs the CLI with Homebrew using the package loft-sh/tap/vcluster, then creates a tenant cluster with vcluster create my-vcluster --namespace team-x against an existing Kubernetes cluster. If you have no cluster, the same create command accepts --driver docker to run it in Docker.

is vcluster open source

Yes. This repository is licensed Apache-2.0 and is written in Go, with the engine source under cmd/ and pkg/. The vCluster Platform that sits on top is a separate product.

vcluster vs capsule

Capsule applies tenancy on a single shared API server through namespaces grouped under a tenant object, so there is no second control plane to operate. vCluster runs a separate API server per tenant, which is what gives each tenant its own CRDs and cluster-scoped RBAC.

vcluster vs kind

kind runs Kubernetes nodes in containers as a local test harness. The README describes vind as "like kind, but with the full vCluster feature set", listing UI, sleep and resume, LoadBalancer, image cache and external nodes joining over VPN.

Official sources

  1. License: Apache-2.0
  2. loft-sh/vcluster on GitHub
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/loft-sh-vcluster.svg)](https://hysenlabs.com/projects/loft-sh-vcluster)