Open-source project
ray-project/kuberay avatar
ray-project/kuberay

KubeRay: running Ray applications on Kubernetes with three custom resources

Project brief: A toolkit to run Ray applications on Kubernetes. Kubectl Plugin (Beta): Starting from KubeRay v1.3.0, you can use the kubectl ray plugin to simplify common workflows when deploying Ray on Kubernetes.

2,709 stars877 forksGoApache-2.0

At a glance

What is it?
KubeRay is the Kubernetes operator that turns RayCluster, RayJob and RayService into first-class cluster objects. It suits teams already on Kubernetes who need Ray, and it is a poor fit if you have no cluster to run it on.
Who is it for?
Adopt KubeRay if you already run Kubernetes and want Ray clusters, jobs and Serve deployments managed as custom resources rather than by hand. Do not adopt it if you have no cluster or no Ray workload, because the operator only manages Ray; it is not a general scheduler and not a substitute for Ray itself.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository received new commits within the last day.
What is it written in?
Mainly Go, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The gap KubeRay fills between Ray and a Kubernetes cluster

Ray is a distributed compute engine. Kubernetes is a container orchestrator. Neither knows how to create the other's objects, so without an operator you are left writing manifests and reconciling them yourself. KubeRay is the Kubernetes operator that closes that gap: the README describes it as a toolkit that "simplifies the deployment and management of Ray applications on Kubernetes", and its core component supplies three custom resource definitions, RayCluster, RayJob and RayService.

The audience is specific. You need a Kubernetes cluster, a Ray workload, and a reason to want that workload expressed as a Kubernetes object rather than as a shell command. Teams that already run Kubernetes and want Ray to live inside the same control plane, with the same kubectl habits, are the intended users. If your Ray work happens on a laptop or on a managed Ray offering, the operator adds a layer you will not use.

One structural detail matters for anyone reading the repository. Since September 2023, user-facing KubeRay documentation lives on the Ray documentation site, not in this repository. The repository keeps only development and maintenance documentation. That split means the README is a signpost, not a manual, and the quickstarts it links to are hosted elsewhere.

How the operator reconciles RayCluster, RayJob and RayService

KubeRay core is the official, fully maintained component, and it is built on the Kubernetes controller pattern. The repository's go.mod lists sigs.k8s.io/controller-runtime alongside k8s.io/client-go and k8s.io/api, which is the standard stack for a Go operator that watches custom resources and reconciles cluster state toward the declared spec.

The three resources divide the work. RayCluster is the base object: KubeRay manages its lifecycle, including creation and deletion, autoscaling, and fault tolerance. RayJob wraps that: KubeRay creates a RayCluster, waits for it to be ready, submits the job, and can be configured to delete the cluster once the job finishes. RayService combines a RayCluster with a Ray Serve deployment graph, and the README states it offers zero-downtime upgrades for the RayCluster plus high availability. That last claim is the one worth testing against your own traffic patterns, because zero downtime depends on how your Serve graph handles in-flight requests during a rollout.

Around the core sit optional pieces, each with its own maturity label. The kubectl ray plugin is Beta from KubeRay v1.3.0 and is aimed at people who are not fluent in Kubernetes. The KubeRay APIServer is Alpha and gives a simplified configuration layer that some organizations use to back their own user interfaces. The KubeRay Dashboard arrived in v1.4.0 and the README calls it Experimental and not yet production-ready, while inviting feedback. Read those labels literally: Beta and Alpha components are not the same commitment as the core.

Getting the operator running and submitting a first RayJob

The repository carries a helm-chart/ directory at the top level, and the README's quickstart links point at the Ray documentation for RayCluster, RayJob and RayService. The chart installs the operator and its CRDs into your cluster. The README's quickstart section is the entry point for the install and for each resource type, and it directs readers to the Ray documentation rather than reproducing the manifests here.

What you should confirm after the operator is up is that the custom resource kinds are registered. The repository layout gives you the names to look for: ray-operator/ holds the operator, and the three kinds it manages are RayCluster, RayJob and RayService. The kubectl ray plugin, Beta from v1.3.0, exists to shorten these workflows, and the README points readers who are unfamiliar with Kubernetes toward it.

For a first real use, pick RayJob rather than RayCluster. The README describes the flow directly: KubeRay creates a RayCluster, submits the job when the cluster is ready, and can be configured to delete the cluster once the job finishes. That means one object gives you a cluster, a job and cleanup, which is a shorter path to a working result than managing the cluster yourself. The exact spec fields for the job entrypoint and the head group are documented in the RayJob quickstart on the Ray documentation site, not in this repository, so copy the structure from there.

Where KubeRay is the wrong tool

The clearest failure mode is a mismatch of assumptions. KubeRay manages Ray on Kubernetes; it does not give you Kubernetes, and it does not give you Ray. If your team has no cluster, adopting KubeRay means adopting an orchestrator, a container registry workflow, and a networking model before the first Ray task runs. That is a large amount of work for a problem the operator did not create.

A second boundary is component maturity. The README labels the kubectl plugin Beta, the APIServer Alpha, and the Dashboard Experimental and explicitly not production-ready. If your platform depends on a stable management API or a UI, those components are not the part of KubeRay to build on yet. The core CRDs are the maintained surface; everything else is optional and carries a caveat in the project's own words.

A third is documentation location. Because user-facing docs moved to the Ray documentation site, the repository will not answer operational questions about the custom resources. Anyone who expects to find the operator manual in the repository, or who pins a workflow to files under docs/ here, will be reading development material instead. That is a design decision by the project, not an oversight, but it changes where you look when a reconcile loop misbehaves.

Finally, KubeRay is not a scheduler for arbitrary workloads. It integrates with queuing systems such as Volcano, Apache YuniKorn and Kueue, as the README's ecosystem list shows, but the queueing logic lives in those systems. If your problem is fair-share scheduling across mixed workloads, KubeRay is a participant, not the answer.

KubeRay compared with KServe and with plain Ray

The two comparisons people search for are KubeRay against KServe and KubeRay against Ray itself, and they are different questions.

KubeRay versus Ray is a question about layers, not competitors. Ray is the distributed compute engine; KubeRay is the operator that runs Ray clusters on Kubernetes. You do not choose one instead of the other. You choose whether Ray runs under Kubernetes at all. Without Kubernetes, Ray runs directly and KubeRay has nothing to manage.

KubeRay versus KServe is a genuine architectural choice for model serving, and the difference is what each treats as the unit of deployment. KServe is a Kubernetes-native model serving layer built around its own inference service abstractions. KubeRay's RayService is a RayCluster plus a Ray Serve deployment graph, so the serving logic is expressed in Ray Serve and the operator's job is to keep that graph and its cluster alive, with zero-downtime upgrades and high availability as stated goals. If your serving code is already Ray Serve, RayService keeps it in one system. If your serving stack is built around KServe's abstractions and you only need Ray for training, adding RayService means maintaining two serving paths. The deciding question is where your inference code already lives, not which project is larger.

Maintenance, release cadence and the Apache-2.0 licence

KubeRay is not archived, and the last push was on 2026-08-20, which is the same date as the v1.7.0 release. The release history shows v1.7.0-rc.0 on 2026-08-14, v1.7.0 on 2026-08-20, and v1.6.2 before that on 2026-06-18. That is a project shipping release candidates and patch releases on a regular cadence, and the README's own framing calls the core component "official, fully-maintained".

Upgrade cost is where the practical work sits. The repository's go.mod declares go 1.26.0 and pins Kubernetes libraries at v0.37.0 for k8s.io/api, k8s.io/apimachinery and k8s.io/client-go, with controller-runtime at v0.24.1. Those versions tell you what the operator is built against and therefore roughly which Kubernetes API surface it expects. Before upgrading the operator, check the release notes for the version you are moving to and confirm the Kubernetes and Ray versions you run are in range. Treat the operator and the Ray image as two things that must be compatible, not one.

The licence is Apache-2.0, a permissive licence that permits commercial use and modification and includes a patent grant. That is a summary of what the licence is, not legal advice; if your organization has policies about copyleft, attribution or patent clauses, route the LICENSE file through whoever handles that.

Editorial conclusion

Adopt KubeRay if you already run Kubernetes and want Ray clusters, jobs and Serve deployments managed as custom resources rather than by hand. Do not adopt it if you have no cluster or no Ray workload, because the operator only manages Ray; it is not a general scheduler and not a substitute for Ray itself. Before committing, verify which KubeRay version matches your Kubernetes and Ray versions, check whether the Beta kubectl ray plugin and the Experimental dashboard are acceptable for your environment, and read the Ray documentation pages for the RayCluster, RayJob and RayService quickstarts since the repository itself only carries development documentation.

Frequently asked questions

What is the difference between Ray and KubeRay?

Ray is the distributed compute engine, and KubeRay is a Kubernetes operator that deploys and manages Ray applications on Kubernetes. They are layers rather than alternatives: KubeRay manages Ray clusters, jobs and Serve deployments as custom resources.

What are the key differences between KServe and KubeRay?

KServe is a Kubernetes-native model serving layer with its own inference abstractions, while KubeRay's RayService is a RayCluster plus a Ray Serve deployment graph, so the serving logic lives in Ray Serve. The choice depends on where your inference code already lives.

What is KubeRay?

KubeRay is an open-source Kubernetes operator that simplifies deploying and managing Ray applications on Kubernetes. Its core component provides three custom resource definitions: RayCluster, RayJob and RayService.

What is the KubeRay operator?

The operator is the official, fully-maintained component of KubeRay. It manages the lifecycle of RayCluster, including creation, deletion, autoscaling and fault tolerance, and it creates and cleans up clusters for RayJob and RayService.

What is the Kubernetes platform?

KubeRay does not provide a Kubernetes platform; it assumes you already have one. It is a Kubernetes operator that runs Ray applications on that cluster, and the README's quickstarts point to the Ray documentation for the install and for each custom resource.

Is KubeRay the same as Ray?

No. Ray is the distributed compute engine, and KubeRay is a Kubernetes operator that manages Ray applications on Kubernetes. The operator supplies RayCluster, RayJob and RayService, and it manages the RayCluster lifecycle including autoscaling and fault tolerance.

Official sources

  1. Official README
  2. Project repository
  3. Release notes
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/ray-project-kuberay.svg)](https://hysenlabs.com/projects/ray-project-kuberay)