Self-hosted service
kubeflow/spark-operator avatar
kubeflow/spark-operator

Kubeflow Spark Operator: running spark-submit through Kubernetes custom resources

Kubernetes operator for managing the lifecycle of Apache Spark applications on Kubernetes.

3,155 stars1,535 forksPythonApache-2.0

At a glance

What is it?
The Kubeflow Spark Operator turns Spark applications into Kubernetes objects called SparkApplication, and the controller runs spark-submit for you. It is for teams already on Kubernetes who want cron, retries and pod customisation declared in YAML rather than in shell scripts.
Who is it for?
Adopt it if your Spark jobs already run on Kubernetes and you want submission, cron, restart policy and pod customisation expressed as custom resources the cluster reconciles. Do not adopt it if you only need to run spark-submit occasionally from a laptop, or if you cannot take on CRD lifecycle work.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository received new commits within the last day.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The submission problem the Spark Operator removes

Running Spark on Kubernetes without an operator means someone, or something, has to call spark-submit with a long argument list, watch the driver pod, and decide what happens when it fails. That logic usually ends up in a CI job, a cron container or a script that holds credentials. The Spark Operator moves it into the cluster. The README describes the project as making Spark applications "as easy and idiomatic as running other workloads on Kubernetes", using Kubernetes custom resources to specify, run and surface the status of applications.

The audience is narrow and specific: platform or data engineers who already have a Kubernetes cluster and want Spark jobs to look like every other workload there. The README lists the supported range as Spark 2.3 and above, which must support Kubernetes as a native scheduler backend. If your Spark version predates that, this project is not the tool.

How the controller, CRDs and webhook fit together

Two custom resources carry the model: SparkApplication and ScheduledSparkApplication. The README states the operator automatically runs spark-submit on behalf of users for each SparkApplication eligible for submission, so the controller is the thing that actually launches work, not a wrapper around your own submit call.

Scheduling is handled natively through cron support for scheduled applications, and the repository depends on github.com/robfig/cron/v3 for that. Restart behaviour is also declarative: the README lists automatic application restart with a configurable restart policy, automatic retries of failed submissions with optional linear back-off, and automatic re-submission when a SparkApplication object's specification is updated. A mutating admission webhook handles pod customisation that Spark itself cannot express, such as mounting ConfigMaps and volumes and setting pod affinity or anti-affinity. Metrics are exported to Prometheus at both application level and for drivers and executors.

The current API version is v1beta2 and the project describes its own status as beta. That matters more than the feature list: the CRD schema is the contract you build against, and it is explicitly not stable yet.

Installing the Spark Operator with the Helm chart

The README gives Helm as the first install path. You add the chart repository, install into a dedicated namespace, and wait for deployments to become ready before creating anything. The namespace in the example is spark-operator.

bash
helm repo add --force-update spark-operator https://kubeflow.github.io/spark-operator

helm install spark-operator spark-operator/spark-operator \
    --namespace spark-operator \
    --create-namespace \
    --wait

After the install returns, the operator's deployments should be ready in that namespace. Then apply the example application the README points at, which creates a SparkApplication named spark-pi in the default namespace.

bash
kubectl apply -f https://raw.githubusercontent.com/kubeflow/spark-operator/refs/heads/master/examples/spark-pi.yaml

kubectl get sparkapp spark-pi

The kubectl get output is the first real signal that the loop works: the operator submitted the job and the status field is being updated by the controller. Cleanup is a normal delete.

bash
kubectl delete sparkapp spark-pi

The prerequisites are Kubernetes 1.16 or newer, Helm 3 or newer, and kubectl 1.16 or newer. If you prefer manifests over Helm, the README documents a kustomize path: clone the repository and apply config/default with --server-side --force-conflicts, then apply config/spark-rbac in each namespace where Spark applications will run. The README's example uses the spark-operator-spark service account in the default namespace. Skipping that second step is a common way to end up with applications that never start.

The v1beta1 to v1beta2 migration is not optional

The README is direct about this: if your manifests use the v1beta1 API version, change apiVersion to sparkoperator.k8s.io/v1beta2, delete the previous CustomResourceDefinitions named sparkapplications.sparkoperator.k8s.io and scheduledsparkapplications.sparkoperator.k8s.io, and replace them with the v1beta2 version by installing the latest operator or running kubectl create -f config/crd/bases.

That is a CRD replacement, not a patch. Anyone upgrading an existing cluster should treat it as a planned change with a window, because the custom resources are the interface every manifest, dashboard and pipeline in the cluster depends on. The README does not document a rollback path for the CRD swap, so plan the change assuming you need your own backup of the resource definitions before you start.

When the Spark Operator is the wrong layer

The operator assumes Kubernetes is already your compute substrate and that you want the cluster to own job lifecycle. If your Spark work runs on YARN or on a managed service where you never touch pods, adding this project means running a controller, a webhook and CRDs to reproduce scheduling you already have.

There is also a version coupling to respect. The README's version matrix ties operator versions to API versions, Kubernetes versions and a base Spark version, and the Dockerfile shows the operator image built on docker.io/apache/spark:4.0.4. Upgrading Spark is therefore not a local decision: you are choosing an operator release that ships the matching base image. Teams that upgrade Spark on their own schedule will find that constraint annoying, and it is a real one rather than a documentation gap.

Finally, the mutating webhook is on the admission path for the pods it matches. The README does not describe webhook failure policy here, so verify how your installation handles the webhook being unavailable before you rely on it in a busy cluster.

Spark Operator compared with Airflow's SparkSubmitOperator

Airflow is the alternative most teams already have, and it solves a different half of the problem. The SparkSubmitOperator runs spark-submit from an Airflow worker and reports success or failure back into a DAG. The Spark Operator does not orchestrate a graph of tasks; it reconciles a single application object, and the README's cron support covers periodic runs, not dependencies between jobs.

The practical difference shows up in ownership. With Airflow, the schedule, retries and alerting live in the DAG, and the cluster only runs pods. With the Spark Operator, the restart policy, retry back-off and re-submission on spec change live in the SparkApplication, and the cluster reconciles them. If you need to chain Spark jobs with other systems, keep Airflow. If you want the cluster to keep a job running without a scheduler process watching it, the operator is the layer that does that. The two are not mutually exclusive: Airflow can submit a SparkApplication through the Kubernetes API instead of calling spark-submit itself.

Maintenance, licence and upgrade cost

The repository is not archived and the last push was on 2026-09-18, so it is under current development. Releases are frequent enough to matter: v2.5.0 on 2026-03-19, v2.5.1 on 2026-06-15 and v2.5.2 on 2026-07-31. The Go module path is github.com/kubeflow/spark-operator/v2 and go.mod requires Go 1.25.0 with a toolchain of go1.25.11, so building the operator from source tracks a recent Go release. The module depends on controller-runtime v0.23.3 and Kubernetes libraries at v0.35.4, which means operator upgrades pull Kubernetes API changes along with them.

On licensing: the project is Apache-2.0, the same licence as Apache Spark, and the README carries an OpenSSF Best Practices badge and a FOSSA licence-scan badge. Apache-2.0 includes a patent grant and requires you to preserve notices; it is permissive enough to embed in a commercial platform. That is a description of the licence, not legal advice: if you redistribute a modified operator image, check the notice requirements with your own counsel. The Dockerfile pins base images by digest, which helps reproducibility but also means base image updates arrive as explicit changes you have to make.

Editorial conclusion

Adopt it if your Spark jobs already run on Kubernetes and you want submission, cron, restart policy and pod customisation expressed as custom resources the cluster reconciles. Do not adopt it if you only need to run spark-submit occasionally from a laptop, or if you cannot take on CRD lifecycle work. Before installing, confirm your existing manifests use apiVersion sparkoperator.k8s.io/v1beta2, because the README states v1beta1 manifests must be converted and the old CRDs deleted and replaced.

Frequently asked questions

What is the Kubeflow Spark Operator?

It is a Kubernetes operator for Apache Spark that uses custom resources to specify, run and surface the status of Spark applications, and runs spark-submit on behalf of users for each SparkApplication eligible for submission. The README describes its goal as making Spark applications as easy and idiomatic as running other workloads on Kubernetes.

What is the Spark Operator in Kubernetes?

It is a controller that watches SparkApplication and ScheduledSparkApplication custom resources and reconciles them, adding cron scheduling, restart policy, retry with optional linear back-off, a mutating admission webhook for pod customisation, and Prometheus metrics. It requires Spark 2.3 and above with Kubernetes as a native scheduler backend.

How do I install the Kubeflow Spark Operator with Helm?

The README's quick start adds the chart repository at https://kubeflow.github.io/spark-operator, then installs the spark-operator chart into the spark-operator namespace with --create-namespace and --wait. A kustomize install path is also documented, and it requires applying config/spark-rbac in each namespace where Spark applications run.

Which Kubernetes and Spark versions does the Spark Operator require?

The README lists Kubernetes 1.16 or newer, Helm 3 or newer and kubectl 1.16 or newer as prerequisites, and states that Spark 2.3 and above is supported where Kubernetes is a native scheduler backend. A version matrix in the README maps operator versions to API versions, Kubernetes versions and a base Spark version.

Does the Spark Operator support scheduled Spark jobs?

Yes. The README lists native cron support for running scheduled applications, exposed through the ScheduledSparkApplication custom resource, and the repository depends on github.com/robfig/cron/v3. That covers periodic runs, not dependencies between separate jobs.

Official sources

  1. kubeflow/spark-operator on GitHub
  2. License: Apache-2.0
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/kubeflow-spark-operator.svg)](https://hysenlabs.com/projects/kubeflow-spark-operator)