Open-source project
spark-jobserver/spark-jobserver avatar
spark-jobserver/spark-jobserver

spark-jobserver: a REST job server for Apache Spark

REST job server for Apache Spark

2,836 stars967 forksScalaNOASSERTION

At a glance

What is it?
spark-jobserver puts a REST interface in front of Spark job, jar and context management, so applications can submit work without a Spark client. Here is how it is put together, how to run it, and where it stops being the right tool.
Who is it for?
Adopt spark-jobserver if you need a long-lived HTTP endpoint that submits Spark jobs and keeps contexts warm, and if you are prepared to build and run the server yourself from the repository. Do not adopt it if you need a currently released, versioned artifact that matches your Spark version, since the newest release listed in the repository is v0.11.1 from 2021-05-06 and the README's own version table stops at Spark 2.4.4.
Can I use it commercially?
Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
Is it still maintained?
Activity is slowing. The repository last received commits 7 months ago.
What is it written in?
Mainly Scala, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 24, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The gap spark-jobserver fills between an application and a Spark cluster

Submitting a Spark job normally means having a Spark distribution, a submission script and a jar on the machine that starts the work. That is fine for a scheduled batch pipeline. It is awkward when the thing that wants to run Spark is a web service, a notebook backend or an internal tool that already speaks HTTP and has no business carrying a Spark client.

spark-jobserver answers that with a RESTful interface for submitting and managing Apache Spark jobs, jars and job contexts. The README describes the model as "Spark as a Service". A client uploads a jar once, asks the server to create a context, and then posts job requests against that context. The server owns the Spark side.

The audience is narrow and specific: platform teams who want to hand application developers a stable HTTP contract instead of a Spark installation, and teams running several applications that should share warm Spark contexts rather than each paying startup cost. The README lists Ooyala, Netflix, Datadog, Target and others as users, and notes that the project is included in Datastax Enterprise. That is a statement about who has used it, not a measure of how well it fits your stack.

Contexts, jars and the DAO: how a request actually travels

The central object is the context. A context is a long-running SparkContext that the server keeps alive, and the README is explicit that this is what enables sub-second low-latency jobs. Two modes are documented. Ad-hoc mode creates a transient context for a single, unrelated job. Persistent context mode keeps the context around, which the README calls faster and required for related jobs.

That distinction drives the whole design. In ad-hoc mode you pay context startup on every request. In persistent mode you pay it once and then reuse RDDs and DataFrames through Named Objects, which the README describes as a way to cache and retrieve RDDs or DataFrames by name to improve sharing and reuse among jobs. If your jobs have nothing to share, persistent contexts buy you little and cost you memory that stays allocated.

Jars are uploaded separately, which the README frames as a step that makes job startup faster. Job and jar metadata is persisted through a pluggable DAO interface, so the server is not purely in-memory about what has been submitted. For isolation, the project offers a separate JVM per SparkContext, marked EXPERIMENTAL in the feature list, alongside an EXPERIMENTAL supervise mode and a beta HA deployment. Those three labels are worth reading as a group: the isolation, supervision and multi-server paths are the parts of the system the project itself flags as less settled.

The API surface is grouped into binaries, contexts, jobs and data, with a separate data API example in the documentation. The server also supports Spark SQL, Hive and streaming contexts, and has Python, Scala and Java job APIs.

Building and running spark-jobserver from the repository

The repository ships a Dockerfile that builds the server in one stage and runs it in another. The build stage installs sbt, downloads a Spark distribution, copies the source and runs the assembly task. The run stage copies the assembled jar and the Spark directory into a smaller image, exposes ports 8090 and 9999, and starts the server with spark-submit.

The image build arguments are declared at the top of the Dockerfile, so you can see exactly which versions the container targets:

dockerfile
ARG SBT_VERSION=1.2.8
ARG SCALA_VERSION=2.11.8
ARG SPARK_VERSION=2.4.5
ARG HADOOP_VERSION=2.7

The assembly command the Dockerfile runs is:

bash
sbt clean job-server-extras/assembly

The runtime stage sets three environment variables that matter if you are wiring the server into an existing deployment:

dockerfile
ENV SPARK_JOBSERVER_MEMORY=1G \
    LOGGING_OPTS="-Dlog4j.configuration=file:/opt/sparkjobserver/config/log4j.properties" \
    MANAGER_JAR_FILE=/opt/sparkjobserver/bin/spark-job-server.jar \
    MANAGER_CONF_FILE=/opt/sparkjobserver/config/jobserver.conf

Port 8090 carries the API and 9999 is exposed for JMX; the README links a separate JMX tips document. The container starts the server through spark-submit with the class spark.jobserver.JobServer, passing SPARK_JOBSERVER_MEMORY as driver memory and LOGGING_OPTS as both driver java options and executor extra java options.

If you are consuming the client API from an sbt project instead of running the server, the README gives this resolver, noting that earlier release binaries were removed after the sunset of Bintray and that only recent releases are available on the JFrog platform:

scala
resolvers += "Artifactory" at "https://sparkjobserver.jfrog.io/artifactory/jobserver/"

For a first real use, the README's development walk-through is the WordCountExample: package the jar, send it to the server, then run it either in ad-hoc mode with a transient context or in persistent context mode. The README also documents a giter8 template for creating a job server project from scratch, and a manual path if you already have an sbt project structure.

Where spark-jobserver is the wrong choice

The version table is the first thing to read, and it is a constraint rather than a feature. It maps 0.8.1 to Spark 2.2.0, 0.10.2 to Spark 2.4.4, and 0.11.1 to Spark 2.4.4 with Scala 2.11 and 2.12. The newest release listed is v0.11.1 from 2021-05-06. The repository's last push is 2026-03-03, and the repository is not archived, so there is activity, but the README's own table does not extend to Spark 3.x. If your cluster runs a Spark version outside that table, you are in territory the documentation does not cover.

The second constraint is operational. A long-running context holds cluster resources between requests. That is the point, and it is also the failure mode: an idle context is not free, and the resource profile of a context is fixed until you restart it. The README acknowledges this by describing context restart as the moment when you change resources. There is no documented autoscaling of a context in place.

Third, the parts you would want in production are the parts labelled EXPERIMENTAL or beta: separate JVM per context for isolation, supervise mode, and HA deployment. That is a coherent engineering choice for a project this age, but it means the isolation story is not a settled guarantee you can lean on without reading the source.

Finally, if all you need is to run a scheduled Spark job, spark-jobserver adds a server, a DAO and an HTTP layer between you and spark-submit for no benefit. The REST interface earns its place when something other than a Spark operator needs to trigger the work.

How it compares with Livy and with plain spark-submit

The closest alternative in the same problem space is Apache Livy, which also exposes a REST interface for submitting Spark jobs and managing sessions. The difference is in the shape of the abstraction. Livy's model centres on sessions and batches over the Spark interpreter and supports interactive code submission. spark-jobserver's model centres on compiled job jars plus named objects: you upload a jar, register a job class, and the server runs it against a context. That makes spark-jobserver a better fit when jobs are packaged artifacts with a defined interface, and a worse fit when you want to send ad-hoc code fragments to a cluster.

The other alternative is no server at all: run spark-submit from your application or scheduler. That keeps the Spark version coupling in your hands and removes a moving part. It costs you the warm context, so every invocation pays startup, and it means every caller needs a Spark distribution. The README's claim of sub-second low-latency jobs through long-running contexts is the specific thing you would be giving up.

Between the two, the deciding question is whether the callers can be trusted with a jar-based contract. If they can, spark-jobserver's jar upload plus named object sharing is a tighter model than interactive sessions. If they cannot, a session-oriented server fits better.

Maintenance, upgrade cost and the licence question

The release cadence visible in the repository is slow: v0.11.1 in May 2021, v0.11.0 in February 2021, v0.10.1 in November 2020. The last push to the default branch is 2026-03-03, so the repository is not abandoned, but the tagged releases are old and the README's version table stops at Spark 2.4.4. Anyone adopting this should plan on building from source rather than consuming a release artifact that matches a modern Spark, and should budget for the possibility of carrying patches.

The Dockerfile pins sbt 1.2.8, Scala 2.11.8, Spark 2.4.5 and Hadoop 2.7 as build arguments. Those pins are the upgrade surface: moving to a different Spark means changing the argument, rebuilding the assembly, and then checking whether the job API and the extras modules still compile. The repository has a scalastyle-config.xml and a run_tests.sh, plus integration tests under job-server-integration-tests, so there is a test harness to lean on during that work.

On licensing, the repository metadata reports the licence as NOASSERTION, which means GitHub could not map the LICENSE.md file to a known licence identifier. The README has a License section but its contents are not stated in the repository description. Read LICENSE.md directly and have someone qualified assess it for your distribution model; nothing in this article should be read as legal advice.

Editorial conclusion

Adopt spark-jobserver if you need a long-lived HTTP endpoint that submits Spark jobs and keeps contexts warm, and if you are prepared to build and run the server yourself from the repository. Do not adopt it if you need a currently released, versioned artifact that matches your Spark version, since the newest release listed in the repository is v0.11.1 from 2021-05-06 and the README's own version table stops at Spark 2.4.4. Before committing, verify that the Scala and Spark combination you need is covered by build.sbt and the version table, and check the LICENSE.md file directly, because the repository metadata reports the licence as NOASSERTION rather than naming one.

Frequently asked questions

What is spark-jobserver?

It is a REST job server for Apache Spark that provides a RESTful interface for submitting and managing Spark jobs, jars and job contexts. It was originally started at Ooyala and this repository is now the main development repo.

What does spark-jobserver do when you submit a job?

The client posts the job against a context, which is a long-running SparkContext the server keeps alive. In ad-hoc mode the context is transient and serves a single unrelated job, while persistent context mode keeps it around and is described as required for related jobs.

Which Spark versions does spark-jobserver support?

The README's version table maps 0.8.1 to Spark 2.2.0, and both 0.10.2 and 0.11.1 to Spark 2.4.4 with Scala 2.11 and 2.12. The table does not list a Spark 3.x release.

How do I install spark-jobserver?

The repository provides a Dockerfile that builds the assembly with sbt and runs the server via spark-submit, exposing port 8090 for the API and 9999 for JMX. The README also documents a manual deployment path and a context-per-JVM option.

Can spark-jobserver share data between jobs?

Yes, through Named Objects, which the README describes as a way to cache and retrieve RDDs or DataFrames by name to improve object sharing and reuse among jobs. This requires a persistent context rather than an ad-hoc one.

Official sources

  1. Issues
  2. README
  3. Releases
  4. spark-jobserver/spark-jobserver on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/spark-jobserver-spark-jobserver.svg)](https://hysenlabs.com/projects/spark-jobserver-spark-jobserver)