# docker-spark: running a Spark standalone cluster on Docker Compose or Kubernetes

> The big-data-europe/docker-spark repository packages Apache Spark into Docker images that set up a standalone cluster under Docker Compose or Kubernetes, with templates for building application containers on top. Spark 3.3.0 is the highest supported version, the repository carries no GitHub releases, and the INIT_DAEMON_STEP environment variable works only inside a BDE pipeline.

**big-data-europe/docker-spark** — Apache Spark docker image 

- Repository: https://github.com/big-data-europe/docker-spark
- Stars: 2,048 · Forks: 682
- Language: Shell
- License: not declared
- Published: 2026-10-09 · Updated: 2026-10-09 · Language: en
- Canonical page: https://hysenlabs.com/projects/big-data-europe-docker-spark

## One master, two workers and a history server form the default Compose cluster

The docker-compose.yml in the repository defines four services: spark-master, spark-worker-1, spark-worker-2 and spark-history-server. spark-master publishes port 8080 for its web UI and port 7077 for the Spark protocol. spark-worker-1 maps its internal 8081 to the same host port, and spark-worker-2 maps its internal 8081 to host port 8082. spark-history-server exposes port 18081.

All four services use the bde2020 image set. Tags follow a structured scheme: bde2020/spark-master:3.3.0-hadoop3.3 encodes both the Spark version and the Hadoop variant, so changing the version means updating the tag in every service definition simultaneously.

Workers locate the master through the SPARK_MASTER environment variable, which the compose file sets to spark://spark-master:7077. Docker's bridge networking resolves spark-master as a hostname because the container name and the internal DNS record match. No other configuration in the compose file handles this discovery.

Spark job history is a separate concern. The history server mounts /tmp/spark-events-local on the host to /tmp/spark-events inside its container. Docker creates an empty directory at that host path if it does not exist before the cluster starts, so the history server comes up cleanly but shows no completed jobs until workers write event logs to that location.

## The INIT_DAEMON_STEP variable in spark-master belongs to the BDE pipeline, not Spark

spark-master carries one environment variable absent from the worker and history-server definitions: INIT_DAEMON_STEP. The compose file sets it to setup_spark, and the README says to fill it in as configured in your pipeline, referring to the BDE (Big Data Europe) pipeline. Compose instructions appear in the context of integrating into a BDE pipeline.

For teams using docker-spark outside the BDE ecosystem, the variable's purpose depends on pipeline tooling that is not part of this repository. Standard Spark configuration keys use a different naming convention, so INIT_DAEMON_STEP is not a Spark setting. The README does not say what to configure it to when no BDE orchestrator is present, and it does not describe consequences of leaving the compose file value unchanged.

The repository links the Compose instructions to the app-bde-pipeline GitHub project for more context. Anyone adopting docker-spark as a standalone cluster image without the surrounding pipeline should read that link to understand what the variable signals and whether it needs to be present.

## Starting master and worker with docker run uses --link, a flag Docker deprecated years ago

Outside of Compose, the README gives two commands for starting the cluster manually:

```bash
docker run --name spark-master -h spark-master -d bde2020/spark-master:3.3.0-hadoop3.3
```

```bash
docker run --name spark-worker-1 --link spark-master:spark-master -d bde2020/spark-worker:3.3.0-hadoop3.3
```

The -h spark-master flag sets the master container's hostname. The --link flag on the worker creates an /etc/hosts entry pointing spark-master at the master's current IP address. Those two together make the address spark://spark-master:7077 resolvable from inside the worker container.

--link is legacy container linking that Docker marked for eventual removal after user-defined networks were introduced as the recommended alternative. It does not update automatically if the master container is recreated, so the worker would lose contact with a replacement master. For a single local test, the commands work. For a CI environment or any setup that recreates containers regularly, the Compose file is the correct path, since Docker Compose uses a named bridge network and DNS resolution rather than static /etc/hosts entries.

## The version matrix spans Spark 1.5.1 to 3.3.0, with OpenJDK 11 in only two builds

The README lists supported combinations spanning from Spark 1.5.1 for Hadoop 2.6 up to Spark 3.3.0 for Hadoop 3.3 with OpenJDK 8 and Scala 2.12. Every Spark 3.x combination specifies Scala 2.12. Older 2.x builds that target Hadoop 2.7+ list only Spark version, Hadoop version and JDK without naming a Scala version.

OpenJDK 8 covers the large majority of the entries. OpenJDK 11 appears in exactly two: Spark 3.1.1 for Hadoop 3.2, and Spark 3.0.0 for Hadoop 3.2. OpenJDK 7 appears once, in Spark 2.0.0 for Hadoop 2.7+ with Hive support. JDK 11 is the highest Java version in the list; nothing above it is included.

Hive support appears in only two entries, both for Spark 2.0.0, one with OpenJDK 8 and one with OpenJDK 7. No Spark 3.x build carries Hive support. What the Hive support configuration enables, and how to activate it, is not described in the README.

Spark 3.2.0 and 3.2.1 use Hadoop 3.2 in this image set. Spark 3.3.0 moved to Hadoop 3.3. If your existing environment runs Hadoop 3.2, upgrading to the Spark 3.3.0 image changes the Hadoop version at the same time.

## Application containers layer over bde2020/spark-base rather than bundling a full Spark distribution

Building and deploying your own Spark job to this cluster follows the template pattern. The repository holds templates for three build systems in the template/ directory: Maven, Python and Sbt. Completed examples for Maven and Python appear under examples/maven/ and examples/python/. Sbt has a template but no example.

An application container extends bde2020/spark-base rather than including a fresh Spark installation of its own. Spark binaries come from the base image, and the application adds its code on top. This keeps the Spark version consistent between the application container and the cluster images, which are all built from the same base. Packaging Spark separately inside the application image would produce a second Spark installation that might differ from the cluster's version.

One gap is documented by absence. The README points to each template's own README for the full build and run steps, but that document is inside the template/ directory rather than shown in the main README. Python is the more self-contained starting point, since the examples/python/ directory gives a working example to adapt.

## Kubernetes runs one manifest while Compose requires naming each worker service explicitly

Kubernetes deployment gives a different topology than Docker Compose. Compose requires you to define spark-worker-1 and spark-worker-2 as separate named services and choose their port mappings. Kubernetes reads one manifest and schedules a worker pod on every node in the cluster:

```bash
kubectl apply -f https://raw.githubusercontent.com/big-data-europe/docker-spark/master/k8s-spark-cluster.yaml
```

After applying, the master is reachable at spark://spark-master:7077 within the default namespace. An interactive spark-shell connects to the cluster through a short-lived pod with the client label:

```bash
kubectl run spark-base --rm -it --labels="app=spark-client" --image bde2020/spark-base:3.3.0-hadoop3.3 -- bash ./spark/bin/spark-shell --master spark://spark-master:7077 --conf spark.driver.host=spark-client
```

For spark-submit, the README shows the same kubectl run pattern with --class naming the application class and the application URL passed as the final argument. The --deploy-mode client and --conf spark.driver.host=spark-client flags appear in both commands. The --conf spark.driver.host=spark-client flag is required because workers need to route responses back to the driver. The label app=spark-client on the pod matches the headless service named spark-client, which the manifest creates so that hostname is resolvable from the worker pods. The README notes that custom application images must be reachable from the workers as well, and creating a headless service for the pod with the matching --conf spark.driver.host value is the documented way to achieve this.

## No GitHub releases, an unlisted licence and a last push on 2026-04-20

The repository carries no GitHub releases. There is no tagged version to pin in a dependency file, no release notes documenting what changed between image builds, and no semantic version number to quote in a bug report. Tracking a specific build means pinning the Docker image tag directly, which the README shows as 3.3.0-hadoop3.3 for the current image set.

Licence information is absent. The repository metadata does not list a licence identifier, and the top-level entries (.github/, .gitignore, README.md, base/, build.sh, docker-compose.yml, examples/, history-server/, k8s-spark-cluster.yaml, master/, submit/, template/, worker/) include no LICENCE or LICENSE file. Organisations that require licence review before adopting open source images have no declared licence to review here.

The last push to the repository was on 2026-04-20. Spark 3.3.0 is the current ceiling in the supported version list. Whether newer Spark versions will be added is not stated, and the gap between the latest available image and current Spark releases grows over time.

## Conclusion

Use docker-spark if you need a standalone Spark cluster on Docker Compose for local development or a pipeline, and Spark 3.3.0 is recent enough for your workload. It handles Kubernetes too, with one manifest putting a worker on every available node. Skip it if your organisation requires a declared licence before adopting open source code, because the licence field is unknown and no LICENCE file appears in the top-level entries. Before committing, apply the Kubernetes manifest in a test namespace and run the spark-shell command to confirm the driver-to-worker path resolves at spark://spark-master:7077.

## FAQ

### How do I start a docker-spark cluster with Docker Compose?

The docker-compose.yml in the repository defines spark-master on ports 8080 and 7077, spark-worker-1 on 8081, spark-worker-2 on 8082 and spark-history-server on 18081. All four services use bde2020 images tagged 3.3.0-hadoop3.3. Before starting, create /tmp/spark-events-local on the host; if it does not exist, the history server starts with no completed jobs to show.

### What Spark versions does docker-spark support?

Supported versions span from Spark 1.5.1 for Hadoop 2.6 up to Spark 3.3.0 for Hadoop 3.3. Almost all builds use OpenJDK 8; OpenJDK 11 appears only in the Spark 3.1.1 and Spark 3.0.0 builds for Hadoop 3.2, and OpenJDK 7 in one Spark 2.0.0 entry with Hive support.

### Can docker-spark be deployed on Kubernetes?

Yes. One kubectl apply command reads k8s-spark-cluster.yaml from the repository and sets up a master plus one worker pod on every available node. The master is reachable within the default namespace at spark://spark-master:7077.

### What is the INIT_DAEMON_STEP environment variable in docker-spark?

It is a pipeline step marker for the BDE (Big Data Europe) pipeline and is not a standard Spark configuration key. The README says to fill it in as configured in your BDE pipeline. Outside that pipeline context, the README gives no guidance on what value to use.

### How do I run my own Spark application on a docker-spark cluster?

The project provides templates for Maven, Python and Sbt under the template/ directory. Application containers extend bde2020/spark-base rather than packaging their own Spark installation, and the examples/maven/ and examples/python/ directories hold completed examples to adapt.

## Sources

- [big-data-europe/docker-spark on GitHub](https://github.com/big-data-europe/docker-spark)
- [Issues](https://github.com/big-data-europe/docker-spark/issues)
- [README](https://github.com/big-data-europe/docker-spark/blob/master/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/big-data-europe-docker-spark
