Open-source project
databricks/learning-spark avatar
databricks/learning-spark

learning-spark: Example Code from the O'Reilly Learning Spark Book

Example code from Learning Spark book

3,896 stars2,389 forksJavaMIT

At a glance

What is it?
databricks/learning-spark is the companion code repository for the O'Reilly book Learning Spark. It contains Java, Scala, and Python examples updated to run against Spark 1.3, along with build files for SBT and Maven.
Who is it for?
databricks/learning-spark is useful for readers of the O'Reilly Learning Spark book who want to run the examples locally. It is the wrong starting point for new Spark development: the code targets Spark 1.3 and Scala 2.10, which are several major versions behind current Spark releases.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 92 days ago.
What is it written in?
Mainly Java, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What This Repository Contains

databricks/learning-spark is the example code repository for the Learning Spark book published by O'Reilly Media. The book covers Apache Spark, the distributed data processing framework. The repository holds working code examples for the topics covered in each chapter, written in Java, Scala, and Python.

The README notes that the examples have been updated to run against Spark 1.3, which means they may differ slightly from the printed versions in some copies of the book. The last push was on 2026-06-30. The repository is licensed under MIT.

The book this code accompanies is Learning Spark by Holden Karau, Andy Konwinski, Patrick Wendell, and Matei Zaharia. The repository was maintained by Holden Karau, whose Travis CI link appears in the README.

Requirements and Dependencies

The README lists the following requirements for building and running the examples:

- JDK 1.7 or higher - Scala 2.10.3 - Spark 1.3 - A Protobuf compiler (on Debian: `sudo apt-get install protobuf-compiler`) - R and the CRAN package Imap for the Chapter Six example - Python's urllib3 for the Python examples

These requirements reflect the state of the Spark ecosystem circa 2014-2015. Spark 1.3 predates the DataFrame API becoming stable, predates Spark SQL's full integration, and predates the Dataset API introduced in Spark 1.6. The Scala version 2.10.3 is also no longer supported by current Spark releases, which require Scala 2.12 or 2.13.

Attempting to run these examples against a current Spark installation will likely require code modifications.

Building with SBT or Maven

The repository supports two build systems. SBT is configured through `build.sbt` and the `sbt/` wrapper directory. Maven is configured through `pom.xml`. Both produce an assembly JAR containing the examples and all their dependencies, which can then be submitted to a Spark cluster.

To build with SBT and then submit a job:

code
./sbt/sbt assembly

To build with Maven:

code
mvn package

After building, jobs are run with the Spark submit script:

code
cd $SPARK_HOME; ./bin/spark-submit --class com.oreilly.learningsparkexamples.[lang].[example] ../learning-spark-examples/target/scala-2.10/learning-spark-examples-assembly-0.0.1.jar

Replace `[lang]` and `[example]` with the appropriate language and class name. The assembly JAR path includes `scala-2.10`, confirming the Scala version requirement. The `build-project`, `setup-project`, and `run-all-examples` scripts in the repository root automate parts of this workflow.

Running Python Examples

The Python examples live under `src/python/`. They are run directly with PySpark rather than through an assembly JAR:

code
./bin/pyspark ./src/python/[example]

This command is run from within the `$SPARK_HOME` directory. Each Python example is a standalone script that creates a SparkContext and performs operations. The README lists Python urllib3 as a requirement for the Python examples.

The Python examples cover the same conceptual territory as the Java and Scala examples but use the PySpark API. In the 2014-2015 era, PySpark had a smaller feature set compared to the Scala API, so some capabilities demonstrated in Java or Scala were not available through Python.

The mini-complete-example Directory

The README highlights the `mini-complete-example/` directory as a standalone example with minimal dependencies and a smaller build file. The main examples directory has long build files because the full suite of examples requires many libraries. The mini example is aimed at developers who want a simple starting point without the overhead of the complete build.

The `mini-complete-example/` is the recommended entry point for verifying that your Spark setup can compile and run these examples before tackling the full build. It reduces the number of things that can go wrong during initial setup.

The `DESCRIPTION` file at the repository root and the `files/` directory contain supporting resources referenced by specific examples.

Limitations: Spark 1.3 and a Decade of API Changes

The most significant limitation of this repository is its age relative to current Spark. Spark 1.3 was released in 2015. The current Spark major version is 4.x, which includes the fully mature DataFrame and Dataset APIs, Spark SQL improvements, structured streaming, and a significantly different configuration model.

Code written for Spark 1.3 uses the RDD API as the primary abstraction. While RDDs still exist in current Spark, the recommended programming model has shifted to DataFrames and Datasets, which offer better performance through the Catalyst query optimizer and Tungsten execution engine. Learning RDD patterns from this repository and then adapting to current Spark requires understanding what changed between Spark 1.x and the current release. The API surface for transformations and actions is similar in intent, but the execution model and optimization semantics differ significantly between the two.

The Scala 2.10.3 requirement is also a barrier. Current Spark requires Scala 2.12 or 2.13. The SBT build files in this repository cannot be used directly with a current Spark or Scala setup.

Databricks publishes official example notebooks and documentation for current Spark versions. The Apache Spark project also maintains example code in the Spark source repository under `examples/src/main/`. Both are maintained against current APIs and are more appropriate references for new development than this repository.

Editorial conclusion

databricks/learning-spark is useful for readers of the O'Reilly Learning Spark book who want to run the examples locally. It is the wrong starting point for new Spark development: the code targets Spark 1.3 and Scala 2.10, which are several major versions behind current Spark releases. Developers learning Spark for current use should work from official Spark documentation and examples rather than this 2014-era companion code.

Frequently asked questions

Which Spark version does the learning-spark repository target?

The README states the examples have been updated to run against Spark 1.3. This is an older version; current Spark releases are in the 4.x range. The code uses RDDs as the primary API rather than the DataFrames and Datasets that are standard in current Spark.

Can I run the learning-spark examples with a current version of Spark?

The examples target Spark 1.3 and Scala 2.10.3. Current Spark requires Scala 2.12 or 2.13 and uses a different build model. Running these examples against a current Spark installation would require significant code and build file modifications.

Is there a simpler starting point than the full example suite?

Yes. The README highlights the mini-complete-example/ directory as a standalone example with minimal dependencies and a smaller build file. It is the recommended entry point for verifying a basic Spark setup before working through the full example collection.

Official sources

  1. databricks/learning-spark on GitHub
  2. Issues
  3. License: MIT
  4. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/databricks-learning-spark.svg)](https://hysenlabs.com/projects/databricks-learning-spark)