Scala Machine Learning Projects: Packt's Eleven-Project Code Repository
Scala Machine Learning Projects, published by Packt
At a glance
- What is it?
- This repository holds the Scala source code for eleven end-to-end machine learning projects that accompany the Packt book 'Scala Machine Learning Projects,' using Spark ML, H2O, DeepLearning4j, and MXNet. The last commit was in July 2023, and the dependency stack includes Apache Spark, Hadoop, and specific Scala versions that require careful environment setup.
- Who is it for?
- This repository suits a data scientist or engineer who has purchased 'Scala Machine Learning Projects' and wants to run the eleven examples in a local Scala and Spark environment. It is not a standalone reference: the code is unexplained outside the book.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Probably not. The repository last received commits 39 months ago, on July 16, 2023.
- What is it written in?
- Mainly Scala, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 28, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What the Repository Contains and Who Should Use It
PacktPublishing/Scala-Machine-Learning-Projects is the official code repository for the Packt book 'Scala Machine Learning Projects.' The README describes it as containing all supporting project files necessary to work through the book from start to finish.
The intended audience is data scientists, data engineers, and deep learning practitioners who have a solid background in Scala and its functional programming concepts. The README also notes that some prior familiarity with Spark ML, H2O, Zeppelin, DeepLearning4j, and MXNet would be an advantage, though not strictly required. Developers who have never used Scala will find the code hard to follow without both the book and independent Scala learning resources.
The repository contains chapters 01 through 11, covering all eleven projects described in the book.
Eleven Projects Across Spark ML, H2O, DeepLearning4j, and MXNet
The README states the book covers eleven end-to-end projects using four major machine learning libraries. Spark ML is Apache Spark's built-in machine learning pipeline library. H2O provides distributed machine learning algorithms that can run alongside Spark. DeepLearning4j is a JVM-native deep learning library. MXNet supports both symbolic and imperative deep learning workloads.
The README also references MDP (Markov Decision Processes) and Zeppelin, indicating that at least some chapters involve reinforcement learning or interactive notebook-style computation. The exact project topics are not listed in the repository; the book provides the descriptions.
Each chapter folder contains a Maven project with its own pom.xml. Maven handles dependency resolution and compilation, so there is no single root build file for the whole repository.
System Requirements: A Stack That Takes Time to Assemble
The README lists the required components explicitly. Setting up a working environment requires all of the following:
Apache Spark 2.0.0 or higher, Hadoop 2.7 or higher, Java JDK and JRE version 1.7 or 1.8, Scala 2.11.x or higher, MXNet, Zeppelin, DeepLearning4j, and H2O with details in the chapter files and supplied pom.xml files. The IDE setup requires Eclipse Mars or Luna with the Maven plugin version 2.9 or higher, the Maven compiler plugin version 2.3.2 or higher, and the Maven assembly plugin version 2.4.1 or higher, or IntelliJ IDE with the SBT plugin and Scala Play Framework installed.
This is a substantial stack. Spark 2.0.0 was released in 2016, and current Spark releases are in the 3.x series. Running the examples with current dependency versions may require updating pom.xml files and resolving API changes. The README does not document how to do this.
A Scala Code Example from the Repository
The README includes a Scala snippet that illustrates the coding style used throughout the projects. The example constructs a unique identifier string from a genomic genotype object:
def variantId(genotype: Genotype): String = {
val name = genotype.getVariant.getContigName
val start = genotype.getVariant.getStart
val end = genotype.getVariant.getEnd
s"$name:$start:$end"
}The function calls `getContigName`, `getStart`, and `getEnd` on a `Variant` object accessed through the `Genotype` type. This is an example of one of the eleven projects, likely from a genomics or bioinformatics chapter. The code uses Scala's string interpolation syntax (`s"..."`), a basic feature of Scala 2.11.
This snippet is a sample of what the projects look like, not a runnable entry point. The full project for each chapter is a Maven application with its own main class and dependencies.
What the Repository Does Not Provide
The README does not include per-chapter descriptions of what each project does or which dataset it uses. There are no README files inside the individual chapter folders visible from the repository listing.
There are no GitHub releases and no version tags. The repository has no continuous integration configuration, meaning there is no external record of whether the examples ever compiled successfully against the declared dependency versions.
The README lists four related Packt products but gives no troubleshooting guidance. Developers who hit dependency conflicts or API changes will need to resolve them without guidance from the repository.
Maintenance Status: Last Push July 2023
The last push to this repository was on 2023-07-16. The repository is open and not archived, but it has received no updates since then. Three years into a Spark ecosystem that moved from 2.x to 3.x with breaking API changes, the example code targets versions that may not install cleanly in an environment built around current tooling.
The MIT licence allows readers to fork, modify, and redistribute the code. Engineers who update the pom.xml files and resolve API incompatibilities are legally free to share their updated versions, but the original repository shows no record of such contributions being merged.
The Packt website offers a DRM-free PDF to purchasers of the print or Kindle edition. That download also does not include dependency-updated code.
How This Compares to Python-Based Machine Learning Courses
Python dominates the machine learning ecosystem. Libraries such as scikit-learn, PyTorch, and TensorFlow have larger community support, more recent documentation, and lower setup friction than the JVM stack this repository uses. A developer starting from scratch will find it easier to get a first model running in Python than to assemble the Spark, Hadoop, Scala, and Maven environment described in the README.
The reason to use this Scala repository is integration with existing JVM infrastructure. Teams that run Spark on a JVM-heavy data platform, or that deploy machine learning models inside Java services, have a genuine reason to work through examples in Scala rather than Python. DeepLearning4j and Spark ML are production tools for those environments. The book's eleven projects cover that production-oriented Scala path, which Python tutorials do not.
Editorial conclusion
This repository suits a data scientist or engineer who has purchased 'Scala Machine Learning Projects' and wants to run the eleven examples in a local Scala and Spark environment. It is not a standalone reference: the code is unexplained outside the book. The last push was on 2023-07-16, and the dependency versions listed (Spark 2.0.0, Scala 2.11.x, Hadoop 2.7) are several major releases behind current. Before using the examples, verify that the pom.xml versions in each chapter folder compile against the Spark and Scala releases in your environment.
Frequently asked questions
What build tool does Scala Machine Learning Projects use for its examples?
Each chapter folder contains a Maven project with its own pom.xml file. The README lists Maven plugins as required components and does not mention SBT for the examples, though IntelliJ with the SBT plugin is listed as an IDE option.
What Scala version is required for the Scala Machine Learning Projects examples?
The README specifies Scala 2.11.x or higher. All examples in the repository were developed on Ubuntu 16.04 LTS 64-bit and Windows 10 64-bit, according to the README.
Does the Scala Machine Learning Projects repository cover all eleven chapters of the book?
Yes. The repository contains Chapter01 through Chapter11, corresponding to all eleven projects described in the book. The README states the repository contains all supporting project files necessary to work through the book from start to finish.