Angel-ML/angel: A Parameter Server for High-Dimensional Models on YARN
A Flexible and Powerful Parameter Server for large-scale machine learning
At a glance
- What is it?
- Angel is a Java and Scala parameter server platform from Tencent and Peking University, aimed at models whose parameter count outgrows a single machine. Here is what the repository documents, how to build it, and where it stops being the right tool.
- Who is it for?
- Adopt Angel if you already run Spark or YARN clusters and your models have grown past what a single driver can hold, and you want the parameter server to be the storage layer rather than a Python process. Do not adopt it if you need a pip-installed framework, a managed service, or a project with a documented compatibility matrix; the README links to deployment docs but gives no supported-version table.
- Can I use it commercially?
- Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
- Is it still maintained?
- Yes. The repository last received commits 12 days ago.
- What is it written in?
- Mainly Java, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The problem Angel-ML/angel was built to solve
Wide models break the worker-centric pattern. When a logistic regression or factorization machine has hundreds of millions of features, the parameter vector no longer fits comfortably in the memory of the process that computes gradients, and shipping it between workers every iteration dominates the runtime. Angel takes the parameter server position on this: parameters are partitioned across dedicated server nodes, and workers hold only the slices they are updating. The README describes the design as model-centered, with parameters of complex models partitioned into multiple parameter-server nodes and a flexible consistency model for synchronization. The intended user is not a notebook user. It is an engineer with a Spark or YARN cluster who already trains on data too large for one machine and has hit the point where the model, not the data, is the bottleneck. Angel is written in Java and Scala, and the README states it supports running on YARN.
How the parameter server, PS Service and Spark on Angel fit together
The architecture separates three roles. Parameter server nodes hold the partitioned model. Workers compute gradients and push updates. A coordinator drives synchronization according to the consistency model you configure. Angel exposes this through a PS Service abstraction, and that abstraction is what makes Spark on Angel possible: a Spark job keeps its normal RDD and DataFrame execution, while the model lives in Angel servers reached through PS Service instead of being broadcast or repartitioned across executors. The README points to separate design documents for the Model Partitioner, the SyncController and psFunc, which is where the actual partitioning and update semantics are specified. Two things are worth flagging. First, the README says graph computing and deep learning framework support is under development, yet the algorithm index already lists graph algorithms such as PageRank, Louvain and LINE under Spark on Angel, and links several deep models to a separate PyTorch-On-Angel repository. That is a documentation inconsistency you should resolve by reading the linked pages rather than the summary paragraph. Second, the algorithm list is split between plain Angel and Spark on Angel, and the repository includes a document titled Angel or Spark On Angel to help you choose between them. That choice is not cosmetic; it determines which API and which deployment path you follow.
Building Angel-ML/angel and running a first job
The repository ships a Dockerfile that pins the toolchain, which is the least ambiguous installation path in the tree. The dev stage is based on maven:3.6.1-jdk-8 and installs curl, g++, make and unzip, then downloads and builds protobuf 2.5.0 from source before setting PATH and LD_LIBRARY_PATH. If you build outside Docker, those are the versions to match: JDK 8, Maven 3.6.1, protobuf 2.5.0. The java_builder stage then copies the tree and runs the package goal with tests skipped:
FROM maven:3.6.1-jdk-8 as dev
RUN apt-get update \
&& apt-get install -y --no-install-recommends \
curl=7.52.1-5+deb9u9 \
g++=4:6.3.0-4 \
make=4.1-9.1 \
unzip=6.0-21+deb9u1 \
&& rm -rf /var/lib/apt/lists/*
RUN curl -fsSL --insecure -o /tmp/protobuf-2.5.0.tar.gz https://github.com/protocolbuffers/protobuf/releases/download/v2.5.0/protobuf-2.5.0.tar.gz \
&& tar -xzf /tmp/protobuf-2.5.0.tar.gz -C /tmp \
&& cd /tmp/protobuf \
&& ./configure \
&& make -j4 \
&& make installThe equivalent step inside the builder stage is a single Maven invocation, which is what you would run locally after protobuf is on the path:
mvn -e -B -Dmaven.test.skip=true packageExpect the build to produce the assembly artifacts under the assembly module. From there the README sends you to the Compilation Guide, the Running on Local guide and the Running on Yarn guide for the actual launch, and to the Quick Start Example for a first Spark on Angel program. Those documents are the ones to follow for configuration keys and resource settings; the README itself does not list ports or environment variables, so do not guess at them. For algorithm selection, the model configuration document is the reference the README points to.
Where Angel-ML/angel is the wrong tool
The strongest limitation is the deployment surface. Angel is built for YARN, and the README names YARN as the supported runtime. If your infrastructure is Kubernetes-native with no YARN footprint, there is no documented path here, and the Dockerfile is a build image rather than a supported production container. The second limitation is the toolchain. Pinning protobuf 2.5.0 and JDK 8 in a 2026 build image means you are adopting a dependency set that most of your other services have moved past; that is a real integration cost, not a stylistic one. Third, the consistency model is described as flexible, but the README does not enumerate the available modes or their trade-offs, so you have to read the SyncController design document before you can reason about staleness in training. Fourth, the README's own summary of graph and deep learning support conflicts with its algorithm index, which means you should not rely on the summary when scoping a project. Finally, if your model fits in memory on one machine, or your team's working language is Python, the whole parameter server layer is overhead you are paying for nothing.
How Angel-ML/angel differs from Petuum and from plain Spark MLlib
Petuum is the closest architectural relative: another parameter server design, also from an academic and industrial collaboration, also aimed at scaling model parameters beyond a single machine. The difference in practice is the surrounding ecosystem rather than the idea. Angel's distinguishing move is the PS Service abstraction, which lets an existing Spark job keep Spark for data processing while the model sits in Angel servers. Plain Spark MLlib takes the opposite approach: it keeps the model in the driver or in broadcast variables and relies on Spark's own aggregation primitives. That works until the model is large enough that broadcasting it every iteration becomes the cost. Angel's answer is to never broadcast it at all. If you are already committed to Spark and your models are moderate, MLlib is less machinery to operate. If your feature space is genuinely high-dimensional, the parameter server split is the reason to look at Angel, and the Angel or Spark On Angel document is the repository's own guidance on which of Angel's two execution modes to pick.
Maintenance, licensing and what to check before upgrading
The last push to the default branch was on 2026-09-09, and the most recent release is Release-3.3.0 from 2025-09-29. Before that, Release-3.2.0 landed in August 2021 and v3.1.0 in May 2020, so the release cadence is uneven: a four-year gap sits between 3.2.0 and 3.3.0. Plan upgrades around that rhythm rather than assuming frequent point releases, and read the release notes for 3.3.0 before moving a cluster, because the README itself still advertises release 3.2.0 in its badge. On licensing, the README carries an Apache 2.0 badge and links to LICENSE.TXT, but the repository metadata reports the licence as NOASSERTION, meaning an automated classifier could not resolve it. Those two signals disagree, and the file in the tree is the one that governs; if the licence matters to your legal review, read LICENSE.TXT directly rather than trusting either badge or metadata. The project is under the Linux Foundation's deep learning umbrella per the community section, with a mailing list and a committer list in the tree.
Editorial conclusion
Adopt Angel if you already run Spark or YARN clusters and your models have grown past what a single driver can hold, and you want the parameter server to be the storage layer rather than a Python process. Do not adopt it if you need a pip-installed framework, a managed service, or a project with a documented compatibility matrix; the README links to deployment docs but gives no supported-version table. Verify first that you can build the tree against protobuf 2.5.0 and JDK 8, because the Dockerfile pins both and a mismatched toolchain is the first thing that fails.
Frequently asked questions
What is Angel-ML/angel used for?
It is a distributed machine learning and graph computing platform based on the parameter server model, written in Java and Scala and designed to run on YARN. It partitions the parameters of large models across parameter-server nodes and supports Spark on Angel through a PS Service abstraction.
How do I build Angel-ML/angel from source?
The repository Dockerfile builds protobuf 2.5.0 from source on top of maven:3.6.1-jdk-8, then runs the Maven package goal with tests skipped. Building outside Docker means matching JDK 8, Maven 3.6.1 and protobuf 2.5.0 yourself.
Does Angel-ML/angel run on Kubernetes instead of YARN?
The README states that Angel supports running on YARN, and the repository's Dockerfile is a multi-stage build image rather than a documented production deployment. No Kubernetes runtime is described.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/angel-ml-angel)