Self-hosted service
apache/singa avatar
apache/singa

Apache SINGA: a distributed deep learning platform in C++ for cluster training

a distributed deep learning platform

3,606 stars1,265 forksC++Apache-2.0

At a glance

What is it?
Apache SINGA is an Apache-licensed distributed deep learning system written mainly in C++, with Python and Java bindings and a set of example programs. The design question worth asking is whether its cluster-first training model fits your stack, and the release history suggests you should check the build path before committing.
Who is it for?
Adopt Apache SINGA if you need a C++-first distributed training stack you can build from source, and you are willing to pin to the 3.0.0 release line and the documented installation path. Do not adopt it if you expect a fast-moving framework with frequent releases or a large third-party model zoo; the latest release is 3.0.0 from 2020-04-21, so plan for that.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 28 days ago.
What is it written in?
Mainly C++, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What Apache SINGA is for, and who it is actually aimed at

Apache SINGA is described in its README as a distributed deep learning system, and the repository is hosted under the Apache Software Foundation with the Apache-2.0 licence. The primary language is C++, and the top-level tree contains include/, src/, python/, java/, examples/, doc/ and a rafiki entry alongside build files such as CMakeLists.txt and setup.py. That layout says a lot about the intended user. This is not a notebook-first framework. It is a system you build, link against, or wrap, and the Python and Java directories exist as bindings over a native core rather than as the core itself.

The audience follows from that. If your training job already lives in a Python-only world and you never touch a compiler, SINGA is probably not the shortest path. If you have a cluster, a C++ codebase, or a requirement to run training inside an existing native service, the shape of the project makes more sense. The examples directory is the clearest signal of scope: cnn, rnn, rbm, gan, mlp, onnx, cifar_distributed_cnn, largedataset_cnn, model_selection, model_slicing_psql, hfl, singa_peft and trans are all present as separate example programs. That range covers classic supervised models, a distributed CNN variant, and more recent-looking entries such as PEFT and transformer examples.

The name is also a practical problem. Searching for it returns results about a courtesy lion mascot and unrelated businesses, and the related searches for this project are dominated by those meanings rather than by the software. When you look for help, search for the full string apache singa, not the short name.

How the distributed training mechanism is laid out in the repository

The mechanism is visible in the directory structure rather than in a single architecture document. There is a native C++ core under src/ and include/, a Python binding layer under python/, and a Java binding layer under java/. The build is driven by CMake, with a cmake/ directory of modules and a top-level CMakeLists.txt. Python packaging goes through setup.py.

The distributed part shows up in the examples. There is a cifar_distributed_cnn example sitting next to a plain cnn example, which tells you the project treats single-node and multi-node training as separate reference programs rather than one program with a flag. There is also a model_slicing_psql example, where the name suggests model slicing combined with a PostgreSQL store, and an hfl example, which in the deep learning context normally means hierarchical or hybrid federated learning. The README itself does not spell out the communication backend or the parameter synchronisation strategy; for that you have to read the source under src/ and the documentation site at singa.apache.org.

That is a real cost. A distributed training system lives or dies on how gradients move between workers, and the README gives you a link to the docs and a link to the examples, nothing more. If you need to know whether the system does synchronous parameter averaging, asynchronous updates, or something else, the repository will answer it in code, not in prose. Treat that as a due diligence step, not a detail.

Installing Apache SINGA and running a first example

The README does not carry install instructions inline. It points to two places: the installation page at singa.apache.org/docs/installation/ and the examples directory in the repository. That is the documented path, and it is the one to follow, because the build has native dependencies and the exact package names matter.

There is also a published Docker image. The README carries a Docker pulls badge pointing at the apache/singa repository on Docker Hub, so pulling that image is a documented route to a working environment. The exact tag is not given in the README, so check the Docker Hub tags page rather than guessing one.

For a pip install, the setup.py file documents how wheels for SINGA are produced. It states that the script must be launched from the root directory of the project inside a container created from tool/docker/devel/centos/cudaxx/Dockerfile.manylinux2014, and gives this sequence:

bash
nvidia-docker run -v <local singa dir>:/root/singa -it apache/singa:manylinux2014-cuda10.2
/opt/python/cp36-cp36m/bin/python setup.py bdist_wheel
/opt/python/cp37-cp37m/bin/python setup.py bdist_wheel
/opt/python/cp38-cp38/bin/python setup.py bdist_wheel

What you should see is a wheel file produced per Python version. The same file notes that the generated wheel must be repaired with the auditwheel tool to comply with PEP513, otherwise dependent libraries are not bundled and the upload is rejected by PyPI on a file name error. The repair step is exposed as a setup.py subcommand:

bash
/opt/python/cp36-cp36m/bin/python setup.py audit

That matters if you are building wheels for your own deployment rather than installing a released one. For a first real use, the examples directory is the intended entry point. Each example is a directory with its own CMakeLists.txt or Python entry point, so the sequence is: pick the example closest to your model (cnn for a plain convolutional network, mlp for a baseline), read that directory, and follow whatever run instructions it contains. The README does not describe a single uniform run command across all examples, so do not assume one.

Release cadence and the version you will actually be running

The release list is short and the gap is wide. Apache SINGA 3.0.0 was released on 2020-04-21, after a 3.0.0.rc1 on 2020-04-08. Before that, 2.0.0 landed on 2019-04-07. The last push to the default branch was on 2026-09-02, so the repository is being touched, but the most recent tagged release is still 3.0.0 from 2020.

That combination is unusual and worth stating plainly. Code is moving on master while the release line has not advanced in years. If you install from a release artifact, you are installing a 2020 codebase. If you build from master, you are building something that has no matching release tag and no changelog describing what changed since 3.0.0. Neither is wrong, but they are different commitments, and the README does not tell you which one the documentation site describes.

The examples that look newest, singa_peft and trans, are the ones most likely to depend on master-only code. If you want PEFT or transformer examples, check whether the example references APIs that exist in the 3.0.0 release before you plan around it.

Where Apache SINGA is the wrong tool

The clearest failure mode is expecting the ecosystem conveniences of a mainstream framework. There is no indication in the README of a hosted model hub, a pretrained checkpoint registry, or a large set of third-party tutorials. The examples directory is the model zoo, and it is a set of reference programs, not a catalogue of downloadable weights.

The second limitation is build friction. The setup.py file describes a manylinux2014 container, CUDA 10.2, and per-Python-version wheel builds with an auditwheel repair step. That is a real toolchain, not a pip install and go. If your team does not already build native extensions, the installation page is going to be the hardest part of the project, and the README does not offer a simpler fallback beyond the Docker image.

The third is documentation depth on the distributed path. The README links to the docs site and to the examples, and stops there. For a system whose entire premise is distributed training, the README does not document the communication layer, the failure behaviour of a worker, or how to resume a job. If you need those answers before you can adopt anything, get them from the source tree first.

Finally, if your workload is a single GPU and a standard model, the distributed machinery is overhead. A plain cnn or mlp example does not need a cluster, and the build cost is the same either way.

How it differs from PyTorch and TensorFlow in approach

The honest comparison is with PyTorch and TensorFlow, because those are what most teams evaluate against. The difference is where the centre of gravity sits. PyTorch and TensorFlow put the Python API at the centre and treat the native runtime as an implementation detail. SINGA inverts that: the core is C++ under src/ and include/, and python/ and java/ are binding layers over it. That is why the repository contains a java/ directory at all, which is not something you find in a Python-first framework.

The second difference is distribution as a first-class example rather than a wrapper. In PyTorch, distributed training is a module you import and configure. In SINGA, cifar_distributed_cnn exists as its own example directory next to the single-node cifar example. The project's own materials present distribution as a mode you pick a program for.

The third difference is ecosystem weight. PyTorch and TensorFlow have years of releases, third-party packages and pretrained models. SINGA's last release is 3.0.0 from 2020-04-21. That is not a reason to dismiss it if you specifically want a C++-native, Apache-governed training system, but it is the reason a general-purpose team should not pick it. If you want the largest pool of pretrained models and community answers, the mainstream frameworks win on that axis and SINGA does not compete there.

Licence, governance and the cost of upgrading

SINGA is licensed under Apache-2.0, and the repository carries the standard Apache files: LICENSE, NOTICE, DISCLAIMER, KEYS and SECURITY.md. It is an Apache Software Foundation project, with development and commits mailing lists listed in the README and JIRA used for issue tracking rather than GitHub issues. That governance model is a genuine difference from a single-vendor project: contributions go through the ASF process, and the project has a documented security reporting path in SECURITY.md.

For licence implications, Apache-2.0 is a permissive licence with an explicit patent grant and a NOTICE file requirement. If you redistribute SINGA or a derivative, the NOTICE file travels with it. This article is not legal advice; read LICENSE and NOTICE in the repository and talk to your own counsel if you are embedding it in a product.

The upgrade cost is the part to think about. With 3.0.0 as the newest release tag and master still receiving pushes, the distance between a released artifact and the current source is unknown and undocumented in the README. There is no migration guide referenced. If you build against master today and a release appears later, you have no changelog to diff against. The practical approach is to pin to a specific commit, record it, and rebuild deliberately rather than tracking master.

Editorial conclusion

Adopt Apache SINGA if you need a C++-first distributed training stack you can build from source, and you are willing to pin to the 3.0.0 release line and the documented installation path. Do not adopt it if you expect a fast-moving framework with frequent releases or a large third-party model zoo; the latest release is 3.0.0 from 2020-04-21, so plan for that. Before writing any code, verify that the installation page for your platform produces a working build, then run one example from the examples directory end to end on your actual hardware.

Frequently asked questions

Is Apache SINGA free to use?

Yes. The repository is licensed under Apache-2.0, which is a permissive open source licence, and the LICENSE and NOTICE files are present at the top level of the repository.

What is Apache SINGA used for?

The README describes it as a distributed deep learning system. The examples directory covers convolutional networks, recurrent networks, RBMs, GANs, MLPs, ONNX, a distributed CNN, and PEFT and transformer examples.

How do I install Apache SINGA?

The README points to the installation page at singa.apache.org/docs/installation/ and to the examples directory. There is also a published Docker image at apache/singa on Docker Hub; the README does not give a specific tag, so check the Docker Hub tags page.

What is the latest release of Apache SINGA?

The most recent release listed is 3.0.0, published on 2020-04-21, following a 3.0.0.rc1 on 2020-04-08. The previous release was 2.0.0 on 2019-04-07.

Which programming languages can I use with Apache SINGA?

The primary language of the repository is C++, and the top-level tree includes both a python/ directory and a java/ directory, so Python and Java bindings are part of the project layout.

Official sources

  1. apache/singa on GitHub
  2. Issues
  3. License: Apache-2.0
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/apache-singa.svg)](https://hysenlabs.com/projects/apache-singa)