Library / SDK
deepspeedai/DeepSpeed avatar
deepspeedai/DeepSpeed

DeepSpeed: ZeRO under the hood, and what a 0.x patch release can break

DeepSpeed is a deep learning optimization library that makes distributed training and inference easy, efficient, and effective.

43,163 stars5,016 forksPythonApache-2.0

At a glance

What is it?
DeepSpeed is an Apache-2.0 Python library for distributed training and inference, best known for the ZeRO family. Looking at the build scripts and release history rather than the feature list, it is a research project with a research project's release discipline.
Who is it for?
DeepSpeed fits teams that have decided to build on ZeRO, Ulysses sequence parallelism or the MoE path and can absorb a library that ships breaking API changes inside a 0.x version. It does not fit anyone who needs a stable surface across patch releases, a one-command install on Windows, or a stated matrix of supported accelerators and CUDA versions, because none of that is written down.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 1 day ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 28, 2026, and from our analysis. They are not legal advice.

Editorial analysis

ZeRO is the part everyone meets first, and the rest is research

The repository description calls DeepSpeed a deep learning optimization library that makes distributed training and inference easy, efficient and effective. The system innovations it names are ZeRO, ZeRO-Infinity, 3D-Parallelism, Ulysses Sequence Parallelism and DeepSpeed-MoE, and the claim attached to them is that the library enabled training of MT-530B and BLOOM. The adoption list is concrete about scale: Megatron-Turing NLG at 530B, Jurassic-1 at 178B, BLOOM at 176B, GLM at 130B, xTrimoPGLM and YaLM at 100B, GPT-NeoX and AlexaTM at 20B, Turing NLG at 17B and METRO-LM at 5.4B. Integrations exist for the usual suspects, with entries for Transformers with DeepSpeed and Accelerate. Read that list as a resume of who has run the code at scale, not as a compatibility matrix for what your hardware will do.

setup.py turns a missing torch into a warning, not an error

The installer script wraps its torch import in a try block and degrades quietly. If the import fails it sets torch_available to False and prints:

code
[WARNING] Unable to import torch, pre-compiling ops will be disabled. Please visit https://pytorch.org/ to see how to properly install torch on your system.

Installation continues. What you lose is pre-compilation of the custom ops, which the script enumerates through op_builder: get_default_compute_capabilities, the OpBuilder base class, ALL_OPS and accelerator_name, plus accelerator.get_accelerator to pick a backend. The script also asks OpBuilder.is_rocm_pytorch() whether it is building against ROCm. A related trap is the local abort helper, which prints a red [ERROR] line and then raises through a bare assert False, so any build failure surfaces as an assertion rather than a clean exception. Installing into an environment without torch gives you a package that installs successfully and then behaves differently at runtime.

There is no Windows wheel for you, only a batch file

Windows support is documented as a procedure rather than a download. The header of setup.py spells out the four steps for building wheels on Windows: install pytorch, such as pytorch 2.3 + cuda 12.1, install the Visual C++ build tool, include the CUDA toolkit, and launch a cmd console with Administrator privilege so the required symlink folders can be created. The build itself is then invoked as a batch file:

code
build_win.bat

The result lands in dist/*.whl, which you install yourself. In other words the Windows path assumes you already have a matching PyTorch and CUDA stack, a compiler, and permission to create symlinks, and it asks you to assemble all of that before the package manager will help. A separate MANIFEST_win.in sits alongside the ordinary MANIFEST.in, so the Windows wheel genuinely is a distinct build. If your workstation is Windows and your cluster is Linux, this is the step where the mismatch surfaces.

make test forks every unit test, and make format diffs against master

The Makefile sets .DEFAULT_GOAL to help, so a bare make prints the target list generated by awk from the comment annotations. There are two real targets. The test target runs:

code
pytest --forked tests/unit/

The --forked flag runs each test in its own process, which keeps one test's global state from poisoning the next, at the cost of a process launch per test. The format target is more opinionated: if a venv directory is missing it creates one, installs pre-commit with -U, then runs pre-commit clean, uninstall and install to rebuild the hook environment. It finishes with:

code
pre-commit run --files $(git diff --name-only master)

That means lint runs against files that differ from the master branch, not against staged or committed work. Contribute from a branch and it works as intended. Run it on a detached checkout and the file list is computed against a branch you may not have.

Patch releases land every few weeks and can carry API changes

Three of the most recent tags are v0.19.7 on 2026-09-16, v0.19.6 on 2026-08-27 and v0.19.5 on 2026-08-10, and the master branch was pushed on 2026-09-28. The version number is tracked in its own version.txt at the repository root, alongside release/ and ci/ directories. The reason to care about the cadence is that the project is comfortable changing its API inside the 0.x line. A December 2025 post in the news section is titled DeepSpeed Core API updates: PyTorch-style backward and low-precision master states, which is an API shape change, not a bug fix. Treat a patch number as a moving target: pin the exact version, read the release notes for the jump you are making, and expect to touch call sites when you upgrade rather than only a requirements line.

A research lab is shipping inside the package

Look at the directory list and the shape of the project is clear. Alongside deepspeed/ and csrc/ there are blogs/, benchmarks/, examples/, azure/, docker/, op_builder/, accelerator/, requirements/, scripts/, tests/ and ci/, and examples/README.md sits next to a directory for the System DMA allgather work. The news section tracks active research threads rather than features: SuperOffload for superchips, ZenFlow as a stall-free offloading engine, a follow-up study on ZenFlow with CPU core binding, DeepNVMe for I/O scaling, DeepCompile, AutoTP for automatic tensor parallelism, Ulysses-Offload, Arctic Long Sequence Training, Muon optimizer support and SDMA collectives for ZeRO-3 on AMD GPUs. Governance is written down too, with GOVERNANCE.md, CODEOWNERS, CODE_OF_CONDUCT.md, CONTRIBUTING.md, THIRD_PARTY_NOTICES.md and a SECURITY.md. The practical consequence is that features you do not use arrive in the same release train as experiments you do.

Support for AMD and ROCm is real but assembled from separate pieces

There is no support matrix in the repository, so you have to assemble the picture. Setup asks OpBuilder.is_rocm_pytorch() at build time, which means a ROCm path exists. The accelerator package exposes get_accelerator so the backend is chosen rather than hardcoded, and op_builder.get_default_compute_capabilities produces the list of GPU architectures the custom ops are compiled for, which is why a new card generation means a rebuild. AMD shows up in the research track too, with the May 2026 post on System DMA for ZeRO-3 that offloads collectives off the compute units on AMD GPUs. What you cannot read anywhere is which accelerator versions are supported per release, or what happens on a mixed-architecture node. On a new GPU, budget for compiling ops yourself and test the collective path before you schedule a long run.

What the repository does not promise, and who that hurts

Several things a library consumer would look for are simply absent, and each absence has a cost. No minimum hardware or GPU memory requirement is stated for any of the ZeRO stages, so a planner sizing a run has to derive it. No list of supported CUDA, ROCm or PyTorch versions appears beyond the single worked example in the Windows build notes, and no compatibility table maps DeepSpeed versions to frameworks. No backward compatibility guarantee is written down for the 0.x line, and the December 2025 core API post shows the surface does move. The consequence lands hardest on the person who inherits the training job: they get a pinned version, a cluster configuration and a knowledge of which parts of the recipe are load-bearing, and the repository will not tell them whether a bump is safe. Read the news posts as change notes, and keep your own record of what you pinned and why.

Editorial conclusion

DeepSpeed fits teams that have decided to build on ZeRO, Ulysses sequence parallelism or the MoE path and can absorb a library that ships breaking API changes inside a 0.x version. It does not fit anyone who needs a stable surface across patch releases, a one-command install on Windows, or a stated matrix of supported accelerators and CUDA versions, because none of that is written down. Before you commit, pin an exact release, build the Windows wheel yourself if that is your platform, and read the core API update note, since the PyTorch-style backward change landed after years of stable calling conventions.

Frequently asked questions

What is DeepSpeed used for?

Distributed training and inference of deep learning models. The library names ZeRO, ZeRO-Infinity, 3D-Parallelism, Ulysses Sequence Parallelism and DeepSpeed-MoE among its innovations, and lists models from METRO-LM at 5.4B up to Megatron-Turing NLG at 530B as trained with it.

Is DeepSpeed developed by Microsoft?

The source files carry a Microsoft Corporation copyright header and the DeepSpeed Team name, and the project is described as having been an important part of the AI at Scale initiative. The repository itself now lives under the deepspeedai GitHub organisation.

What is DeepSpeed?

A deep learning optimization library for distributed training and inference, written in Python and released under Apache-2.0. It integrates with existing frameworks, with documented paths for Transformers with DeepSpeed and for Accelerate.

What is DeepSpeed ZeRO?

ZeRO is one of the system innovations DeepSpeed is built around, alongside ZeRO-Infinity, 3D-Parallelism, Ulysses Sequence Parallelism and DeepSpeed-MoE. Separate posts cover extensions of it, such as SDMA collectives for ZeRO-3 on AMD GPUs and ZeRO++ used at LinkedIn for recommendation model distillation.

How do you install DeepSpeed on Windows?

You build the wheel yourself. setup.py instructs you to install pytorch, such as pytorch 2.3 + cuda 12.1, install the Visual C++ build tool, include the CUDA toolkit, and run from a cmd console with Administrator privilege to create the symlink folders, then run build_win.bat and take the result from dist/*.whl.

Official sources

  1. Official documentation
  2. Official README
  3. Project repository
  4. Release notes
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/deepspeedai-deepspeed.svg)](https://hysenlabs.com/projects/deepspeedai-deepspeed)