# Colossal-AI v0.5.0: the benchmark table compares hardware, not frameworks

> Colossal-AI is an Apache-2.0 Python project for making large model training cheaper and faster, with a benchmark table, a curated examples tree, and a setup.py that carries real constraints. Read closely, the benchmark compares NVIDIA B200 against H200 rather than Colossal-AI against anything else, the 70B rows contain no B200 at all, and the install path refuses Windows outright while keeping compiled extensions behind an environment variable.

**hpcaitech/ColossalAI** — Making large AI models cheaper, faster and more accessible. Build your AI agents, chatbots, and RAG applications with HPC-AI Model APIs!

- Repository: https://github.com/hpcaitech/ColossalAI
- Website: https://www.colossalai.org
- Stars: 41,443 · Forks: 4,499
- Language: Python
- License: Apache-2.0
- Published: 2026-08-04 · Updated: 2026-08-18 · Language: en
- Canonical page: https://hysenlabs.com/projects/hpcaitech-colossalai

## Every row in the table changes more than the GPU

The Colossal-AI benchmark is four rows wide, and the columns that matter move together. The 7B case on 8 cards runs zero2(dp8) with a batch size of 36 per data parallel rank and reports 17.13 samples per second at 534.18 TFLOPS per GPU with a peak of 119040.02 MiB. The 7B case on 8 B200 cards runs zero1(dp2)+tp2+pp4 with a batch of 128 and reports 25.83 samples per second at 805.69 TFLOPS per GPU with a peak of 100119.77 MiB. Sequence length is 4096 in every row.

Consequence for you: the quoted 50 percent throughput gain for the 7B model is a hardware comparison in which the parallel strategy and the batch size also changed, so it cannot be read as the contribution of the B200 alone. Peak memory moves in the opposite direction of the extra parallelism, which tells you the memory figure is a property of the strategy, not of the card.

## The 70B claim names a B200 that is not in the table

The text under the table says that for the 70B model on 16 cards the B200 showed over 70 percent higher throughput and TFLOPS per GPU. The table has no such row. Both 16-card 70B entries are H200: one at zero2 with a batch of 48, reporting 3.27 samples per second and 469.1 TFLOPS per GPU at a peak of 150032.23 MiB, and one at zero1(dp2)+tp2+pp4 with a batch of 128, reporting 5.66 samples per second and 811.79 TFLOPS per GPU at 100072.02 MiB.

Consequence for you: the improvement that is actually in the table at 70B comes from moving off zero2 onto tensor and pipeline parallelism while quadrupling the batch, on the same GPU. If you are budgeting a 70B training run, this table tells you what a parallelism change bought, and tells you nothing about what a B200 would buy. Anyone citing the B200 advantage at 70B should ask for the missing row.

## setup.py refuses Windows before the build begins

The platform check is the first thing that happens when the package is built, and it is an exception rather than a documented note:

```python
if sys.platform == "win32":
    raise RuntimeError("Windows is not supported yet. Please try again within the Windows Subsystem for Linux (WSL).")
```

Consequence for you: there is no native Windows install, and the failure arrives as a raised error during installation rather than as a line in a support table, which means a pip run on Windows stops before any of the project is placed. The message names WSL as the alternative, so a Windows user who intends to use the project has to plan for a Linux environment from the start, and any tooling that assumes a cross-platform wheel is wrong about this package.

## BUILD_EXT=1 decides whether you get the compiled extensions

The extension build is opt-in through an environment variable, and the default is off:

```python
BUILD_EXT = int(os.environ.get("BUILD_EXT", "0")) == 1
```

setup.py imports torch and BuildExtension from torch.utils.cpp_extension inside a try block, and sets TORCH_AVAILABLE to False when that import fails, so the torch dependency is detected rather than declared for this path. The repo carries .cuda_ext.json, a .clang-format file, and an extensions/ directory, which is where the native side of the project would live.

Consequence for you: a default install gives you the Python layer without the compiled parts, so anyone expecting CUDA kernels has to set the variable to exactly 1 and have a torch toolchain present at build time. Since the check is an integer comparison against the string 1, a value of true or yes leaves the build in its default state, and you get no compiled extension and no warning that you asked for one.

## The version is written into the package during the build

Version handling is a side effect of setup.py. get_version reads version.txt from the project root, strips it, and writes colossalai/version.py with a single assignment line before returning the value, so the installed version string is generated rather than committed. The same file reads requirements files line by line and strips each entry, and it reads README.md as utf-8 for the package metadata.

Consequence for you: a plain source checkout has no version.py until something builds it, so anything that reads the version at runtime depends on a build having run. More importantly for planning, the tagged releases stop at v0.5.0 on 2025-06-04, with v0.4.9 on 2025-03-04 and v0.4.8 on 2025-02-20 before it, while the last push to main is 2026-09-27. What pip installs and what the default branch contains are different things, and the changelog is in CHANGE_LOG.md rather than in the release notes.

## The two sections above the benchmark are paid cloud offers

Before any project content, the front page carries two commercial sections. The first advertises HPC-AI Cloud with pre-configured environments, NVIDIA Blackwell B200s from $2.47 per hour and an H200 cluster from $1.99 per hour, under the line Skip the setup. The second advertises HPC-AI Model APIs naming Kimi 2.5, MiniMax 2.5, and GLM 5.1 with 2M plus context windows and pricing up to 50 percent cheaper than OpenRouter, plus four dollars in free credits.

Consequence for you: the two most prominent entry points on the page are offers from a commercial vendor rather than instructions for running the open source code, and the install steps are not on the front page at all. If you came to evaluate the Apache-2.0 project, the readthedocs documentation and the documentation site are where the local install is described, and the free credits are not the same thing as the free software.

## Submodules and a requirements directory shape a source install

The top level carries .gitmodules, MANIFEST.in, a requirements/ directory, docker/, applications/, extensions/, tests/, docs/, pytest.ini, .coveragerc, .isort.cfg, .clang-format, and a .pre-commit-config.yaml. setup.py reads its requirements through a helper that opens a requirements file and strips every line, which is how the directory is consumed rather than a single pinned file.

Consequence for you: a plain git clone does not fetch what .gitmodules points at, so a source install from a fresh clone can be missing pieces that a build expects, and the failure will look like a missing module rather than a missing submodule. The examples tree is also split by purpose, with community/, inference/, language/, and tutorial/ directories alongside an images/ folder, so the example closest to your task is a directory you have to look for rather than a file listed at the top.

## Four homes for documentation, two of them commercial

The header links a paper on arXiv, a documentation site at colossalai.org, an examples directory, GitHub discussions as the forum, a GPU cloud playground page, and a company blog at hpc-ai.com. The badges add a readthedocs build, a scheduled build workflow, a Slack invite, a Hugging Face organisation, and a WeChat image, and the language toggle points at a Chinese README under docs/README-zh-Hans.md.

Consequence for you: the licence question is simple under Apache-2.0, but the operational answers are split across a documentation site, a readthedocs build, a discussion forum, and a company blog, and the news list on the front page is mostly the company blog, dated back to 2025/02. Anyone maintaining a deployment should record which of those sources they actually read, because the project repository itself is not the place where any of it is written down.

## Conclusion

Colossal-AI is worth reading if you train at 7B to 70B scale and want the parallelism vocabulary in one place, since the examples tree separates inference, language, community, and tutorial work and the parallel strategies in the benchmark are the ones real multi-node jobs are argued about. It is not the place to look for a framework comparison, because its table holds no baseline from another system, and it is not a Windows project, because setup.py raises before a build starts. Before you plan around it, check the tagged release you would install against the branch you would read, since the newest tag is v0.5.0 from 2025-06-04 while the last push is 2026-09-27, and decide whether BUILD_EXT=1 with a torch toolchain is something your build environment can hold.

## FAQ

### What is Colossal-AI?

Colossal-AI is a Python project described as making large AI models cheaper, faster, and more accessible, distributed under the Apache-2.0 license. Its own benchmark table reports throughput, TFLOPS per GPU, and peak memory for 7B and 70B models across parallel strategies including zero2(dp8), zero2, and zero1(dp2)+tp2+pp4 on H200 and B200 hardware.

### How does Colossal-AI compare with DeepSpeed?

The repository does not name DeepSpeed or any other framework as a comparison point. Its benchmark table varies GPU model, parallel strategy, and batch size together and contains no baseline row from a different training system, so a framework comparison cannot be drawn from it.

### Does Colossal-AI install on Windows?

No. setup.py raises a RuntimeError on win32 with the message that Windows is not supported yet and suggests trying again within the Windows Subsystem for Linux. The check runs while the package is being built, so a Windows install attempt stops with that error rather than completing partially.

### How do you build the compiled extensions in Colossal-AI?

The build reads the BUILD_EXT environment variable and enables extensions only when the value is 1, because the check compares an integer conversion of it against 1 with a default of 0. That path also imports torch and BuildExtension from torch.utils.cpp_extension, and a failed import sets TORCH_AVAILABLE to False rather than failing the whole script.

## Sources

- [Official documentation](https://www.colossalai.org)
- [Official README](https://github.com/hpcaitech/ColossalAI#readme)
- [Project repository](https://github.com/hpcaitech/ColossalAI)
- [Release notes](https://github.com/hpcaitech/ColossalAI/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/hpcaitech-colossalai
