# xLLM: a C++ inference engine aimed at Chinese AI accelerators

> xLLM is an Apache-2.0 inference engine for LLM, VLM, DiT and REC models, built around Ascend NPU, Cambricon MLU, MUSA, DCU, MACA and Iluvatar hardware. This review covers what it does, how to install it, and where it is the wrong choice.

**xLLM-AI/xllm** — A high-performance inference engine for LLM, VLM, DiT and REC models, optimized for diverse AI accelerators. It is hosted in OpenAtom Foundation.

- Repository: https://github.com/xLLM-AI/xllm
- Website: https://xllm-ai.com/
- Stars: 1,587 · Forks: 304
- Language: C++
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/xllm-ai-xllm

## What xLLM is for, and who it is aimed at

The README describes xLLM as an efficient LLM inference framework, specifically optimized for Chinese AI accelerators, and frames the goal as enterprise-grade deployment with reduced cost. The hardware table lists Ascend NPU (A2, A3), Cambricon MLU, Moore Threads MUSA (S5000), Hygon DCU (BW1000), MetaX MACA (MXC500) and Iluvatar CoreX (BI150). That list is the product. An engineer running NVIDIA GPUs already has vLLM, TensorRT-LLM and SGLang competing for the job, and xLLM offers no reason to switch. An engineer holding a rack of Ascend or MLU cards has far fewer options, and that is the reader this project is written for.

The stated scope is wider than text generation. The repository description covers LLM, VLM, DiT and REC models, and the examples directory contains generate.py, generate_beam_search.py, generate_embedding.py, generate_vlm.py and sample.py. So the intended user is not only someone serving a chat model, but someone who needs one engine to cover vision-language, embedding and diffusion-style workloads on the same non-NVIDIA fleet. The README also claims the engine is battle-tested at scale across JD.com's core retail business, which is a deployment claim rather than a benchmark, and should be read as such: it says the project has survived production traffic, not that it is faster than anything else. Anyone weighing xLLM against a CUDA-first stack should start by asking which of those six accelerator families they actually own.

## Service-engine decoupling and the rest of the architecture

The README's highlights name a service-engine decoupled architecture: the service layer handles scheduling and availability, the engine layer handles computation. That split is the central design decision, and it is visible in the repository layout. The HTTP surface is built on brpc, according to the acknowledgment section, so the scheduling and request-handling side is a C++ service rather than a Python process. The engine side draws its graph construction method from ScaleLLM and references that project's runtime execution. Tokenization is a C++ implementation built on tokenizers-cpp, and weight loading relies on the C binding of safetensors. A C++ JSON parser is implemented with insights from Python and Go implementations of a partial JSON parser.

The other architectural piece worth naming is the KV cache. The news entries state that xLLM builds hybrid KV cache management on top of Mooncake, supporting global KV cache management with intelligent offloading and prefetching. Offloading and prefetching imply a cache that can live outside accelerator memory and be pulled back on demand, which is the mechanism that lets a decoupled service layer share state across engine processes. The README does not document the eviction policy, the tier layout, or how a cache miss is handled, so anyone planning a multi-node deployment should treat that as an open question to resolve in the docs site rather than something the README settles.

The build system is setuptools-driven. setup.py imports a set of environment helpers named set_cuda_envs, set_dcu_envs, set_ilu_envs, set_maca_envs, set_mlu_envs, set_musa_envs and set_npu_envs from scripts/build_support/env.py, which is a compact statement of intent: the build is parameterised per accelerator vendor, and the device type is detected at build time. On the NPU path, setup.py can compile TileLang kernels through xllm/compiler/tilelang_launcher.py with the ascend target and a device platform argument. That is a source build with vendor toolchains, not a portable wheel, and it explains why the install instructions live on a docs site rather than in the README.

## Installing xLLM and running a first generation

The README does not carry install commands. It points to a Quick Start page at docs.xllm-ai.com/en/getting_started/quick_start/, a Launch xLLM page, and a Docker image at quay.io/repository/jd_xllm/xllm-ai. The package metadata in pyproject.toml declares the distribution name as xllm, requires Python 3.10 or newer, and registers a console entry point: xllm maps to xllm.launch_server:main. So the launch path is a Python command, even though the engine underneath is C++.

What is not given is the exact pip invocation, the wheel index, or whether a prebuilt wheel exists for any given accelerator. Because setup.py detects the device type and sets vendor-specific environment variables before building, the honest instruction is to follow the Quick Start page for your hardware rather than guessing a pip line. The entry point itself is the one command the repository files do confirm, since pyproject.toml declares it under the project.scripts table:

```bash
xllm = "xllm.launch_server:main"
```

That line is the console script declaration, not a shell command to paste. It tells you that after installation the executable is named xllm and it starts the launch server. For a first real generation, the examples directory is the place to look, and the repository lists these files: examples/generate.py, examples/generate_beam_search.py, examples/generate_embedding.py, examples/generate_vlm.py and examples/sample.py. The README does not document their arguments, their model path flags, or their default output, so treat the files themselves as the reference. The same applies to examples/generate_vlm.py for vision-language models and examples/generate_embedding.py for embeddings. For a serving deployment rather than a script, the launch server entry point is the one to read, and the Launch xLLM page is where the flags are documented. The point of showing the entry point is to make clear where the documentation stops and the source begins, and to keep a first run from turning into an afternoon of guessing flag names.

## The supported model list moves faster than the release tags

The news section is a timeline of day-0 model support: GLM-5.3-Flash on 2026-08-27, MiniMax-M3 on 2026-06-13, DeepSeek-V4 on 2026-04-24, GLM-5 on 2026-02-12, GLM-4.7 on 2025-12-21, GLM-4.6V on 2025-12-08, and the GLM-4.5/GLM-4.6 series plus VLM-R1 on 2025-12-05. Several of those entries link to a script under a preview branch, for example preview/glm-5.3-flash/testspace/run_glm_53_flash.sh. That detail matters more than the headline. Day-0 support for a model that lands in a preview branch means the deployment recipe is a branch-specific shell script, not a stable, versioned feature. If you need reproducibility, pin to a release tag and check whether the model you want is covered there.

The release cadence visible in the tags is roughly monthly to quarterly: v0.10.1 on 2026-07-14, v0.10.0 on 2026-07-01, v0.9.1 on 2026-04-14. The gap between v0.9.1 and v0.10.0 is under three months, and v0.10.1 followed v0.10.0 by two weeks, which suggests patch releases are used to fix regressions rather than holding them for the next minor. The repository's last push was on 2026-09-10, so the project is not dormant. That said, the README does not document an upgrade procedure, a compatibility matrix between engine versions and model checkpoints, or a rollback path. For an engine that loads multi-gigabyte weights and compiles vendor kernels, that is a real gap, and it is the first thing to raise with the maintainers if you plan to run it in production.

## Where xLLM is the wrong tool

The clearest limitation is the hardware table read in reverse. xLLM is purpose-built for Chinese AI accelerators; the README says so directly. If your fleet is NVIDIA, the ecosystem around vLLM and SGLang is larger, the documentation is deeper, and the community answers more questions. Choosing xLLM there buys nothing and costs you the ability to search for solutions to your problems. The same applies to AMD.

A second limitation is the build. setup.py is a substantial file that imports vendor environment helpers and, on the NPU path, can invoke a TileLang kernel compiler with a target platform and a job count. That is a source build against vendor SDKs. Teams that expect a single pip install with prebuilt binaries for every accelerator will be disappointed, and teams without the matching driver and toolkit versions will fail at build time rather than at runtime with a clear message. The README's hardware table gives the driver floor for Ascend only, which implies the other five vendors have their own undocumented prerequisites.

A third is documentation depth. The README is a landing page. It links out to a docs site for Quick Start, Launch, Online Service, Offline Inference and Supported Models, and the local docs/ directory exists, but the README itself does not explain the request lifecycle, the scheduling policy, or the KV cache tiering. If your evaluation depends on understanding the scheduler before you trust it, budget time for reading the source and the technical report on arXiv rather than the README. And if your workload is a single small model on a single GPU, none of this machinery is aimed at you.

## How it differs from vLLM

vLLM is the obvious comparison, and it appears in the related searches for this project. The difference is not a feature list; it is a target. vLLM's design center is NVIDIA GPUs, with a PagedAttention block manager and a Python-centric serving stack. xLLM's design center is a set of non-NVIDIA accelerators, with a C++ service layer on brpc and an engine layer that draws its graph construction from ScaleLLM. Practically, that means xLLM's build pipeline is organised around vendor SDK detection, and its performance work is aimed at the memory hierarchies and kernel libraries of Ascend, MLU, MUSA, DCU, MACA and Iluvatar parts.

The KV cache story also differs. xLLM builds hybrid KV cache management on Mooncake, with offloading and prefetching, which is a distributed-cache design rather than a purely local block manager. Whether that translates into better behaviour under multi-tenant load is not something the README establishes. What it does establish is that the two projects make different bets about where the hard problem is: vLLM bets on GPU memory management and a Python ecosystem, xLLM bets on vendor portability and a decoupled service tier. A team already running vLLM on NVIDIA hardware has no migration story here, and the README does not offer one.

## Licence, governance and upgrade cost

xLLM is Apache-2.0, stated in the README badge and in pyproject.toml as the text form of the same licence. The repository carries a THIRDPARTYNOTICES.md file, which is where the obligations for the bundled dependencies live. The acknowledgment section names ScaleLLM, Mooncake, brpc, tokenizers-cpp, safetensors, a partial JSON parser and concurrentqueue, and the third_party/ directory exists at the top level. If you redistribute xLLM inside a product, read THIRDPARTYNOTICES.md and the individual licences of those dependencies; Apache-2.0 on the outer project does not automatically settle the terms of everything linked into the binary. This is not legal advice, and the file is the authoritative source.

Governance changed in a way that matters for procurement. A news entry dated 2026-07-06 states that xLLM was officially donated to the OpenAtom Foundation, and the repository description confirms it is hosted there. For enterprise buyers, foundation hosting is usually easier to justify than single-vendor hosting, but it also means the project's roadmap is now partly a foundation question. The upgrade cost itself is the practical concern: releases arrive every few weeks, model support often lands on preview branches first, and the README documents neither a rollback path nor a version compatibility matrix. Pinning a release tag and testing a model checkpoint against it before upgrading is the only procedure the project's own files support.

## Conclusion

Adopt xLLM if your serving fleet is Ascend NPU, Cambricon MLU, Moore Threads MUSA, Hygon DCU, MetaX MACA or Iluvatar CoreX, and you want a C++ engine with a Python launcher instead of a CUDA-first stack. Do not adopt it if your accelerators are NVIDIA or AMD, or if you need a project whose rollback and upgrade procedure is written down; the README does not document rollback. Before committing, verify on your own hardware that the driver level matches the HDK Driver 25.2.0 or newer the README requires for Ascend A2 and A3, and check the supported models page for the exact checkpoint you intend to serve.

## FAQ

### What is xLLM and what models does it support?

xLLM is a C++ inference engine for LLM, VLM, DiT and REC models, hosted in the OpenAtom Foundation and licensed Apache-2.0. The README's news entries record day-0 support for GLM-5.3-Flash, MiniMax-M3, DeepSeek-V4, GLM-5, GLM-4.7, GLM-4.6V, the GLM-4.5/GLM-4.6 series and VLM-R1.

### Which hardware does xLLM run on?

The README's hardware table lists Ascend NPU (A2, A3), Cambricon MLU, Moore Threads MUSA (S5000), Hygon DCU (BW1000), MetaX MACA (MXC500) and Iluvatar CoreX (BI150). It states HDK Driver 25.2.0 or newer for Ascend A2 and A3.

### How do I install xLLM?

The README does not give install commands; it points to the Quick Start page at docs.xllm-ai.com and to a Docker image at quay.io/repository/jd_xllm/xllm-ai. The distribution is named xllm, requires Python 3.10 or newer, and registers an xllm console entry point that maps to xllm.launch_server:main.

### Does xLLM work with vLLM or replace it?

The README does not mention vLLM at all. xLLM is a separate engine built on brpc for its HTTP service and drawing graph construction from ScaleLLM, and it targets Chinese AI accelerators rather than NVIDIA GPUs.

## Sources

- [License: Apache-2.0](https://github.com/xLLM-AI/xllm/blob/main/LICENSE)
- [Project website](https://xllm-ai.com/)
- [README](https://github.com/xLLM-AI/xllm/blob/main/README.md)
- [Releases](https://github.com/xLLM-AI/xllm/releases)
- [xLLM-AI/xllm on GitHub](https://github.com/xLLM-AI/xllm)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/xllm-ai-xllm
