# vLLM Ascend plugin: running vLLM on Ascend NPU hardware

> vllm-ascend is the community maintained hardware plugin that lets vLLM target Ascend NPUs through the hardware-pluggable interface. It is the right layer for Ascend operators who want vLLM's serving stack, and the wrong one for anyone without that silicon.

**vllm-project/vllm-ascend** — Community maintained hardware plugin for vLLM on Ascend

- Repository: https://github.com/vllm-project/vllm-ascend
- Website: https://docs.vllm.ai/projects/ascend
- Stars: 2,916 · Forks: 2,386
- Language: C++
- License: Apache-2.0
- Published: 2026-09-09 · Updated: 2026-09-09 · Language: en
- Canonical page: https://hysenlabs.com/projects/vllm-project-vllm-ascend

## What vllm-ascend takes off your plate

vLLM's engine was written against CUDA, and the kernels, memory allocator and collective paths all assume it. Running the same engine on an Ascend NPU means replacing that lower layer without forking the scheduler. vllm-ascend exists to be that lower layer. The README calls it "a community maintained hardware plugin for running vLLM seamlessly on the Ascend NPU" and says it follows the hardware-pluggable RFC (vllm-project/vllm issue 11162), which is the interface that lets a backend be swapped without the upstream project carrying vendor code.

The audience is narrow and specific. You need Ascend NPUs, a working CANN install, and a reason to prefer vLLM's serving behaviour over whatever came with the hardware. If you are evaluating Ascend hardware for LLM serving, this plugin is the piece that decides whether you can reuse the vLLM ecosystem (its OpenAI-compatible server, its continuous batching, its model zoo) or whether you are locked into a vendor-only runtime. The README points readers at a support matrix for models and features rather than listing them inline, which tells you the coverage is moving and should be checked per release.

## How the plugin sits between vLLM and the NPU

The repository is a Python package with a compiled C++ core. The vllm_ascend/ directory holds the Python side that vLLM's plugin mechanism loads; csrc/ holds the C++ sources built through CMakeLists.txt and the cmake/ directory. setup.py builds extensions with setuptools and CMake and derives the version from setuptools-scm, so an install from a git checkout gets a version string from git metadata rather than from a hardcoded constant.

That split matters when something breaks. Scheduler-level symptoms (batch shapes, preemption, prefix cache hits) surface in the Python layer, while kernel-level symptoms (unsupported dtype, wrong output on a specific op) come from csrc/. The build requirements pin the NPU-side stack tightly: torch==2.10.0 with torch-npu==2.10.0.post4 and triton-ascend==3.2.2. Those are not loose lower bounds. A mismatch between the installed torch-npu and the one the extension was compiled against is the most likely cause of an import-time failure.

The repository also carries several Dockerfiles rather than one: Dockerfile, Dockerfile.openEuler, and variants suffixed 310p, a3 and a5, each with an openEuler counterpart. That layout is a direct statement that the supported hardware is not a single target and that the base OS pairing is part of the support surface.

## Installing vllm-ascend and running a first inference

The plugin depends on a CANN base image and a specific PyTorch NPU build, so the container path is the one the repository is structured around. The default Dockerfile starts from quay.io/ascend/cann at tag 9.1.0-910b-ubuntu22.04-py3.12 and installs clang-15, which triton-ascend needs.

```dockerfile
ARG CANN_QUAY_URL="quay.io/ascend/cann"
ARG CANN_VERSION="9.1.0"
FROM ${CANN_QUAY_URL}:${CANN_VERSION}-910b-ubuntu22.04-py3.12
```

If you build from source instead, the runtime requirements are listed in requirements.txt and mirrored in pyproject.toml. The pins to respect are the ones that tie the Python packages to the NPU runtime.

```bash
pip install -r requirements.txt
```

The file pins torch==2.10.0, torch-npu==2.10.0.post4, triton-ascend==3.2.2 and transformers==5.14.1. Installing a newer torch-npu on top of a wheel built against these will not fix anything.

The repository ships runnable scripts under examples/. The plainest entry point is examples/offline_inference_npu.py, with tensor-parallel and long-sequence variants beside it. Running the offline example is the fastest way to confirm the plugin loaded and the NPU is visible, before you touch the server path. If you want the HTTP server, that comes from vLLM itself once the plugin is registered; the plugin supplies the backend, not a separate CLI.

## Where vllm-ascend is the wrong choice

The first constraint is hardware. Nothing here helps on a GPU or a CPU. The C++ extensions target Ascend, and the Dockerfiles pull a CANN base image. If you do not have Ascend NPUs and a matching CANN install, this project has no use for you at all.

The second is version coupling. The repository tracks vLLM's release cadence, and the release list shows a rapid sequence: v0.13.0, v0.18.0, v0.23.0, then v0.26.0rc1 on 2026-09-03. The newest tag is a release candidate, and the README's news entries point each release at a versioned copy of the documentation (docs.vllm.ai/projects/ascend/en/v0.26.0rc1/ and so on). That structure is honest about the fact that instructions are version-specific. It also means an upgrade is not a one-line change: you move the plugin, the vLLM version it plugs into, and the torch-npu and triton-ascend pins together, or you spend an afternoon on import errors.

The third is the support matrix. The README directs readers to it for supported models and features instead of promising broad coverage, and it lists Transformer-like, Mixture-of-Experts, Embedding and multi-modal LLMs as categories that can run. Category-level support is not the same as support for the specific checkpoint you have. Check the matrix for your model before planning a deployment around it.

## vllm-ascend against MindIE and against upstream vLLM

The comparison people actually make is with MindIE, the vendor-side inference stack for Ascend. The difference is architectural, not cosmetic. MindIE is a complete serving product: it owns the model loading, the scheduler, the API surface and the deployment tooling. vllm-ascend owns none of that. It plugs Ascend into vLLM's existing scheduler, so the serving behaviour you get is vLLM's, including its continuous batching and its OpenAI-compatible server. Choosing vllm-ascend means choosing the vLLM ecosystem and accepting that you are one plugin deep in a stack with several moving version pins.

The other comparison is with upstream vLLM. Upstream does not ship Ascend support in its own tree; the hardware-pluggable RFC exists precisely so that support can live in a separate repository like this one. So the two are not competitors. You install vLLM, then you install this plugin, and the plugin's job is to make the engine treat the NPU as a valid device. That also explains why the plugin's release notes are tied to vLLM versions rather than being independently numbered in a different scheme.

## Maintenance, licence and the cost of staying current

The project is not archived, and the last push was on 2026-09-09, so the repository is receiving changes. The release history shows a steady cadence through 2025 and 2026, with the most recent tag being v0.26.0rc1 and the most recent stable tag v0.23.0 on 2026-08-16. That cadence is the upgrade cost: a project that mirrors vLLM's releases will ask you to move roughly as often as vLLM does, and each move drags the torch-npu and triton-ascend pins with it.

The licence is Apache-2.0, the same as vLLM itself. That is permissive: you can ship it in a commercial product, modify it, and redistribute it, provided you keep the notices and state changes. The one thing worth reading carefully is that the plugin builds against CANN, which is a separate Huawei component with its own terms, and against torch-npu and triton-ascend, which are separate packages. The Apache-2.0 grant in this repository covers this repository's code, not those dependencies. That is a fact about the boundary, not legal advice; if the combination is going into a product, the dependency licences are what your counsel needs to see.

## Conclusion

Adopt vllm-ascend if you already own Ascend NPUs and want vLLM's scheduler, batching and OpenAI-compatible server on top of them, and you are prepared to track two version chains at once. Do not adopt it if you have no Ascend hardware, if you need a stable long-term tag rather than a release candidate, or if you want a vendor stack that manages the model lifecycle for you. Verify first that your CANN version and the pinned torch-npu and triton-ascend versions match the support matrix, and that the model you intend to serve appears in that matrix rather than only in the examples directory.

## FAQ

### What is vLLM Ascend?

It is a community maintained hardware plugin that runs vLLM on the Ascend NPU. The README describes it as following the hardware-pluggable RFC so that Ascend support lives outside the upstream vLLM tree.

### How do I install vLLM-ascend?

The repository is structured around its Dockerfiles, which start from a CANN base image such as quay.io/ascend/cann at tag 9.1.0-910b-ubuntu22.04-py3.12. A source install uses requirements.txt, which pins torch==2.10.0, torch-npu==2.10.0.post4 and triton-ascend==3.2.2.

### What are the latest releases of vLLM Ascend?

The most recent tag is v0.26.0rc1, released on 2026-09-03, and the most recent stable tag is v0.23.0 from 2026-08-16. The README links each release to a versioned copy of the documentation.

### vllm ascend vs mindie: what is the difference?

MindIE is a complete vendor serving stack that owns the scheduler and the API surface. vllm-ascend owns neither; it plugs Ascend into vLLM's existing engine, so you get vLLM's serving behaviour and take on its version pins.

### vllm ascend vs vllm: are they alternatives?

No. Upstream vLLM does not carry Ascend support in its own tree; the hardware-pluggable RFC is what allows that support to live in a separate repository. You install vLLM and then this plugin.

## Sources

- [License: Apache-2.0](https://github.com/vllm-project/vllm-ascend/blob/main/LICENSE)
- [Project website](https://docs.vllm.ai/projects/ascend)
- [README](https://github.com/vllm-project/vllm-ascend/blob/main/README.md)
- [Releases](https://github.com/vllm-project/vllm-ascend/releases)
- [vllm-project/vllm-ascend on GitHub](https://github.com/vllm-project/vllm-ascend)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/vllm-project-vllm-ascend
