Open-source project
baidu/vLLM-Kunlun avatar
baidu/vLLM-Kunlun

vLLM-Kunlun: running vLLM on Kunlun XPU through a platform plugin

vLLM Kunlun (vllm-kunlun) is a community-maintained hardware plugin designed to seamlessly run vLLM on the Kunlun XPU.

472 stars103 forksPythonApache-2.0

At a glance

What is it?
Baidu's community-maintained hardware plugin lets vLLM target the Kunlun XPU without patching vLLM itself. It is narrow by design: one accelerator family, a short OS list, and a dependency set you have to reconcile by hand.
Who is it for?
Adopt vLLM-Kunlun if you already own Kunlun3 P800 hardware on Ubuntu 20.04 and want the standard vLLM OpenAI server rather than a vendor fork. Do not adopt it for CUDA or ROCm hosts, for CPU-only serving, or for any model outside the README tables.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 6 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The problem vLLM-Kunlun solves for Kunlun XPU owners

vLLM ships with a platform abstraction, and the project cites RFC Hardware Pluggable (vllm-project/vllm issue 11162) as the model it follows. In practice that means a hardware vendor can register a backend through Python entry points instead of maintaining a patched vLLM fork. vLLM-Kunlun is that registration layer for the Kunlun XPU, described in the README as community-maintained and as the recommended way to integrate the Kunlun backend.

The audience is narrow. You need a Kunlun3 P800 accelerator, Ubuntu 20.04, Python 3.10 or newer, and a KL3-customized xpytorch build of PyTorch 2.5.1 or newer. If you are serving on NVIDIA or AMD hardware, this project has nothing to offer you. If you have Kunlun silicon and want to keep using upstream vLLM's request handling, scheduler and OpenAI-compatible HTTP layer rather than a bespoke inference server, the plugin is the point of contact.

The README claims support for more than twenty mainstream models across Transformer, MoE, embedding and multimodal families, with a table naming Qwen, Llama, DeepSeek, GLM, Gemma4 and Kimi-K2 among others. Those tables are the real specification. They also encode the gaps: LoRA columns are blank for Qwen3.5, Qwen3.5-MoE and every DeepSeek entry, and the Kunlun Graph column is blank for gpt-oss.

How the plugin hooks into vLLM's platform layer

The mechanism is entirely entry points. pyproject.toml declares one platform plugin and three general plugins:

toml
[project.entry-points."vllm.platform_plugins"]
kunlun = "vllm_kunlun:register"

[project.entry-points."vllm.general_plugins"]
kunlun_model = "vllm_kunlun:register_model"
kunlun_reasoning_parser = "vllm_kunlun:register_reasoning_parser"
kunlun_tool_parser = "vllm_kunlun:register_tool_parser"

When vLLM enumerates its platform plugins, it calls vllm_kunlun:register. The general plugins add model definitions, a reasoning parser and a tool parser to vLLM's registries. Because this happens at import time through packaging metadata, no vLLM source file is edited. setup.py duplicates the same four entry points, so both build paths produce the same registration surface.

The package also installs a console script, vllm-kunlun = "vllm_kunlun.cmdline:main", declared in pyproject.toml. The README does not document what that command does, which is a gap worth noting before you build tooling around it. Data files under vllm_kunlun/conf and vllm_kunlun/data ship inside the wheel via package-data, so configuration and operator artifacts travel with the install rather than being fetched at runtime.

The repository layout reinforces the patch-oriented design: a vllm_kunlun/patches/ directory holds assets including the logo and a performance image referenced by the README. Whether runtime patching happens beyond registration is not spelled out in the README, so treat the entry points as the documented contract and inspect vllm_kunlun/ if you need more.

Installing vllm-kunlun and serving a model over the OpenAI API

The README points to https://vllm-kunlun.readthedocs.io/en/latest/installation.html for installation rather than giving inline steps, and it lists a KL3-customized xpytorch build as a prerequisite. Install that PyTorch build first, then vLLM at the matching version, then the plugin. The repository includes setup_env.sh and build.sh at the top level, and setup.py reads the version from vllm_kunlun/platforms/version.py, so an editable install from a checkout is the path the source tree is arranged for:

bash
pip install -e .

Dependencies are the part that bites. requirements.txt pins transformers==5.5.3 with a comment stating that vLLM 0.18.0 metadata pins transformers<5,>=4.56.0, but Qwen3.5 needs transformers 5.x. The file tells you the override is deliberate and suggests installing vLLM with --no-deps or applying requirements.txt last if the resolver objects. It also pins pydantic==2.12.0, ray==2.48.0, compressed-tensors==0.17.0 and a long list of runtime packages. Note that pyproject.toml declares an empty dependencies list; the pinned set lives in requirements.txt, so installing only the wheel will not pull it in.

Once the environment is in place, the README's quick start starts the OpenAI-compatible server with vLLM's own entry point, which the plugin has redirected to Kunlun:

bash
python -m vllm.entrypoints.openai.api_server \
    --host 0.0.0.0 \
    --port 8356 \
    --model <your-model-path> \
    --gpu-memory-utilization 0.9 \
    --trust-remote-code \
    --max-model-len 32768 \
    --tensor-parallel-size 1 \
    --dtype float16 \
    --max_num_seqs 128 \
    --max_num_batched_tokens 32768 \
    --block-size 128 \
    --no-enable-prefix-caching \
    --no-enable-chunked-prefill \
    --distributed-executor-backend mp \
    --served-model-name <your-model-name>

Replace the model path and served name with your own values. The server binds port 8356 and exposes the standard chat completions route, which the README exercises with curl:

bash
curl http://localhost:8356/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "<your-model-name>",
    "messages": [{"role": "user", "content": "Hello!"}],
    "max_tokens": 512
  }'

If the server starts and returns a completion, the plugin registered correctly. If vLLM reports no usable platform, the entry points did not resolve and the install is the first thing to check.

Where vLLM-Kunlun is the wrong tool

The hardware requirement is absolute. Kunlun3 P800 is the only accelerator named in the README, and Ubuntu 20.04 is the only OS listed. A container built on a newer Ubuntu base is outside the documented configuration, and the README does not describe a supported workaround.

The dependency override is the second constraint. Forcing transformers to 5.5.3 while vLLM 0.18.0 wants transformers<5 means the resolver will complain, and the README's own advice is to install vLLM with --no-deps or apply requirements.txt last. That is a manual, order-sensitive environment. Reproducible builds require you to pin the whole set yourself.

Feature coverage is uneven, and the tables say so. gpt-oss has no Kunlun Graph entry. LoRA is available only for the Qwen series models whose LoRA column is marked. InternLM2 and Seed-OSS list no quantization support. DeepSeek and Kimi-K2 entries show no LoRA. If your workload depends on adapters for a DeepSeek model, this plugin does not cover it.

The documented quick start also disables two vLLM features explicitly: --no-enable-prefix-caching and --no-enable-chunked-prefill. Workloads that rely on prefix caching for shared system prompts will not get that behavior on this backend as documented. The README does not explain whether these are temporary constraints or architectural ones, and it does not document rollback if an upgrade regresses.

vLLM-Ascend, vLLM-MPS and vLLM-Metax: same pattern, different silicon

The plugin model is not unique to Kunlun. vLLM-Ascend, vLLM-MPS and vLLM-Metax follow the same idea: a vendor or community package registers a platform backend so upstream vLLM can run on non-NVIDIA accelerators. The difference between them is which hardware they target and how much of the model matrix they cover, not the integration shape.

That matters for evaluation. The relevant question is not whether a plugin architecture exists, since several do, but whether the specific plugin supports the models and quantization schemes you need on the silicon you own. vLLM-Kunlun's answer is its two tables plus the quantization column. If you are choosing hardware rather than software, those tables are the comparison document, and they are more informative than any general claim about plugin support.

The release cadence is also worth reading. v0.11.0 shipped on 2026-03-13, preceded by v0.11.0rc1 on 2025-12-26 and v0.10.1.1 on 2025-12-24. The repository's last push was on 2026-09-10, and the default branch is v0.25.1-dev, with the README marking v0.25.1 as under development. The latest tagged release and the development branch are therefore far apart, which is normal for a project tracking upstream vLLM versions but means the README's newest features may not correspond to any tagged artifact.

Licence, upgrade cost and what to check before you commit

The project is Apache-2.0, declared in both the repository metadata and setup.py, with the full text in LICENSE. The README does not discuss model licences, and the plugin ships no weights. If you serve Llama, Gemma or DeepSeek models through it, those terms come from the model publisher, not from this repository. Nothing here is legal advice; read the model licence you actually deploy.

Upgrade cost is driven by the version coupling. The version matrix lists v0.25.1 as the latest development line, and requirements.txt carries comments keyed to vLLM 0.18.0 and vLLM 0.25.1, which shows the dependency set is rewritten as upstream moves. Because the plugin hooks vLLM's platform and registry APIs, a vLLM bump can change the entry-point contract. Pinning both vLLM and vllm-kunlun together, and re-reading requirements.txt at each upgrade, is the practical approach.

Two things are undocumented and worth confirming on your own hardware. First, the README does not document rollback, so keep the previous environment around before upgrading. Second, the performance section shows a chart for 16-way concurrency with input and output size 2048 on Kunlun3 P800, but the numbers live in an image and the README states no methodology, so do not treat it as a benchmark you can extrapolate from.

Editorial conclusion

Adopt vLLM-Kunlun if you already own Kunlun3 P800 hardware on Ubuntu 20.04 and want the standard vLLM OpenAI server rather than a vendor fork. Do not adopt it for CUDA or ROCm hosts, for CPU-only serving, or for any model outside the README tables. Before committing, verify three things: that your xpytorch build matches the plugin version, that your chosen model appears in the supported-model tables with the quantization mode you plan to use, and that you can satisfy requirements.txt, where transformers is deliberately pinned to 5.5.3 against vLLM 0.18.0's own metadata.

Frequently asked questions

What is vLLM-Kunlun used for?

It runs vLLM on the Kunlun XPU as a hardware plugin, registering a platform backend so models such as Qwen, Llama, DeepSeek and GLM can be served through vLLM's OpenAI-compatible API. The README describes it as the recommended way to integrate the Kunlun backend within the vLLM community.

What hardware and software does vLLM-Kunlun require?

The README lists Kunlun3 P800 hardware, Ubuntu 20.04, Python 3.10 or newer, a KL3-customized xpytorch build of PyTorch 2.5.1 or newer, and a matching vLLM version. Qwen3.5 additionally requires transformers 5.x.

What is the current version of vLLM-Kunlun?

The version matrix lists v0.25.1 as the latest development line, and the default branch is v0.25.1-dev. The most recent tagged release shown is v0.11.0 from 2026-03-13.

Official sources

  1. baidu/vLLM-Kunlun on GitHub
  2. License: Apache-2.0
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/baidu-vllm-kunlun.svg)](https://hysenlabs.com/projects/baidu-vllm-kunlun)