# ZhiLight: Zhihu's CUDA Inference Engine for Llama and DeepSeek on PCIe GPUs

> ZhiLight is an Apache-2.0 C++ inference engine from Zhihu and ModelBest, tuned for PCIe-based GPUs and shipped with an OpenAI-compatible server. Its own benchmarks show large TTFT gains on some configurations and losses on others, so the choice depends on your hardware.

**zhihu/ZhiLight** — A highly optimized LLM inference acceleration engine for Llama and its variants.

- Repository: https://github.com/zhihu/ZhiLight
- Stars: 908 · Forks: 104
- Language: C++
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/zhihu-zhilight

## The PCIe gap ZhiLight was built to close

Most open source serving stacks were designed with high-bandwidth interconnect in mind. On a node of A800 or H20 cards connected over PCIe, the communication between GPUs becomes a first-order cost, and a kernel schedule that assumes NVLink leaves throughput on the table. ZhiLight is Zhihu's answer to that specific problem. The README states the engine "can accelerate the inference of models like Llama and its variants, especially on PCIe-based GPUs", and the performance section is organized around PCIe devices: AD102 as a consumer card for experimental work, A800 and H20 as data center cards for production.

The target user is an infrastructure engineer who already has a rack of PCIe GPUs and cannot change that. The project is not a general-purpose serving framework for every model family. The supported list in the README is Llama and Llama2, Llama 3.1 and 3.2 and 3.3, Mixtral, Qwen2 series, DeepSeek-VL2 multimodal models, and DeepSeek-V3 and DeepSeek-R1 in AWQ, GPTQ and FP8 block-quantized form. Anything outside that list is unverified territory.

## Dual streams, host all-reduce and the memory model

Three mechanisms in the README explain where the speed comes from. The first is what the project calls "dual streams": encode and all-reduce are overlapped, so the communication step runs concurrently with computation rather than after it. The README also notes support for Int8-quantized all-reduce, which cuts the bytes moved per collective. The second is a host-side all-reduce built on SIMD instructions, which matters when the interconnect is PCIe and the GPU-to-GPU path is the bottleneck. The third is a custom tensor type with unified global memory management, which is the foundation the other two sit on: a single allocator has to know about every buffer before it can schedule overlap correctly.

On top of that layer sit the usual serving features. Fused kernels for qkv, residual and layernorm. Fused batch attention for decoding using tensor core instructions. Flash attention prefill, chunked prefill, prefix cache and dynamic batching. Tensor parallelism and pipeline parallelism are both supported on one node, with the README recommending TP. MoE is supported, including DeepseekV2 MoE and DeepseekV2 MLA, which is what makes the DeepSeek-R1 numbers possible at all. The server side is an OpenAI-compatible interface the README describes as "adapted from vllm", which is a useful hint about the request schema and the response shape.

## Installing ZhiLight and serving a model

The README gives build commands directly. The wheel is produced by setup.py, which drives a CMake build; the setup.py file shows CMAKE_CXX_STANDARD=17 and a TESTING environment variable that toggles the test configuration. To build concurrently and skip unit tests, the README shows this:

```bash
CMAKE_BUILD_PARALLEL_LEVEL=32 TESTING=0 python setup.py bdist_wheel
```

The same file lists CMAKE_GENERATER as the variable read for selecting a backend, and the README gives the ninja variant. Note the spelling in the README: CMAKE_GENERATER, not CMAKE_GENERATOR, which is the name setup.py reads internally. Copy the README form exactly when following its instructions.

```bash
CMAKE_GENERATER="Ninja" python setup.py bdist_wheel
```

For a development install from a checkout, the README uses pip directly:

```bash
cd ./ZhiLight && pip install -e .
```

If you would rather not build at all, the project publishes a container image. The README states ZhiLight depends only on the CUDA runtime, cuBLAS, NCCL and the packages in requirements.txt, and points at docker/Dockerfile as the reference build.

```bash
docker pull ghcr.io/zhihu/zhilight/zhilight:0.4.8-cu124
```

The requirements.txt pins torch>=2.4.1, transformers>=4.46.2, fastapi>=0.115.5, uvicorn>=0.32.0, prometheus_client>=0.21.0, psutil>=6.1.0, pybind11==2.13.6 and flash-attn>=2.7.0.post2. Once installed, the server starts through a module path:

```bash
python -m zhilight.server.openai.entrypoints.api_server [options]
```

The README writes the options as a placeholder and does not enumerate the flags. That is a real gap: you will have to read the entrypoint source to learn which model path, tensor-parallel size and port arguments it accepts. The examples directory is the better starting point for a first run. It contains offline_inference.py for batch work, online_batch_completion.py for a batch HTTP client and online_stream_chat.py for streaming, all under examples/.

## Where ZhiLight's own benchmarks show it losing

The performance tables are worth reading closely, because they do not support a blanket claim of superiority. On DeepSeek-R1 AWQ with eight A800 cards, ZhiLight reports QPS 0.16 against vLLM's 0.08, and lower TTFT mean and P95, but its TPOT mean of 115.97 and P95 of 139.99 are worse than vLLM's 109.86 and 129.98. Doubling throughput while degrading per-token latency is a trade-off, not a clean win, and interactive chat is exactly the workload where TPOT is felt.

The Qwen2-72B-Instruct-GPTQ-Int4 results are more uneven. On four AD102 PCIe cards ZhiLight reports TTFT mean of 1111.8 against SGLang's 2276.1 and vLLM's 3493.97, with the best TPOT of the three. On four A800 cards the ordering flips: SGLang reports QPS 0.36 and the best TTFT, while ZhiLight sits at QPS 0.18 with TPOT mean 31.95, behind both vLLM at 22.14 and SGLang at 30.41. The Qwen1.5-110B-Chat-GPTQ-Int4 table shows the same pattern: ZhiLight leads on AD102 and trails on A800. MiniCPM-2B-sft-bf16 on a single AD102 card is a near-tie on QPS at 1.67 for all three engines, with ZhiLight best on TTFT and SGLang best on TPOT P95.

The honest reading is that the advantage is configuration-dependent, and the README says as much when it lists test purpose as demonstrating "performance, applicable scenarios and limitations". The benchmarks live in docs/benchmarks/benchmarks.md, and the README does not describe the load pattern, concurrency level or prompt lengths behind the numbers. Without that, the tables tell you which configurations to test on your own hardware, not which engine to pick.

## Build cost, version lag and the upgrade question

ZhiLight is a C++ project with a Python packaging layer, and the cost of that shows up in three places. First, the build itself: a CMake compile against CUDA, cuBLAS and NCCL, with flash-attn and a pinned pybind11==2.13.6 in requirements.txt. A pinned pybind11 is a common source of conflicts when it shares an environment with other extension modules. Second, the release cadence. The most recent release listed is v0.4.8 from 2024-12-10, while the README news entries run through 2025/05 and the last push to the repository was on 2026-03-18. Development has continued past the release tag, so a user who installs the published 0.4.8 image is not running the code described in the newest news items. The Docker tag itself, 0.4.8-cu124, encodes CUDA 12.4, which constrains the driver version on the host.

Third, there is no documented rollback or downgrade procedure in the README. If a new build regresses on your workload, the README does not describe how to return to a previous wheel or image. The practical implication is that you should pin the image tag you deploy and keep the previous one available, because the project documentation will not tell you how to go back.

On licensing, the repository carries Apache-2.0 and includes a NOTICE file. Apache-2.0 permits commercial use and modification and includes a patent grant, but it also requires that you preserve the licence and NOTICE text in redistributions and state significant changes. If you ship ZhiLight inside a product, those obligations travel with it. That is a summary of the licence text, not legal advice; your counsel should review the NOTICE file and any third-party components under 3rd/ before you redistribute.

## ZhiLight against vLLM and SGLang

The README names vLLM and SGLang as its comparison points, and the difference is one of scope rather than features. vLLM is a general serving framework with a broad model registry and a large ecosystem of integrations; its OpenAI-compatible server is the reference many tools target. SGLang couples a serving runtime with a structured generation language, which is why it appears in the benchmarks as a separate engine rather than a drop-in replacement for a raw HTTP endpoint.

ZhiLight takes the opposite approach: a narrow model list, a CUDA-only build, and kernel-level work aimed at PCIe topologies. Its OpenAI interface is, in the README's words, "adapted from vllm", so clients written against vLLM's API have a reasonable chance of working, but the adaptation is not documented as a compatibility guarantee. If your requirement is breadth of model support or a large plugin ecosystem, vLLM is the safer default and ZhiLight is the wrong tool. If your requirement is squeezing latency out of a fixed set of Llama, Qwen or DeepSeek checkpoints on PCIe cards, ZhiLight is designed for exactly that and the benchmarks are its argument.

## Conclusion

Adopt ZhiLight when you are serving Llama, Qwen, Mixtral or DeepSeek variants on PCIe A800, H20 or AD102 class hardware with tensor parallelism on a single node, and you are willing to build from source against CUDA, cuBLAS and NCCL. Do not adopt it if you need multi-node serving, a documented rollback path, or if your workload matches the A800 Qwen2-72B GPTQ case where its own numbers put SGLang ahead. Verify first that the wheel builds with your CUDA and PyTorch versions, that your quantization format is on the supported list, and that the OpenAI endpoint your client uses is one the api_server actually exposes.

## FAQ

### How does llama.cpp work, and how does it differ from ZhiLight?

The README does not describe llama.cpp. ZhiLight is presented as an LLM inference engine developed by Zhihu and ModelBest Inc. for accelerating models like Llama and its variants, especially on PCIe-based GPUs, and it serves an OpenAI-compatible interface adapted from vllm.

### What is an LLM inference engine?

ZhiLight's README treats an inference engine as the layer that runs model inference and serves requests. It lists features such as dynamic batch, prefix cache, chunked prefill, tensor and pipeline parallelism, and quantization support, and it exposes an OpenAI-compatible server through zhilight.server.openai.entrypoints.api_server.

### Who created llama.cpp, and who created ZhiLight?

The README does not say who created llama.cpp. ZhiLight's README states it was developed by Zhihu and ModelBest Inc., and lists contributors @a710128, @spetrel, @unix1986 and @gnap.

## Sources

- [Issues](https://github.com/zhihu/ZhiLight/issues)
- [License: Apache-2.0](https://github.com/zhihu/ZhiLight/blob/main/LICENSE)
- [README](https://github.com/zhihu/ZhiLight/blob/main/README.md)
- [Releases](https://github.com/zhihu/ZhiLight/releases)
- [zhihu/ZhiLight on GitHub](https://github.com/zhihu/ZhiLight)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/zhihu-zhilight
