Chitu: A Production LLM Inference Framework for NVIDIA and Chinese AI Hardware
High-performance inference framework for large language models, focusing on efficiency, flexibility, and availability.
At a glance
- What is it?
- Chitu (赤兔) is an Apache-2 Python inference framework developed at Tsinghua University that runs large language models at scale on NVIDIA GPUs as well as Chinese domestic chips including Huawei Ascend, Moore Threads, Muxi, and Hygon. Its v0.6.0 release introduced the chitu.run single-file launcher for multi-node and prefill-decode-separated deployments.
- Who is it for?
- Chitu fits infrastructure teams in China that need a single inference engine supporting both NVIDIA hardware and domestic chips such as Ascend and Moore Threads, or that need to run large models like DeepSeek-R1 671B on CPU-GPU heterogeneous hardware. Teams outside China using only NVIDIA hardware will find vLLM and SGLang to have larger communities and more extensive documentation in English.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 1 day ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What Chitu Targets and Who Deploys It
Chitu is positioned as a production-grade large model inference engine that covers the progression from small-scale trials to large-scale enterprise deployment. The README describes three core properties: support for a wide range of hardware including both NVIDIA and domestic Chinese chips, scalability from a single CPU or GPU to a large GPU cluster, and stability under concurrent production traffic.
The primary audience is Chinese enterprise teams deploying large language models who need an inference engine that runs on hardware available in China, including chips that vLLM does not support. The project acknowledges contributions and collaboration from Huawei, Moore Threads, Hygon, Enflame (Suiyuan), ZhipuAI, China Telecom, and Paratera, which indicates that the supported hardware set reflects real deployment partnerships rather than theoretical compatibility.
The name Chitu (赤兔) refers to a legendary horse from Chinese history, which reflects the project's origin at Tsinghua University's PACMAN research group.
Hardware Support: NVIDIA, Ascend, Moore Threads, Muxi, and Hygon
The Dockerfile in the repository uses pytorch/pytorch:2.9.1-cuda13.0-cudnn9-devel as the base image for standard NVIDIA builds:
ARG base_image='pytorch/pytorch:2.9.1-cuda13.0-cudnn9-devel'
ARG optional_deps='flash_attn,flash_mla,flashinfer'Separate Dockerfiles are provided for other hardware targets: ascend.Dockerfile for Huawei Ascend, hygon.Dockerfile for Hygon DCUs, and muxi.Dockerfile for Muxi GPUs. The README milestone list shows Ascend 910B support being added in v0.3.5 (June 2025) with full native support, Moore Threads support in v0.5.1 (February 2026), and a v0.4.0 release (August 2025) that covered Ascend, NVIDIA, Muxi, and Hygon in a single release.
Flash Attention (flash_attn), Flash MLA (flash_mla), and FlashInfer are listed as optional dependencies for the NVIDIA path. These improve attention computation throughput on NVIDIA hardware.
Deploying Chitu with chitu.run and the Development Environment
The recommended deployment path, according to the README, is through chitu.run, an executable available in the Assets section of the GitHub Releases page. The v0.6.0 release (July 2, 2026) introduced this single-file launcher, which the README states can start multi-node, multi-instance, and prefill-decode (PD) separated deployments from a single file without additional configuration.
The README notes that complete installation instructions are in docs/zh/DEVELOPMENT.md. The repository includes a setup.py for building from source, which requires PyTorch to be installed first before attempting a build, as stated in the setup.py: torch must be present with the correct CUDA version before running the setup. The build system uses CUDAExtension rather than CMake because many non-NVIDIA GPU vendors provide custom CUDAExtension support but not custom CMake support.
For teams that need to build from source rather than using chitu.run, the build process involves Cython compilation of performance-critical paths and compilation of CUDA extensions. The build time is longer than a standard pip install due to these compilation steps.
DeepSeek-R1 671B Support and Quantization Options
A notable use case in the README's milestone section is running DeepSeek-R1 671B, a model with 671 billion parameters. The v0.1.0 release (March 2025) was the first to support it. Subsequent releases added more efficient quantization paths.
In v0.2.2 (April 2025), Chitu added CPU+GPU heterogeneous mixed inference, which the README states enables running DeepSeek-R1 671B on a single GPU by offloading some computation to CPU. In v0.3.0 (April 2025), FP4 online quantization was added, allowing the FP4 version of DeepSeek-R1 671B from HuggingFace (nvidia/DeepSeek-R1-FP4) to be run with efficient FP4-to-FP8 or FP4-to-BF16 conversion operators.
These capabilities matter for organizations that need to run very large models without a full-scale multi-GPU server. The CPU+GPU heterogeneous path accepts a significant throughput penalty relative to a pure GPU setup, but it makes large models accessible on hardware that cannot hold the full model in GPU VRAM.
Chitu vs. vLLM
vLLM is the most widely deployed open-source LLM inference engine for NVIDIA hardware and has a large English-language community. Chitu's README lists vLLM as one of the projects it learned from. The key difference is hardware scope.
vLLM focuses primarily on NVIDIA hardware and has limited support for non-NVIDIA accelerators. Chitu covers the same NVIDIA deployment scenarios but also provides native support for Huawei Ascend, Moore Threads, Muxi, and Hygon chips, each with a separate Dockerfile optimized for that hardware. For organizations deploying exclusively on NVIDIA hardware in English-language environments, vLLM has a larger ecosystem and more extensive documentation.
SGLang is another alternative in the README's acknowledgments list. SGLang is optimized for complex multi-call inference patterns and constrained decoding. Chitu and SGLang address similar deployment scales but with different optimization priorities.
Limitations, Licensing, and Maintenance Status
The primary documentation is in Chinese. The README has a parallel English README at docs/en/README.md, and there is an English FAQ at docs/en/FAQ.md, but the development manual (docs/zh/DEVELOPMENT.md) and the performance data (docs/zh/PERFORMANCE.md) appear to be Chinese-only based on the file paths. Teams that do not read Chinese will have a reduced documentation surface.
The last push to this repository was on September 24, 2026. The repository is not archived, and the most recent releases are v0.6.0 (July 2, 2026), v0.5.7 (June 4, 2026), and v0.5.6 (May 21, 2026).
The core repository is licensed under Apache 2.0. However, the README explicitly states that third-party code segments are annotated with SPDX headers and their licenses are in the LICENSES/ directory, and that third-party submodules in the third_party/ directory each carry their own license files. Teams deploying Chitu in commercial products should audit both the LICENSES/ directory and the third_party/ submodule licenses before deployment.
Editorial conclusion
Chitu fits infrastructure teams in China that need a single inference engine supporting both NVIDIA hardware and domestic chips such as Ascend and Moore Threads, or that need to run large models like DeepSeek-R1 671B on CPU-GPU heterogeneous hardware. Teams outside China using only NVIDIA hardware will find vLLM and SGLang to have larger communities and more extensive documentation in English. Before adopting Chitu, verify that your hardware is in the supported list in the docs/zh/SUPPORTED_MODELS.md file and review the third-party submodule licenses in the third_party/ directory for any components your deployment needs.
Frequently asked questions
What hardware does Chitu support for LLM inference?
Chitu supports NVIDIA GPUs (multiple series), Huawei Ascend, Moore Threads, Muxi, and Hygon chips. Separate Dockerfiles are provided for each non-NVIDIA hardware target. The README notes that the NVIDIA build uses pytorch/pytorch:2.9.1-cuda13.0-cudnn9-devel as its base image.
Can Chitu run DeepSeek-R1 671B on a single GPU?
Yes. The v0.2.2 release added CPU+GPU heterogeneous inference, which the README states enables single-card inference of DeepSeek-R1 671B by offloading some computation to CPU. A performance penalty applies compared to a multi-GPU deployment.
What is the chitu.run launcher?
chitu.run is a single-file executable introduced in v0.6.0 (July 2, 2026) that can start multi-node, multi-instance, and prefill-decode-separated deployments without additional configuration. It is available in the Assets section of the GitHub Releases page.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/thu-pacman-chitu)