Model or dataset
alibaba/rtp-llm avatar
alibaba/rtp-llm

alibaba/rtp-llm: Alibaba's Production LLM Inference Engine for High-Throughput CUDA Deployments

RTP-LLM: Alibaba's high-performance LLM inference engine for diverse applications.

1,356 stars278 forksPythonApache-2.0

At a glance

What is it?
rtp-llm is an LLM inference acceleration engine developed by Alibaba's Foundation Model Inference Team and deployed in production across multiple Alibaba Group business units. It combines PagedAttention, FlashAttention, INT8 and INT4 quantization, speculative decoding, and a C++-rewritten scheduling layer to serve large language models at high throughput on NVIDIA and ARM hardware.
Who is it for?
Engineering teams who need a production-proven LLM serving engine that has been validated at Alibaba's scale across Taobao, Tmall, Amap, and other large services will find rtp-llm a credible starting point. The documentation is hosted at rtp-llm.ai but the README acknowledges it as a work in progress for the open-source release.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository received new commits within the last day.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What rtp-llm Does and Where It Runs in Production

rtp-llm is a sub-project of Alibaba's havenask search infrastructure. The README states it is widely used within Alibaba Group, supporting LLM services at Taobao Wenwen, Alibaba's international AI platform Aidge, OpenSearch LLM Smart Q&A Edition, and in production research on query rewriting for Taobao Search.

The engine is built to serve large language models at high throughput on NVIDIA GPUs, with particular optimization noted for the V100. It also supports Qwen series models and BERT embedding models on Yitian ARM CPUs, which was announced in January 2025. An AMD ROCm backend and Intel CPU support were listed as in development as of June 2024.

The architecture is built on FasterTransformer from NVIDIA, with additional kernel implementations from TensorRT-LLM integrated over time. The scheduling and batching framework was rewritten in C++ in a June 2024 major refactor, which also added complete GPU memory management and a new Device backend abstraction.

Performance Features: Kernels, Quantization, and Caching

The engine uses a set of high-performance CUDA kernels that the README identifies as PagedAttention, FlashAttention, and FlashDecoding. PagedAttention is the memory management technique that allows the KV cache to be stored in non-contiguous pages, reducing memory waste from reserved but unused cache space during variable-length generation.

Quantization support covers WeightOnly INT8 with automatic conversion at load time, and INT4 via GPTQ and AWQ, both of which are external quantization tools that produce quantized model weights before serving. The engine also supports adaptive KVCache quantization. Dynamic batching overhead has been specifically optimized at the framework level, which the README identifies as an area of detailed engineering focus.

For multi-turn workloads, the engine implements Contextual Prefix Cache and System Prompt Cache, which allow repeated prompt prefixes to be reused across requests rather than recomputed. Speculative decoding is supported, allowing a smaller draft model to generate token candidates that a larger verifier model then accepts or rejects in batch, reducing the number of full forward passes per output token.

Getting Started, Examples, and Citation

The README directs users to the hosted documentation at rtp-llm.ai for install instructions and a quick start guide. The install page is at rtp-llm.ai/build/en/start/install.html and the quick start covers sending a first request at rtp-llm.ai/build/en/backend/send_request.html.

The repository includes example files in the example/ directory for common operations: async_http_client.py for sending asynchronous HTTP requests to the inference server, auto_model_example.py for automatic model loading and inference, get_server_status.py for checking server status, and pause_restart.py for controlling inference lifecycle. The build system is Bazel, indicated by the BUILD, WORKSPACE, and .bazelrc files at the repository root. Docker support is present in the docker/ directory.

Researchers who reference rtp-llm in published work can find the citation entry in the README at the repository root.

The benchmark/ directory and the performance benchmark documentation at rtp-llm.ai/build/en/benchmark/benchmark.html cover the tooling for measuring throughput, latency, and hardware utilization. The Contribution Guide at rtp-llm.ai/build/en/references/Contributing.html covers the process for submitting changes.

Model Compatibility: HuggingFace, LoRA, and Multimodal

The README describes seamless integration with HuggingFace models and lists SafeTensors, PyTorch, and Megatron weight formats as supported. Loading pruned irregular models is also listed as a feature, which addresses fine-tuned models that have had weights set to zero through unstructured pruning.

LoRA is supported in a multi-service configuration: the README states that multiple LoRA adapters can be deployed against a single base model instance. This is useful for serving personalized or task-specific variants without maintaining separate model replicas for each.

Multimodal inputs combining images and text are listed as supported, as are P-tuning models. Multi-machine and multi-GPU tensor parallelism is available for distributing large models across hardware.

Prefill and Decode separation, added in January 2025 and accompanied by a technical report, allows the prefill phase (processing the input context) and the decode phase (generating output tokens) to run on separate hardware, which can improve throughput when batching requests of varying input lengths.

How rtp-llm Compares to vLLM

vLLM is an open-source LLM serving library that originated at UC Berkeley and also implements PagedAttention. It has a larger contributor community, more extensive third-party documentation, and broader ecosystem integrations including OpenAI-compatible API endpoints. For teams that need documented deployment guides and community support, vLLM is a more accessible starting point.

rtp-llm's differentiation, based on the README, is its direct production validation at Alibaba scale and its specific optimizations for the V100 GPU. The C++-rewritten scheduler and the combination of adaptive KVCache quantization with system prompt caching are presented as engineering investments specific to high-volume serving. The Yitian ARM CPU support is also a capability vLLM does not document.

For teams already deep in the Alibaba or Qwen model ecosystem, rtp-llm aligns naturally because the engine is the serving layer Alibaba itself uses for Qwen series model deployments. Teams using other model families like Llama or Mistral who have no particular Alibaba infrastructure dependency will find vLLM's documentation and ecosystem better suited to their starting point.

The documentation gap for rtp-llm is real: the README acknowledges that documentation is still being developed for the open-source community. The primary reference for deployment and API details is the rtp-llm.ai documentation site rather than the repository README itself.

License and Release History

rtp-llm is licensed under Apache-2.0, which permits commercial use, modification, and distribution without requiring derivative works to carry the same license. The last push to the main branch was on 2026-09-27.

Releases include v0.2.0 published on 2025-10-31 with enhanced performance and new features, and earlier v0.1.x releases from April 2024. The gap between April 2024 and October 2025 releases is notable; the v0.2.0 release announcement mentions a major refactor of the scheduling layer that was developed during that period.

The repository includes a NOTICE file alongside LICENSE, which is typical for Apache-2.0 projects built on top of other Apache-2.0 codebases. The README's acknowledgment section cites FasterTransformer, TensorRT-LLM, vLLM, transformers, LLaVA, and Qwen-VL as projects the engine draws from.

The .gitmodules file at the repository root indicates third-party dependencies are managed as git submodules rather than only package dependencies, which affects the clone step: a full recursive clone is needed to get all source dependencies. The 3rdparty/ and deps/ directories hold these external components. Teams integrating rtp-llm into an existing build pipeline should review the Bazel workspace configuration in the WORKSPACE file to understand which third-party libraries are required.

Editorial conclusion

Engineering teams who need a production-proven LLM serving engine that has been validated at Alibaba's scale across Taobao, Tmall, Amap, and other large services will find rtp-llm a credible starting point. The documentation is hosted at rtp-llm.ai but the README acknowledges it as a work in progress for the open-source release. Teams that need a broad ecosystem of integrations, a large community, and thorough third-party tutorials should evaluate vLLM first, since rtp-llm's documentation and contributor community are smaller. The Apache-2.0 license permits commercial use without restriction.

Frequently asked questions

What is rtp-llm and how is it related to Alibaba's products?

rtp-llm is an LLM inference engine developed by Alibaba's Foundation Model Inference Team and deployed across Taobao Wenwen, Aidge, OpenSearch LLM Smart Q&A, and other Alibaba services. It is a sub-project of the havenask search infrastructure.

How does rtp-llm differ from vLLM?

Both implement PagedAttention, but rtp-llm adds a C++-rewritten scheduler, adaptive KVCache quantization, system prompt caching, and specific V100 GPU optimizations. vLLM has a larger community and more third-party documentation; rtp-llm's documentation is still being developed for open-source users.

Does rtp-llm support HuggingFace models?

Yes. The README describes integration with HuggingFace models and lists SafeTensors, PyTorch, and Megatron weight formats as supported. Multiple LoRA adapters can also be deployed against a single base model instance.

Official sources

  1. alibaba/rtp-llm on GitHub
  2. Issues
  3. License: Apache-2.0
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/alibaba-rtp-llm.svg)](https://hysenlabs.com/projects/alibaba-rtp-llm)