LMDeploy: Deploying LLMs with TurboMind and PyTorchEngine
LMDeploy is a toolkit for compressing, deploying, and serving LLMs.
At a glance
- What is it?
- LMDeploy gives teams running large language models two inference backends, a built-in AWQ quantization pipeline, and a tensor-parallel API server, all under one package. The key decision is whether the C++ TurboMind engine fits your hardware, or whether the pure-Python PyTorchEngine is the safer starting point.
- Who is it for?
- Teams that serve InternLM, DeepSeek, or Qwen models at scale on NVIDIA hardware and need a quantization-to-serving pipeline without switching tools should look at LMDeploy first. Teams on AMD ROCm, MTHREADS MACA, or Cambricon hardware can try the corresponding requirements files, but the documentation for those platforms is thinner and the PyTorchEngine, not TurboMind, is the engine that supports Ascend via a graph mode added in late 2024.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 7 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 25, 2026, and from our analysis. They are not legal advice.
Editorial analysis
TurboMind and PyTorchEngine: Choosing a Backend
LMDeploy ships two inference engines with different engineering tradeoffs. TurboMind is a C++ inference engine built on custom CUDA kernels. Its feature list, as described in the README, includes Paged Attention, Split-K decoding (a form of Flash Decoding), W4A16 inference, and online int8/int4 KV cache quantization. PyTorchEngine is written entirely in Python, added in January 2024 specifically to lower barriers for developers and enable rapid experimentation. Huawei Ascend hardware is supported only through PyTorchEngine, which added graph mode on Ascend in late 2024 according to the release notes, doubling inference speed on that platform.
The two engines are not interchangeable for every model. TurboMind handles the performance-critical production path; PyTorchEngine covers models and hardware that TurboMind does not yet reach. A team evaluating LMDeploy should confirm which engine supports their target model before building a serving stack around it.
AWQ, W4A16, and Online KV Cache Quantization
The quantization pipeline covers two distinct targets: model weights and the key-value cache. For weights, LMDeploy uses the AWQ (Activation-aware Weight Quantization) algorithm to produce 4-bit weight-only quantized models. The README states that 4-bit inference performance is 2.4x higher than FP16. Pre-quantized models are available on the HuggingFace Hub under the lmdeploy organization.
In February 2026, support was added for vllm-project/llm-compressor, covering both 4-bit symmetric and asymmetric quantization. A separate guide at docs/en/quantization/llm_compressor.md covers that path.
For the KV cache, TurboMind supports online int8 and int4 quantization, applied at inference time without reprocessing the model weights. The guide at docs/en/quantization/kv_quant.md covers the setup. The combination of W4A16 model weights and int4 KV cache is the configuration that yields the highest memory compression, though the README is silent on what happens to output quality under joint quantization; the OpenCompass evaluation cited in the README covers weight quantization only.
Installing LMDeploy and Running a First Request
Installation from PyPI requires pip:
pip install lmdeployThe package name and command come directly from the README. PyPI availability was disrupted before April 2026 due to storage quota limits; v0.12.3 was the version that resumed normal PyPI uploads. The current release at time of this writing is v0.18.0, published on 2026-09-28.
The setup.py uses cmake_build_extension as its build backend and detects the installed CUDA major version at build time, pulling the matching NCCL, cuBLAS, and cuRAND libraries from PyPI. On a machine without matching pre-built wheels, pip will attempt to compile TurboMind from source, which requires a C++ toolchain and CMake.
For the first request after installation, the README directs users to the Quick Start guide at https://lmdeploy.readthedocs.io/en/latest/get_started/get_started.html. The README itself is silent on the specific invocation commands, so the documentation site is the required starting point for the actual model-loading and inference calls.
OpenAI-Compatible Serving and Multi-GPU Tensor Parallelism
LMDeploy exposes an OpenAI-compatible API server, documented at docs/en/llm/api_server.md. The proxy_server.md file covers multi-model, multi-machine, multi-card inference services, added in January 2024. A Kubernetes deployment directory (k8s/) and a Docker directory (docker/) are both present in the repository, so containerized or orchestrated deployments have starting materials available.
Tensor parallelism is listed as one of the core mechanisms in the README's introduction. The inference design combines persistent batching (what the README also calls continuous batching), blocked KV cache, and dynamic split-and-fuse to raise throughput. The README claims up to 1.8x higher request throughput than vLLM for the internlm2-20b model at 16+ requests per second.
DeepSeek models added a specific deployment path in June 2025 through integration with DLSlime and Mooncake for PD (Prefill-Decode) disaggregation. That configuration is documented separately from the standard serving path and requires those external libraries.
Hardware Constraints: CUDA First, Others via Separate Requirements
The repository carries five platform-specific requirements files: requirements_cuda.txt, requirements_ascend.txt, requirements_camb.txt, requirements_rocm.txt, and requirements_maca.txt. CUDA is the primary supported platform, and TurboMind's MXFP4 support, added in September 2025, targets NVIDIA GPUs from V100 onward.
For AMD ROCm, Cambricon, and MTHREADS MACA hardware, the requirements files exist but the README provides no guidance on setup steps for those platforms. The news section of the README covers Ascend specifically (September 2024 and October 2024 entries), pointing to a separate getting-started guide at docs/en/get_started/ascend/. For the other non-CUDA platforms, the README does not document what is and is not implemented.
Running LMDeploy on CPU-only machines is not described anywhere in the README or repository documentation. The architectural emphasis on CUDA kernels and tensor parallelism across GPUs makes a CPU path unlikely for TurboMind, though the PyTorchEngine's pure-Python design could in principle reach other backends.
Where vLLM Approaches the Same Problem Differently
vLLM is a Python-first LLM inference server that also implements PagedAttention and continuous batching. The difference the README implies is architectural: LMDeploy invests in a dedicated C++ engine (TurboMind) with CUDA kernels tuned for quantized inference, while keeping a Python engine (PyTorchEngine) for models and hardware outside TurboMind's reach. The README attributes the throughput comparison to the internlm2-20b model specifically, not to a general benchmark.
The practical consequence for teams choosing between them: LMDeploy's quantization pipeline is built in rather than bolted on, covering both weights and the KV cache with matching tooling. Teams already invested in vLLM's ecosystem, or using models that LMDeploy does not yet support, will find switching costs that the repository does not address. The supported model list lives in docs/en/supported_models/supported_models.md; checking it against the target model is the concrete first step.
Editorial conclusion
Teams that serve InternLM, DeepSeek, or Qwen models at scale on NVIDIA hardware and need a quantization-to-serving pipeline without switching tools should look at LMDeploy first. Teams on AMD ROCm, MTHREADS MACA, or Cambricon hardware can try the corresponding requirements files, but the documentation for those platforms is thinner and the PyTorchEngine, not TurboMind, is the engine that supports Ascend via a graph mode added in late 2024. Anyone expecting a zero-config experience similar to Ollama's model-pull workflow will not find it here; LMDeploy assumes familiarity with Hugging Face model identifiers and tensor-parallelism configuration. Before committing, verify that your CUDA version is compatible with the pre-built wheels on PyPI: the setup.py detects the installed CUDA major version and pulls the matching NCCL and cuBLAS libraries, and a mismatch at install time surfaces as a C++ extension build failure.
Frequently asked questions
What is LMDeploy?
LMDeploy is a Python toolkit for compressing, deploying, and serving large language models, developed by the MMRazor and MMDeploy teams. It provides two inference backends: TurboMind, a C++ engine with custom CUDA kernels, and PyTorchEngine, a pure-Python engine. It supports AWQ 4-bit quantization and an OpenAI-compatible API server.
What is the difference between LMDeploy's TurboMind and PyTorch engines?
TurboMind is a C++ inference engine with custom CUDA kernels supporting features like Paged Attention, W4A16 inference, and online KV cache quantization. PyTorchEngine is written entirely in Python, added to lower developer barriers and enable experimentation on hardware TurboMind does not yet support, including Huawei Ascend. The two engines do not support identical model sets.
How does LMDeploy compare to vLLM?
The README states that LMDeploy delivers up to 1.8x higher request throughput than vLLM for the internlm2-20b model. The architectural difference is that LMDeploy's TurboMind engine applies C++ CUDA kernel tuning specifically for quantized inference, while its PyTorchEngine covers hardware and models outside TurboMind's scope. The README does not compare the two across a broader model set.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/internlm-lmdeploy)