PaddlePaddle FastDeploy: an LLM and VLM serving toolkit with PD disaggregation and a vLLM-compatible API
High-performance Inference and Deployment Toolkit for LLMs and VLMs based on PaddlePaddle
At a glance
- What is it?
- FastDeploy is PaddlePaddle's production inference toolkit for large language and vision-language models, covering PD disaggregation, quantisation, speculative decoding and multi-vendor accelerators. The README documents the feature set well but leaves install commands to per-hardware pages, and the project is Linux and Python 3.10 to 3.12 only.
- Who is it for?
- FastDeploy fits teams already running PaddlePaddle models, especially ERNIE and PaddleOCR families, or teams on Kunlunxin XPU, Hygon DCU, Iluvatar, Enflame, Metax or Intel Gaudi hardware where the mainstream CUDA serving stacks are not an option. It is the wrong pick if you need Windows or macOS, if your stack is already built around vLLM or SGLang and you have no PaddlePaddle dependency, or if you want a single pip install with no vendor-specific wheel.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 5 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What FastDeploy is for, and who ends up using it
FastDeploy is the deployment half of the PaddlePaddle stack. The repository describes it as an inference and deployment toolkit for large language models and vision-language models, positioned as a production serving layer rather than a training or experimentation library. The models it names first in its release notes are ERNIE-4.5, ERNIE-4.5-VL and PaddleOCR-VL, which tells you where the engineering effort is concentrated.
The audience is narrower than the feature list suggests. You are the intended user if you have a model checkpoint in Paddle format or a HuggingFace-format model the toolkit already lists, and you need to put it behind an HTTP endpoint with batching, caching and quantisation. The README's own quick-start path is a ten-minute deployment guide, and the online serving documentation sits alongside offline inference, so both batch scoring and a long-running server are first-class cases.
The second audience is hardware. FastDeploy lists NVIDIA GPU, Kunlunxin XPU, Iluvatar CoreX, Enflame S60, Hygon DCU, Metax GPU and Intel Gaudi as supported targets, each with its own installation page and, in several cases, its own requirements file at the repository root. If you are on a domestic Chinese accelerator, this breadth is the reason to look here before anywhere else. If you are on a single NVIDIA card, the breadth is irrelevant and the comparison shifts to what the serving engine does.
PD disaggregation, KV cache transport and the router
The headline mechanism is load-balanced PD disaggregation. Prefill and decode run as separate instance roles, and the README describes dynamic switching of instance roles plus context caching, with the stated goal of holding SLO targets and throughput while improving resource utilisation. That is a real architectural commitment: prefill work is compute-bound and decode work is memory-bandwidth-bound, so splitting them lets you size each pool differently instead of provisioning one shape for both.
Moving the KV cache between those pools is handled by a separate transport library, described as lightweight and able to choose between NVLink and RDMA. A router component performs the load balancing, and setup.py downloads a prebuilt binary named fd-router from a Paddle CI bucket during installation, selecting the x86_64 or aarch64 build by host architecture. That download is a network dependency at install time, not just at runtime, and it is the kind of step that fails quietly behind a restrictive proxy.
Around that core sit the usual serving accelerators: chunked prefill, prefix caching, global cache pooling, speculative decoding and multi-token prediction. Quantisation covers W8A16, W8A8, W4A16, W4A8, W2A16 and FP8, with W4AFP8 added in v2.5.0. The service layer exposes an OpenAI-compatible API and the README states compatibility with the vLLM interface, so existing clients that speak the OpenAI chat completions schema should not need rewriting. The repository also acknowledges borrowing parts of vLLM's code to keep that interface compatible, which is worth knowing when you compare behaviour between the two.
Installing FastDeploy and serving a first model
There is no single install command in the README. Installation is split by hardware, and each page under docs/zh/get_started/installation/ carries the commands for that accelerator. The top-level requirements.txt is the Python dependency set, not a self-contained installer, and it pulls in paddleformers, fastapi, uvicorn, redis, etcd3, triton, cupy-cuda12x and a flashinfer wheel hosted on a Baidu object store. Start from the page that matches your card rather than from pip.
The environment constraints are explicit: Linux only, Python 3.10 through 3.12. The README badges confirm both. There is no Windows or macOS path documented.
Once the environment is in place, the documented entry point for a server is the OpenAI-compatible serving mode. The README points to docs/zh/online_serving/README.md and to docs/zh/get_started/quick_start.md, which is the ten-minute path. The README does not reproduce the full server command, so read the online serving page for the exact flags before running anything. What you should see once it starts is a uvicorn server accepting requests, since uvicorn and fastapi are both in requirements.txt.
Because the service is OpenAI-compatible and the repository pins openai>=1.93.0, the official client is the natural first check. Set the base URL to the endpoint your server printed and send one chat completion. If you get text back, the serving path works end to end. If the request hangs, the problem is usually the router rather than the model, since fd-router is fetched separately during installation.
Where FastDeploy gets in the way
The installation story is the biggest friction point. A prebuilt router binary downloaded from paddle-qa.bj.bcebos.com during setup.py means air-gapped clusters need a mirror or a manual placement step, and the README does not document an offline install for it. Nothing in the repository describes how to verify that binary's provenance beyond the URL itself.
Platform support is genuinely narrow. Linux only, and the Python range stops at 3.12. If your serving fleet is containerised on a base image with Python 3.13, you are rebuilding that image.
The dependency list is heavy. Redis and etcd3 appear as requirements, which implies coordination services for the distributed paths, and the OpenTelemetry instrumentation packages for Redis, MySQL, FastAPI and logging suggest the observability examples under examples/observability/ assume that stack. A single-node deployment that never touches PD disaggregation still installs all of it.
The documentation is also primarily Chinese. The README is the Chinese version with an English translation at README_EN.md, and the installation and quick-start links in the body point at docs/zh/ paths. English readers will be following translated pages for the details that matter most.
Finally, the interface compatibility with vLLM cuts both ways. It lowers migration cost, but it also means the project's own documentation does not always spell out where behaviour diverges, since the assumption is that you already know the vLLM surface.
FastDeploy against vLLM and SGLang
The honest comparison is not about features, because the three overlap heavily. vLLM and SGLang both offer high-throughput serving with paged attention, continuous batching and OpenAI-compatible endpoints, and FastDeploy explicitly aims at interface compatibility with vLLM. Choosing between them comes down to three axes.
The first is model ecosystem. FastDeploy's releases are organised around ERNIE-4.5, ERNIE-4.5-VL, Qwen3-VL, Qwen3-MoE, DeepSeek V3 and PaddleOCR-VL. If your model is in that set, or in Paddle format generally, FastDeploy has the tuned path. If your model is a Llama variant or a community fine-tune, vLLM's broader model coverage is the safer bet.
The second is hardware. vLLM is a CUDA-first project with growing support elsewhere. FastDeploy ships separate requirements files for DCU, Iluvatar and Metax GPUs and dedicated installation pages for Kunlunxin, Enflame and Gaudi. On non-NVIDIA silicon this is the actual differentiator, and it is why the project exists in its current form.
The third is operational shape. FastDeploy's router and its fd-router binary make PD disaggregation a first-class deployment mode with a load-balancing component in front. vLLM supports disaggregated prefill in its own way, and SGLang has its own router, but the components and their failure modes differ. If you already run one of those in production, switching buys you nothing unless the model or hardware argument applies.
Maintenance, releases and the Apache-2.0 licence
The last push to the develop branch was on 2026-08-26, and the repository is not archived. Releases are roughly quarterly: v2.3.0 on 2025-11-11, v2.4.0 on 2026-01-23, v2.5.0 on 2026-04-09. The v2.5.0 notes claim over 170 bug fixes and performance optimisations alongside the Qwen3-VL and W4AFP8 additions. That cadence is regular enough to plan upgrades around, but note that the release notes are the only upgrade documentation the README points to. There is no migration guide, and no rollback procedure is described.
Upgrade cost is dominated by the dependency chain rather than the FastDeploy code itself. The pinned ranges include transformers>=4.55.1,<5.0.0, openai>=1.93.0, uvicorn>=0.38.0 and a specific flashinfer wheel URL. A major transformers release will require FastDeploy to move its ceiling, and until it does, you are held back. The flashinfer wheel is pinned to a URL rather than a version range, so that dependency cannot be resolved from PyPI alone.
Licensing is Apache-2.0, stated in the README and present as LICENSE at the repository root. The README also states that parts of vLLM's code were referenced and borrowed to maintain interface compatibility. vLLM is itself Apache-2.0, so the licences align, but if you redistribute a modified FastDeploy you should read the LICENSE file and the vLLM attribution yourself rather than relying on this summary. Nothing here is legal advice.
Editorial conclusion
FastDeploy fits teams already running PaddlePaddle models, especially ERNIE and PaddleOCR families, or teams on Kunlunxin XPU, Hygon DCU, Iluvatar, Enflame, Metax or Intel Gaudi hardware where the mainstream CUDA serving stacks are not an option. It is the wrong pick if you need Windows or macOS, if your stack is already built around vLLM or SGLang and you have no PaddlePaddle dependency, or if you want a single pip install with no vendor-specific wheel. Before committing, open docs/zh/get_started/installation/nvidia_gpu.md and confirm a wheel exists for your exact accelerator and driver combination, then run the 10-minute quick start on one model to check that the router binary downloads and starts on your network.
Frequently asked questions
What is FastDeploy used for?
It is PaddlePaddle's inference and deployment toolkit for large language models and vision-language models, covering both offline inference and online serving. Its documented features include PD disaggregation, quantisation, speculative decoding, prefix caching and an OpenAI-compatible API.
How do I install FastDeploy for an NVIDIA GPU?
There is no single install command in the README. Installation is split by hardware, and the NVIDIA path is documented at docs/zh/get_started/installation/nvidia_gpu.md, which is the page the README links to from its install section. Linux and Python 3.10 to 3.12 are the stated requirements.
Does FastDeploy work on non-NVIDIA hardware?
Yes. The README lists Kunlunxin XPU, Iluvatar CoreX, Enflame S60, Hygon DCU, Metax GPU and Intel Gaudi alongside NVIDIA GPU, each with its own installation page. It is Linux only, and there is no Windows or macOS path documented.
Is FastDeploy compatible with the vLLM API?
The README states that its OpenAI API service is vLLM-compatible, and it acknowledges referencing and borrowing parts of vLLM's code to keep that interface compatible. It also exposes an OpenAI-compatible API, so clients using that schema should work without changes.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/paddlepaddle-fastdeploy)