vLLM: what PagedAttention buys you, and what it costs to run
A high-throughput and memory-efficient inference and serving engine for LLMs.
At a glance
- What is it?
- vLLM is an Apache-2.0 inference and serving engine for LLMs that trades a heavy CUDA build for paged KV-cache memory management and continuous batching. Here is how it is installed, how the scheduler works, and where it stops being the right tool.
- Who is it for?
- Adopt vLLM if you are serving a Hugging Face model to several concurrent clients on NVIDIA, AMD or Intel GPUs and you want an OpenAI-compatible endpoint without writing a scheduler. Do not adopt it for a single-user laptop demo on CPU, or if you cannot pin torch == 2.13.0 in your environment.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 4 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 25, 2026, and from our analysis. They are not legal advice.
DEEP OPEN-SOURCE ANALYSIS
The memory problem vLLM was written to fix
Serving a decoder-only LLM is mostly a memory management problem. Every in-flight request holds a key/value cache proportional to its context length, and a naive server reserves one contiguous buffer per sequence at maximum length. That reservation is what wastes the GPU: a request that asked for 8K tokens but stops at 300 still occupies the 8K slot until it finishes. vLLM's answer is PagedAttention, described in the project's own paper (Kwon et al., SOSP 2023), which splits the KV cache into fixed-size blocks and maps them through a block table, the same idea as virtual memory paging in an operating system. Blocks are allocated on demand and freed when a sequence ends, so the cache grows with actual token count rather than with the configured ceiling.
The audience is narrow and specific. This is for engineers who already have a model checkpoint and a GPU and need to put an HTTP endpoint in front of it for more than one caller at a time. If your workload is one request at a time from a notebook, the paging machinery buys you nothing and the install cost is real. The README frames the project as "Easy, fast, and cheap LLM serving for everyone", but the repository layout (csrc/, CMakeLists.txt, rust-toolchain.toml, build_rust.sh) makes clear that the fast path is a compiled one.
How the scheduler and the paged KV cache fit together
The architecture visible in the README is a layered one. At the bottom sit optimized attention kernels (FlashAttention, FlashInfer, TRTLLM-GEN, FlashMLA, Triton) and GEMM/MoE kernels built with CUTLASS and CuTeDSL; these live in csrc/ and are compiled at install time. Above them, the execution layer uses piecewise and full CUDA/HIP graphs plus torch.compile for graph-level transformations. Above that, the serving layer does continuous batching, chunked prefill and prefix caching, which is where the throughput story actually comes from: continuous batching means the batch is re-formed at every decoding step instead of waiting for the slowest sequence in a fixed batch to finish.
Quantization is handled as a set of backends rather than a single path: FP8, MXFP8/MXFP4, NVFP4, INT8, INT4, GPTQ/AWQ, GGUF, compressed-tensors, ModelOpt and TorchAO. That breadth is a genuine strength and also a source of confusion, because which formats are usable depends on the GPU generation and the kernel selected. The README lists tensor, pipeline, data, expert and context parallelism for distributed inference, and disaggregated prefill, decode and encode, which separates the compute-heavy prefill stage from the memory-bandwidth-heavy decode stage onto different workers. None of this is configured by default; each is an explicit choice you make at launch.
Installing vLLM with pip and serving a first model
The README recommends uv, with pip as the alternative. Both pull a wheel that already contains the compiled kernels, which is why the install is large and why the wheel is tied to a CUDA or ROCm version.
uv pip install vllmIf you would rather use pip, the README gives the same package name. Note that the build backend pins torch == 2.13.0 in pyproject.toml, so an environment that already holds a different torch will either be rewritten or will conflict. Python support is declared as >=3.10,<3.15.
The repository ships an OpenAI-compatible API server, and the entry point is registered in pyproject.toml as vllm = "vllm.entrypoints.cli.main:main". The README does not print a serve command, so the launch syntax has to come from the quickstart page on docs.vllm.ai rather than from the repository front page. Once the server is listening, clients that already speak the OpenAI chat completions format can point at the local endpoint without code changes, because vLLM also implements the Anthropic Messages API and gRPC. The repository's own examples/ directory is organized by purpose (basic/, deployment/, disaggregated/, features/, observability/, reasoning/, rl/, scale_out/, speech_to_text/, tool_calling/), which is a faster way to find a working configuration than reading the docs linearly. Building from source is a separate path, documented under the GPU installation guide, and it requires the CMake, ninja, Rust and CUDA toolchain that pyproject.toml and rust-toolchain.toml imply.
Where vLLM is the wrong tool
The clearest limitation is the hardware and build surface. The supported targets listed in the README are NVIDIA, AMD and Intel GPUs plus x86/ARM/PowerPC CPUs, with plugins for Google TPUs, Intel Gaudi, IBM Spyre, Huawei Ascend, Rebellions NPU, Apple Silicon and MetaX. Each of those is a separate wheel and a separate set of kernel paths, and the plugin projects (vllm-ascend, vllm-gaudi) are maintained outside this repository. If your accelerator is not on that list, or if its plugin lags the main branch, you are building from source against a moving target.
The second limitation is configuration. PagedAttention and continuous batching decide how memory is allocated, but they do not decide how much. That is the gpu-memory-utilization setting, and getting it wrong produces the classic failure mode: the server starts, the KV cache is sized too small for the concurrency you actually receive, and requests queue or get preempted under load. Nothing in the README tells you what value to pick, because the right value depends on your model's layer count and head dimension. The project's own answer to capacity questions is the benchmarks/ directory, which is a set of scripts you run rather than a published number.
Third, this is not a fine-tuning or training framework. The README describes inference, serving, embedding and reward models. If your task is training a model, vLLM is the wrong layer entirely.
What you give up compared to llama.cpp and TGI
The honest comparison is with llama.cpp and with Hugging Face's Text Generation Inference. llama.cpp is built around GGUF weights and a CPU-first, single-binary design that runs on a laptop with no CUDA toolchain at all; vLLM does list GGUF among its quantization formats, but the engine around it assumes a GPU with a compiled kernel stack. If your deployment target is an Apple Silicon laptop or a CPU-only box with one user, llama.cpp reaches a working server in minutes and vLLM does not.
TGI occupies the middle ground: it is also a server with an OpenAI-compatible surface, but its kernel and scheduling choices are more opinionated and more fixed, whereas vLLM exposes the parallelism strategy, the attention backend, the quantization format and the prefill/decode split as separate decisions. That flexibility is the point. It is also the cost, because a vLLM deployment has more knobs that can be set wrong. The other real difference is breadth of model support: the README claims 200+ architectures on Hugging Face, spanning decoder-only, mixture-of-experts, hybrid attention and state-space models, multimodal, embedding and reward models. A narrower engine will support fewer of those and break less often on the ones it does support.
Release cadence, licence and upgrade cost
The project is not archived, and the last push to main was on 2026-08-26, the same day as the v0.28.0 release. The two prior releases, v0.27.1 and v0.27.0, landed on 2026-08-11 and 2026-08-10. That is a fast cadence, and it has a direct operational consequence: pinning a version is not optional. Because the wheel bundles compiled kernels and pins torch == 2.13.0, an upgrade is not a pure Python change. Moving from v0.27.x to v0.28.0 can change the torch version, the CUDA toolkit expectation and the available attention backends at once, so treat it as a rebuild of the serving image rather than a package bump.
The licence is Apache-2.0, declared in both pyproject.toml and the LICENSE file, with the Apache-2.0 SPDX headers visible in setup.py. Apache-2.0 is permissive and includes an explicit patent grant, which is why it is common in infrastructure. Two things to check yourself rather than assume: the licences of the model weights you serve, which are separate from the engine's licence, and the licences of any hardware plugin you add, since those live in their own repositories. Nothing here is legal advice, and the compatibility of Apache-2.0 with your own distribution model is a question for your counsel.
The README does not document a rollback procedure or a version compatibility matrix, so the practical upgrade path is to keep the previous image tag and the previous pinned wheel available until the new one has served real traffic.
Editorial conclusion
Adopt vLLM if you are serving a Hugging Face model to several concurrent clients on NVIDIA, AMD or Intel GPUs and you want an OpenAI-compatible endpoint without writing a scheduler. Do not adopt it for a single-user laptop demo on CPU, or if you cannot pin torch == 2.13.0 in your environment. Before committing, verify three things on your own hardware: that the wheel for your CUDA or ROCm version resolves, that your model architecture appears on the supported models page, and that the KV cache fits at the gpu-memory-utilization you plan to set.
Frequently asked questions
Why is vLLM so fast?
The README attributes the throughput to PagedAttention, which manages attention key and value memory in blocks instead of one contiguous buffer per sequence, combined with continuous batching of incoming requests, chunked prefill and prefix caching. Optimized attention and GEMM/MoE kernels, CUDA/HIP graphs and torch.compile sit underneath those scheduling choices.
What is vLLM used for?
It is an inference and serving engine for large language models. The README lists decoder-only LLMs, mixture-of-experts models, hybrid attention and state-space models, multimodal models, embedding and retrieval models, and reward and classification models among the 200+ supported architectures.
How can I build vLLM from source?
The README links to a build-wheel-from-source section in the GPU installation guide on docs.vllm.ai rather than giving the steps inline. The repository layout shows what that build needs: CMakeLists.txt, csrc/, rust-toolchain.toml and build_rust.sh, with cmake, ninja and torch == 2.13.0 declared in pyproject.toml.
Does vLLM use PyTorch?
Yes. torch == 2.13.0 is pinned in the build requirements in pyproject.toml, and setup.py imports torch and reads CUDA_HOME and ROCM_HOME from torch.utils.cpp_extension to locate the accelerator toolchain.
Official sources
Where this project is recommended
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/vllm-project-vllm)
Community notes