Model or dataset
microsoft/vattention avatar
microsoft/vattention

vAttention: dynamic KV-cache memory for LLM serving without rewriting attention kernels

Dynamic Memory Management for Serving LLMs without PagedAttention

524 stars46 forksCMIT

At a glance

What is it?
vAttention is a CUDA virtual memory based KV-cache allocator that keeps KV-cache contiguous in virtual memory while allocating physical pages on demand, so unmodified attention kernels keep working. It ships as a memory allocator plus a modified Sarathi-Serve serving stack, and it needs PyTorch 2.3.0 and CUDA 12.1.
Who is it for?
Adopt vAttention if you serve LLMs on A100 class GPUs, your attention kernels are FlashAttention or FlashInfer based, and you want dynamic KV-cache allocation without rewriting kernels for paged memory. Do not adopt it if you are not on CUDA 12.1 with PyTorch 2.3.0, if your serving stack is vLLM and you cannot move to the bundled Sarathi-Serve, or if you need 64KB to 256KB pages and are unwilling to replace the CUDA UVM driver.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 37 days ago.
What is it written in?
Mainly C, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 27, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The KV-cache allocation problem vAttention targets

Serving a large language model means holding a KV-cache that grows and shrinks with the request mix. The standard answer, PagedAttention, treats GPU memory like an operating system treats RAM: it splits the cache into fixed-size blocks and implements demand paging in user space. That works, but it forces a rewrite of every attention kernel so that it can gather from non-contiguous blocks. vAttention takes the opposite route. It keeps the KV-cache contiguous in virtual memory and uses the CUDA virtual memory APIs to map physical pages into that contiguous range only when they are actually needed. From the attention kernel's point of view, nothing changed. The project is aimed at engineers running LLM inference on NVIDIA hardware who want dynamic memory allocation without maintaining a paged kernel variant. The README states that vAttention also improves performance over PagedAttention in many cases, especially for prefill-bound workloads, and points to the paper at arxiv.org/abs/2405.04437 for the numbers.

How the allocator decouples virtual and physical memory

The mechanism is a split between what the kernel sees and what the driver commits. Virtual address space is reserved up front and stays contiguous, so attention kernels index it exactly as they would a statically allocated cache. Physical memory is committed on demand through the CUDA driver's virtual memory API, which the README links as the CUDA VA group in the driver API documentation. The repository is organized around that split: the vattention directory holds the allocator itself, sarathi-lean is a modified Sarathi-Serve that can run either PagedAttention or vAttention style memory management, scripts holds the benchmark runners, microbenchmarks holds smaller tests, and nvidia-vattn-uvm-driver holds a modified NVIDIA UVM driver. There is also a pod_attn entry at the top level, though the README does not describe it. One design choice worth noting: vAttention can overlap memory allocation with compute, and the _sync suffix in the backend knob disables that overlap. The README recommends the asynchronous form, which tells you the overlap is considered the normal path rather than an optimization you opt into.

Installing vAttention and running a first benchmark

The README pins the environment tightly: PyTorch 2.3.0 and CUDA 12.1 or later, with a caveat that other CUDA versions may or may not work. The project was tested on Linux, A100 GPUs and Python 3.10. Start with the conda environment the README gives.

bash
conda create -n vattn python=3.10
conda activate vattn

The allocator is built against libtorch, and the README warns that the libtorch version has to match the torch version, with only v2.3.0 tested. Download and extract it before building anything.

bash
wget https://download.pytorch.org/libtorch/cu121/libtorch-shared-with-deps-2.3.0%2Bcu121.zip
unzip libtorch-shared-with-deps-2.3.0+cu121.zip

Then build Sarathi-Serve and the vAttention allocator. Note the extra index URL, which pulls FlashInfer wheels for the cu121 and torch2.3 combination, and the LIBTORCH_PATH variable, which must point at the extracted libtorch directory.

bash
cd sarathi-lean/
pip install -e . --extra-index-url https://flashinfer.ai/whl/cu121/torch2.3/
cd ../
cd vattention/
LIBTORCH_PATH=<path to libtorch dir> python setup.py install
cd ../

The quickest way to confirm the setup works is the test mode of either benchmark script. The static trace script reproduces the makespan results from the paper; the dynamic trace script runs 256 arxive dataset requests at qps values of 0.4, 0.8, 1, 2, 4 and 6 with Poisson arrivals.

bash
python scripts/benchmark_e2e_static_trace.py --test

A successful run writes results under experiments/e2e_static_eval or experiments/e2e_dynamic_eval, and the matching process_e2e_static.py or process_e2e_dynamic.py script parses them. Model configurations live in scripts/utils.py, and the README says all Yi and Llama family models are expected to work.

Choosing an attention backend and page size

The attention_backends list in the benchmark scripts is where the memory strategy becomes concrete. Supported values are fa_paged_[block_size], fi_paged_[block_size], fa_vattn_[page_size], fi_vattn_[page_size], fa_vattn_[page_size]_sync and fi_vattn_[page_size]_sync, where fa means FlashAttention (tested at v2.5.9) and fi means FlashInfer (tested at v0.0.6). The README recommends block size 256 for FlashAttention and 16 for FlashInfer, based on observed performance. The vAttention variants accept 64KB, 128KB, 256KB and 2MB page sizes, written for example as fa_vattn_256kb or fi_vattn_2mb_sync. The FlashInfer vAttention wrapper is described as experimental and uses FlashInfer's single_prefill_with_kv_cache for non-paged prefill and FlashAttention's flash_attn_with_kvcache for non-paged decode. That combination is the clearest demonstration of the project's main claim: the same contiguous cache layout serves two different kernel libraries. Treat the experimental label seriously if you plan to put FlashInfer on the critical path.

The UVM driver requirement for small pages

This is the sharpest constraint in the project. NVIDIA CUDA drivers allocate memory only at the granularity of large pages, 2MB or above. If you want vAttention with 64KB, 128KB or 256KB pages, you must follow the README in nvidia-vattn-uvm-driver and replace the default CUDA UVM driver with the project's custom one. The README states plainly that replacing CUDA drivers is not required if you use only 2MB pages. That leaves a real decision: run the stock driver and accept 2MB granularity, or take on a driver replacement to get finer pages. Driver replacement is a machine-level change that affects everything on the host, not just this workload, and the README does not document a rollback procedure. If you are on a shared cluster or a managed GPU instance, that alone may rule out the smaller page sizes. The 2MB path avoids the problem entirely and is the configuration to try first.

Where vAttention is the wrong tool

The integration is the limitation. vAttention is not a drop-in library for an existing serving system. To use it you run sarathi-lean, the modified Sarathi-Serve, which is a specific serving stack with its own scheduler and benchmark configuration. If your production system is built on a different engine, adopting vAttention means adopting that stack or porting the allocator yourself. The environment pins are equally strict: PyTorch 2.3.0 and CUDA 12.1, with the README conceding that later CUDA versions may or may not work. There are no releases listed for the repository, so you are building from the default branch rather than tracking versioned artifacts. The last push to that branch was on 2026-08-24. If your hardware is not A100 class, the README does not claim support, only an expectation that other Linux systems running the specified CUDA and PyTorch versions will work. And if your workload is decode-heavy rather than prefill-bound, the performance argument the README makes is weaker, since that is where it says the advantage is clearest.

vAttention compared with PagedAttention in vLLM

The comparison the project itself draws is with PagedAttention, the approach popularized by vLLM. The difference is architectural, not incremental. PagedAttention implements demand paging in user space: the cache is a set of fixed-size blocks, and custom kernels are written to gather KV entries across those blocks. vAttention pushes the paging down to the CUDA virtual memory layer: the cache stays contiguous to the kernel, and the driver maps physical pages in on demand. The practical consequence is that PagedAttention buys dynamic allocation at the cost of kernel complexity, while vAttention buys it at the cost of a driver-level dependency and a tighter environment. PagedAttention's block abstraction also makes memory sharing across sequences straightforward, which is why it became the default in that ecosystem; the README here does not discuss prefix sharing or cache reuse, so do not assume vAttention covers the same ground. If your kernels are already paged and working, the migration cost is the main thing to weigh, not the allocator.

Serving with the OpenAI compatible API

For benchmarking against external tools, the repository includes an OpenAI compatible API. The README notes it can be used with LLM benchmarking tools such as metron. The server starts from inside the sarathi-lean directory.

bash
cd sarathi-lean/
python -m sara

The README does not document the port the server binds to, the endpoint path, or any authentication, so check the code under sarathi-lean before pointing a load generator at it. This API exists to facilitate benchmarking rather than to serve production traffic, and the README frames it that way. The benchmark configuration knobs are listed in sarathi-lean/sarathi/benchmark/config/default.yml, with a detailed explanation in sarathi-lean/sarathi/benchmark/README.md. If you want to reproduce the paper's results rather than measure your own workload, the static trace script with its 50 requests at 32k, 64k and 128k context lengths and prefill to decode ratios of 500, 100 and 50 is the closer match.

Maintenance, licensing and upgrade cost

The repository is not archived, and the last push to the default branch was on 2026-08-24. There are no retrieved releases, so there is no version number to pin and no changelog to read before upgrading. The licence is MIT, which is permissive and imposes no copyleft obligation on your own code, though it also means no warranty. Two dependencies complicate the licence picture without being legal advice: the FlashAttention and FlashInfer packages the backends depend on carry their own licences, and the nvidia-vattn-uvm-driver directory contains a modified version of NVIDIA's UVM driver, which is not covered by this repository's MIT licence. If you intend to replace the system driver, review the terms attached to that driver separately. Upgrade cost is dominated by the version pins: moving to a newer PyTorch or CUDA means rebuilding the allocator against a matching libtorch and re-validating the kernels, and the README only claims testing for the 2.3.0 and 12.1 combination.

Editorial conclusion

Adopt vAttention if you serve LLMs on A100 class GPUs, your attention kernels are FlashAttention or FlashInfer based, and you want dynamic KV-cache allocation without rewriting kernels for paged memory. Do not adopt it if you are not on CUDA 12.1 with PyTorch 2.3.0, if your serving stack is vLLM and you cannot move to the bundled Sarathi-Serve, or if you need 64KB to 256KB pages and are unwilling to replace the CUDA UVM driver. Before committing, verify three things in your own environment: that the libtorch build you download matches the torch version exactly, that your kernels work with the attention_backends knob you intend to use, and whether the 2MB page path alone is acceptable, since that is the only configuration that avoids the custom driver.

Frequently asked questions

Does vAttention require rewriting attention kernels like PagedAttention does?

No. The README states that vAttention provides support for dynamic memory allocation to unmodified attention kernels by keeping the KV-cache contiguous in virtual memory, whereas PagedAttention implements demand paging in user space and requires rewriting custom kernels.

Which PyTorch and CUDA versions does vAttention need?

The README says using the repository requires PyTorch 2.3.0 and CUDA 12.1 or later, with the caveat that other CUDA versions may or may not work. It also states the project was tested with the Linux kernel, A100 GPUs and Python 3.10.

When do I need to replace the CUDA UVM driver for vAttention?

Only when you want page sizes of 64KB, 128KB or 256KB, because NVIDIA CUDA drivers allocate memory only at the granularity of large pages of 2MB or above. The README notes that replacing CUDA drivers is not required if you use vAttention with only 2MB pages.

Which attention backends can vAttention run with?

The README lists fa_paged_[block_size], fi_paged_[block_size], fa_vattn_[page_size], fi_vattn_[page_size] and their _sync variants, where fa denotes FlashAttention and fi denotes FlashInfer. It recommends block size 256 for FlashAttention and 16 for FlashInfer.

How do I start the OpenAI compatible API in vAttention?

The README gives the command as changing into the sarathi-lean directory and running python -m sara. It does not document the port, endpoint path or authentication, so check the code before using it with a benchmarking tool.

Official sources

  1. Issues
  2. License: MIT
  3. microsoft/vattention on GitHub
  4. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/microsoft-vattention.svg)](https://hysenlabs.com/projects/microsoft-vattention)