Model or dataset
rednote-machine-learning/RedKnot avatar
rednote-machine-learning/RedKnot

RedKnot: Head-Aware KV Reuse and SegPagedAttention for Long-Context LLM Inference

Efficient Long-Context LLM Serving with Head-Aware KV Reuse and SegPagedAttention

2,462 stars1,090 forksPythonApache-2.0

At a glance

What is it?
RedKnot is a Python extension to SGLang that reduces redundant computation in long-context LLM serving through three composable mechanisms: head decomposition and aggregation, sparse FFN and MoE execution, and SegPagedAttention. The DeepSeek V4 Flash TP8 path is the primary reproducible benchmark, verified on 8 H200 GPUs.
Who is it for?
RedKnot is a practical choice for teams that serve large language models at long context lengths and need to reduce time-to-first-token without changing the model weights. The DeepSeek V4 Flash TP8 path provides a reproducible baseline on 8xH200 hardware.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 3 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The Long-Context Inefficiency Problem RedKnot Addresses

Standard LLM serving systems recompute all attention at every step or cache the full KV state for every head. At long context lengths (64K tokens and above), this becomes expensive: attention heads that primarily attend to local context are treated the same as heads that attend globally, and every token is treated as equally important when selecting which feed-forward network rows to compute.

RedKnot's approach, as described in its arXiv paper (2606.06256), is to classify heads by behavior and exploit that structure. Heads that can be prepared offline are separated from heads that require online computation. Tokens that contribute little to the residual stream are excluded from expensive FFN rows. KV pages are organized per head and per segment rather than uniformly, so different head types consume different context scopes.

The README states the system is built around 'three composable ideas rather than one model-specific cache shortcut,' which means the mechanisms can be combined or applied independently depending on the model architecture and hardware configuration.

The Three Core Mechanisms in Detail

The first mechanism is head decomposition and aggregation. Attention heads are classified by their long-context behavior. Reusable local heads are prepared offline; global, retrieval, and recovery heads remain online. The README describes these projected contributions as 'merged back into the model without changing the model's external interface.' The same abstraction handles MLA, MHA, GQA, and native sliding-window attention, with model-specific projection and RoPE handling.

The second mechanism is sparse FFN and MoE execution. Token-level importance scores control which tokens enter expensive feed-forward work. Adaptive expert Top-K assigns additional experts only when the router distribution requires them. Dense boundary layers and protected query rows preserve the critical computation path. The README notes that this mechanism operates at the token and expert levels, targeting the FFN rows that have the most impact on output quality.

The third mechanism is SegPagedAttention. KV pages and their visibility are organized per head and per segment, allowing different head types to consume different context scopes without forcing a single uniform cache layout. This differs from standard PagedAttention, which assigns one cache layout to all heads regardless of their behavioral classification.

Performance Targets and the DeepSeek V4 Flash Benchmark

The README documents performance targets with explicit qualifications. On qualified long-context profiles, RedKnot targets quality regression within 1 percentage point, a 2 to 5 times TTFT speedup in the hot state, and 70 to 90 percent arithmetic compute savings. The README also clarifies what the compute ledger measures: it excludes memory traffic, kernel-launch cost, TP communication, and other runtime components, so it is not a claim about total system energy or end-to-end throughput.

The DeepSeek V4 Flash TP8 release is described as the primary reproducible path. The verified hardware configuration uses 8 NVIDIA H200 GPUs with 143,771 MiB per GPU, TP8, CPython 3.11.13, PyTorch 2.9.1 with CUDA 12.8, and Triton 3.5.1. Specific kernel versions are listed in the README: FlashMLA 1.0.0, SGL Kernel 0.3.20, and FlashInfer 0.5.3.

The TTFT protocol uses hot-state measurement: 3 unmeasured paired warmups followed by 10 measured Recomputed/RedKnot pairs per case, with p50 and p95 reported for streaming first output token. This methodology is documented to make the numbers reproducible rather than illustrative.

Getting Started: Repository Layout and Entry Points

RedKnot is built as an extension of SGLang. The README links to the SGLang project as its foundation (github.com/sgl-project/sglang). The repository's Python code lives in `python/sglang/srt/` and follows SGLang's module structure. This means setting up RedKnot requires a compatible SGLang installation alongside RedKnot's additional kernel and memory management components.

The `redknot` and `segpaged` attention backends are exposed as selectable options within the SGLang serving runtime. The README documents an experimental multi-request shared KV manager in `python/sglang/srt/mem_cache/head_kv/`. An integration and lifecycle guide is available at that path within the repository, along with a validation report at `test/srt/redknot/HEAD_KV_VALIDATION.md`.

The DeepSeek V4 Flash release ships with frozen inputs, head policy, sparse-MoE policy, and execution manifests required for the packaged benchmark. The README describes this as 'one-command reproduction' over the four frozen RAG suite lengths (64K, 128K, 256K, and 440K tokens), though the specific command is located in the benchmark and examples directories rather than in the main README.

Benchmark entry points are in the `benchmark/` directory, and model-specific examples are in `examples/redknot/`. The `docker/` directory provides container configurations for running the system without manual dependency resolution.

Experimental Features and Current Limitations

The multi-request shared KV backend is explicitly marked experimental in the README. It supports MHA/GQA request forks, copy-on-write append/repair, and cross-process snapshot sharing over HTTP. The README is direct about its limitations: it does not implement MLA, full scheduler and TP serving integration, or automatic dense-pool replacement. Strict Qwen3-8B FP32 qualification passes, but BF16 model qualification remains unresolved. The README states explicitly that 'the component capacity results do not establish an end-to-end serving speedup.'

The Ascend NPU adaptation is driven by Huawei Cloud. Short-term, the goal is functional parity with the upstream SGLang NPU baseline in RedKnot's Recomputed reference path. Medium-term, the plan includes publishing Ascend-side qualification profiles for the four frozen suite lengths. The current status and known gaps are tracked in `docs/ASCEND.md`, which teams planning an Ascend deployment should read before starting.

The laboratory model adapters for Mistral, Qwen3, Qwen3.5 MoE, and Llama 3.3, released in July 2026, are described as experimental. The README notes that they cover native SWA, GQA/MHA head policies, and sparse-FFN execution for those architectures, but they are not in the same verified state as the DeepSeek V4 Flash TP8 path.

Comparison with Standard vLLM and Maintenance Status

vLLM is the most widely deployed open-source LLM serving framework and uses PagedAttention as its memory management approach. PagedAttention treats all KV cache pages uniformly and does not distinguish between heads based on their long-context behavior. For context lengths in the tens of thousands of tokens, this uniform treatment means all heads pay the full cache cost regardless of whether they could be served with reused local state.

RedKnot's head decomposition mechanism specifically targets this inefficiency: local heads that could be precomputed are separated from global and retrieval heads. The cost of this approach is implementation complexity: each new model requires a head policy that classifies heads and specifies how to handle RoPE relocation and projection merge. The README documents these requirements for DeepSeek V4 Flash and the lab adapters, but teams adopting a model without an existing adapter will need to write one.

A RedKnot-vLLM branch exists in the repository, indicating that the core algorithm is being ported to vLLM's serving stack as a plugin. The README notes that this migration completed in September 2026 and that model-specific runtime integration and end-to-end validation are ongoing.

The last push was on 2026-09-14. The repository has no GitHub releases; the code is developed and distributed directly from the main branch.

Editorial conclusion

RedKnot is a practical choice for teams that serve large language models at long context lengths and need to reduce time-to-first-token without changing the model weights. The DeepSeek V4 Flash TP8 path provides a reproducible baseline on 8xH200 hardware. Teams working with other hardware or models should consult the lab adapters for Mistral, Qwen3, Qwen3.5 MoE, and Llama 3.3, with the understanding that those paths are described as experimental. The multi-request shared KV backend is explicitly experimental, and the README states that it does not establish an end-to-end serving speedup on its own. The Ascend NPU adaptation is in progress; check docs/ASCEND.md for the current status before planning a deployment on Huawei Cloud.

Frequently asked questions

Does RedKnot require specific GPU hardware to run?

The verified DeepSeek V4 Flash TP8 release used 8 NVIDIA H200 GPUs with specific driver, CUDA, and kernel versions documented in the README. The lab-model adapters for Mistral, Qwen3, Qwen3.5 MoE, and Llama 3.3 are described as experimental and may require different configurations. An Ascend NPU adaptation is in progress; the README references docs/ASCEND.md for current status.

What is the difference between RedKnot's SegPagedAttention and standard PagedAttention?

Standard PagedAttention manages KV cache pages uniformly across all attention heads. SegPagedAttention organizes KV pages and visibility per head and per segment, allowing heads classified as global, local, or retrieval to consume different context scopes without a single shared cache layout. The README states this is one of three composable mechanisms rather than a standalone replacement.

Is the multi-request shared KV backend production-ready?

The README explicitly marks the multi-request shared KV backend as experimental. It notes that this component does not implement MLA, full scheduler and TP serving integration, or automatic dense-pool replacement. BF16 model qualification is unresolved. The README states that component capacity results do not establish an end-to-end serving speedup.

Official sources

  1. Issues
  2. License: Apache-2.0
  3. README
  4. rednote-machine-learning/RedKnot on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/rednote-machine-learning-redknot.svg)](https://hysenlabs.com/projects/rednote-machine-learning-redknot)