Open-source project
FutureMLS-Lab/OSCAR avatar
FutureMLS-Lab/OSCAR

OSCAR: 2-bit KV Cache Quantization for SGLang and llama.cpp

OSCAR: Offline Spectral Covariance-Aware Rotation for 2-bit KV Cache Quantization

559 stars84 forksPythonMIT

At a glance

What is it?
OSCAR fits per-layer rotations offline from a small calibration set so the KV cache can sit in INT2 while a BF16 sink and recent window stay in full precision. It targets engineers serving long-context models who are limited by KV memory rather than weights.
Who is it for?
Adopt OSCAR if you serve long-context models on SGLang main or the zhongzhu/llamacpp fork and your bottleneck is KV memory, not weight memory: the repository ships a rotation zoo plus prebuilt GGUF files for Qwen3-32B, Gemma 4 12B and Qwen3-4B-Thinking-2507, so you can try it without calibrating anything yourself.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 1 day ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What OSCAR compresses, and why the KV cache is the right target

Weights are usually the first thing people quantize, but in long-context serving the KV cache is what actually sets the memory ceiling. Every token in the context keeps a key and a value vector per layer, and those vectors grow linearly with sequence length while the model weights stay fixed. OSCAR attacks that side of the problem: the README states it stores the bulk of the KV cache in INT2 while retaining only a small BF16 sink and recent window, which it describes as roughly an 8x reduction compared with BF16. The intended user is someone running long-context inference on a single machine or a small number of GPUs, where the alternative is not a bigger cluster but a shorter context. The repository is Python, MIT licensed, and the README says OSCAR is built directly into SGLang on main and into a llama.cpp fork on the zhongzhu/llamacpp branch.

How the offline rotation is fit from Q/K/V covariance

The mechanism is a calibration step followed by a serving step, and the split matters. During calibration OSCAR captures Q, K and V activations on a small calibration set, then estimates what the README calls attention-aware K/V covariance structures offline. From those covariances it derives two things per layer: a rotation and clipping thresholds. The rotation changes the basis the KV vectors are expressed in, so that the directions attention actually consumes line up with the axes quantization treats well; the clipping thresholds bound the values before they are rounded. This is why the project name says spectral covariance-aware: the rotation is not a fixed Hadamard transform, it is fitted to measured activation statistics. The output is a per-layer rotation file. Because the fit happens offline, serving does not pay for it, and the repository provides a rotation zoo on Hugging Face so you can download calibrated rotations instead of recomputing them. The README also points to a fused mixed-precision Flash-Attention kernel for the Apple Metal path, which is what makes the mixed INT2 plus BF16 layout practical rather than a dequantize-then-attend detour.

Installing OSCAR and running a first Qwen3-8B calibration

There is no standalone pip package described in the README. OSCAR is used through a host framework, and the repository layout reflects that: the top level holds rotation/, sglang-dump-qkv/, sglang-research/ and third_party/, with the framework integration living in the SGLang tree and in the llama.cpp fork. The README's setup and quick start sections use a Qwen3-8B example. The first real use is the calibration pass, which is what sglang-dump-qkv/ exists for: it captures the Q/K/V activations that the covariance estimate consumes. Because the README does not print the exact invocation, treat the directory as the entry point and read its contents before running anything.

bash
ls sglang-dump-qkv/

That listing tells you which script performs the dump for your framework version. Once the activations exist, the rotation fit consumes them and writes a per-layer rotation, which you then point the server at. The README's serving section is titled Serving with the rotation, and the calibration knobs section collects the parameters that change the fit. If you would rather skip calibration entirely, the faster path is the rotation zoo: the README links a Hugging Face repository of already-fitted rotations, and separately links prebuilt GGUF files for Qwen3-32B, Gemma 4 12B and Qwen3-4B-Thinking-2507 that already carry the INT2 KV layout for the llama.cpp fork. For the hybrid-model branch the README gives one concrete setting to use:

bash
export SGLANG_LLOYD_MAX=1

That variable is documented in the release note for qwen3.5 and minimax-m2.7 preview support, where it is required alongside the zhongzhu/hybrid-model branch. Note the branch names in the README: Gemma 4 12B SGLang support is on zhongzhu/gemma4-12b, multimodal work is on zhongzhu/VL, and the llama.cpp integration is on zhongzhu/llamacpp. Building main will not give you those.

Where OSCAR loses, and the cases it is the wrong tool for

The published numbers are not uniformly better than BF16, and the README does not hide it. On MiniMax2.7 with LM_RATIO=1.16, GPQA-Diamond and HumanEval improve, AIME 2025 is unchanged, and MATH500 drops from 0.9379 to 0.9279. On Qwen3.5-4B, GPQA falls half a point. So the honest position is that INT2 KV is close to BF16 on the benchmarks shown, not identical, and the direction of the delta depends on the task. A second limitation is structural: OSCAR keeps a BF16 sink and a recent window in full precision. That is a deliberate design choice, and it means the 8x figure is a property of the whole layout rather than of every token in the cache. Short-context workloads, where the KV cache was never the constraint, get the mixed-precision complexity without much memory back. Third, the calibration is offline and tied to a model: a rotation fitted for one checkpoint is not transferable, which is exactly why the project publishes a zoo instead of a single artifact. Finally, the README's own news entries describe vLLM support as a PR in progress and describe MiniMax 3 and GLM 5.2 testing as upcoming, so those paths are not something you can adopt today.

OSCAR against KIVI, QuaRot and RotateKV on the same benchmark

The README's OCRBench table is the most useful comparison in the repository because it puts several INT2 methods on the same two checkpoints. KIVI reaches 851 on Qwen3-VL-8B and 813 on Qwen3-VL-4B; RotateKV reaches 754 and 638; QuaRot reaches 722 and 773; OTT reaches 850 and 831; OSCAR with Lloyd-Max reaches 854 and 848, against a 16-bit baseline of 858 and 852. The difference in approach is the interesting part. QuaRot and RotateKV apply rotations that are not fitted to the specific model's attention statistics, which is a reasonable trade because it removes the calibration step entirely; the table suggests that on the 4B checkpoint that trade costs a lot. KIVI quantizes per-channel along a fixed axis without a learned rotation. OSCAR's position is that the rotation should be derived from measured K/V covariance, and it pays for that with an offline calibration pass and a per-model artifact. If you cannot run calibration, a rotation-free method is the pragmatic choice; if you can, the table is the argument for spending the compute once.

Maintenance, licence and what an upgrade actually costs you

The repository is not archived and the last push was on 2026-08-30, so it is recent work rather than an abandoned experiment. The licence is MIT, which is permissive and places few obligations on redistribution; the README has a License and acknowledgements section, and since OSCAR is integrated into SGLang and a llama.cpp fork, those upstream projects carry their own licences and you should read them before shipping a combined build. This is not legal advice. The upgrade cost is the part worth thinking about before you start. OSCAR tracks SGLang main and a named fork branch, so your upgrade path is bound to those trees rather than to a versioned release: the README lists no releases. Branch names in the README are model-specific (zhongzhu/gemma4-12b, zhongzhu/VL, zhongzhu/hybrid-model, zhongzhu/llamacpp), which means a model addition can arrive on a branch that is not the one you built. And because the rotation is fitted per model, moving to a new checkpoint means either finding it in the rotation zoo or re-running the calibration dump and fit. Budget for that step every time you change base model, not just when you upgrade the framework.

Editorial conclusion

Adopt OSCAR if you serve long-context models on SGLang main or the zhongzhu/llamacpp fork and your bottleneck is KV memory, not weight memory: the repository ships a rotation zoo plus prebuilt GGUF files for Qwen3-32B, Gemma 4 12B and Qwen3-4B-Thinking-2507, so you can try it without calibrating anything yourself. Do not adopt it if you need the vLLM path today, since the README describes that as a PR in progress, or if you cannot tolerate any benchmark regression: the published MiniMax2.7 table shows MATH500 at 0.9279 against 0.9379 for BF16. Before committing, verify that your model appears in the all-configured-models list and check the branch you intend to build, because Gemma 4 12B SGLang support lives on zhongzhu/gemma4-12b and multimodal support on zhongzhu/VL rather than on main.

Frequently asked questions

What does OSCAR stand for in this project?

The README expands it as Offline Spectral Covariance-Aware Rotation, and the subtitle adds that it is for 2-bit KV cache quantization. The name describes the method: the rotation is fitted offline from spectral covariance statistics of Q/K/V activations.

How much KV cache memory does OSCAR save compared with BF16?

The README states that storing the bulk of the KV cache in INT2 while keeping a small BF16 sink and recent window reduces KV-cache memory by approximately 8x compared with BF16. The 8x applies to the overall layout, not to every token, since the sink and recent window stay in full precision.

Which frameworks does OSCAR run on?

The README says OSCAR is built directly into the open-source SGLang framework on main and into llama.cpp on the zhongzhu/llamacpp branch. It also describes vLLM support as a PR in progress, so that path is not available yet.

Do I have to calibrate a rotation myself to use OSCAR?

No. The project publishes a rotation zoo on Hugging Face so users can download calibrated rotations instead of recomputing them, and it also publishes prebuilt GGUF files with the INT2 KV layout for several models. Calibrating yourself is the alternative when your checkpoint is not in the zoo.

Official sources

  1. FutureMLS-Lab/OSCAR on GitHub
  2. Issues
  3. License: MIT
  4. Project website
  5. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/futuremls-lab-oscar.svg)](https://hysenlabs.com/projects/futuremls-lab-oscar)