# vLLM-Moet: a 2-bit expert patch and a three-tier expert cache on SM120

> vLLM-Moet patches vLLM v0.24.0 for SM120 Blackwell cards, with sign-symmetric 2-bit routed experts, an FP4 delta cache that repairs precision, and a residency ladder that runs down from VRAM to an NVMe pack file. It reports 105 tok/s for GLM-5.2 on four RTX PRO 6000s, 28.3 tok/s on two, and 161 tok/s for DeepSeek-V4-Flash on a single 96 GB card.

**kacper-daftcode/vLLM-Moet** — A vLLM patch + hand‑written SM120 SASS kernels: 2‑bit MoE experts + an FP4 "delta" cache that recovers precision — matching the official (NV)FP4 checkpoint's quality on consumer Blackwell cards

- Repository: https://github.com/kacper-daftcode/vLLM-Moet
- Stars: 542 · Forks: 49
- Language: Sass
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/kacper-daftcode-vllm-moet

## vLLM v0.24.0 does not run on SM120, so the patch is the foundation

vLLM-Moet is not a fork you compile from scratch in the usual sense. It starts from the official vLLM v0.24.0 release, applies a generated runtime patch, and adds hand-written SM120 SASS kernels, which is why the repository's primary language reads as Sass. The reason for the patch is blunt: vLLM v0.24.0 is broken as shipped on SM120, the consumer and workstation Blackwell generation, so nothing upstream runs before this work happens. The layout says how the work is organized. patch/ and kernels/ hold the port, bench/ holds the measurement harnesses, docker/ and Dockerfile.recipes hold image assembly, tools/ holds helpers, docs/ holds a port write-up and a deploy guide, and AGENTS.md sits next to them for coding agents. Four model-specific images exist: Dockerfile.sm120-v024, Dockerfile.sm120-dsv41, Dockerfile.sm120-dsv41-nightly and Dockerfile.sm120-qwen38. Two frontier models skip the patch entirely and run on official checkpoints and official vLLM images on 4x RTX PRO 6000: DeepSeek-V4.1-Flash at 552B with vision and a 512K context, and Qwen3.8-Flash-Next-FP8 at 177B-A6B, where the decode step is rebuilt around small hand-written kernels for 243 and 346 tok/s, three times the stock image. The project sits under Apache-2.0 with 541 stars, 49 forks and 20 open issues, and a last push dated 2026-10-01.

## Sign symmetry, not error size, is what collapses naive 2-bit experts

Only the routed experts get compressed. The dense stack keeps the checkpoint's own precision, FP8 on DeepSeek-V4-Flash and NVFP4 on GLM-5.2, so quantization damage is confined to the part of the model that routing already treats as optional. The load path re-quantizes modelopt NVFP4 experts, e2m1 elements with e4m3 block-16 scales and a per-tensor scale_2, into 2-bit planes, and the GLM-5.2 loader reports f64-exact agreement with the reference pipeline on real shards from the 433 GB nvidia/GLM-5.2-NVFP4 checkpoint. The interesting part is the codebook. Plain 2-bit quantization of these models collapses into degenerate loops, and the cause is not error magnitude but sign asymmetry: the optimal-L2 two-bit codebook drops one sign's tail, and the resulting per-expert bias compounds across dozens of layers. Forcing a sign-symmetric codebook of -4, -1, 1 and 4 at the same L2 error removes the collapse. 33,023 of 33,024 DeepSeek-V4-Flash tensors select the symmetric set, and MTP acceptance lands at or above the FP4 experts, 2.73 against 2.68 in the QUANT_PROBE study. A 180-tensor sweep on GLM-5.2 measures the asymmetry directly: a bias of -0.042 that is negative 99% of the time, with the symmetric codebook 392 times smaller at equal relative RMS.

## Three residency tiers: NVMe pack file, pinned arena, GPU expert cache

When the 2-bit planes are still too large, they stop being a model and become a residency problem. For GLM-5.2 those planes come to roughly 190 GiB, which on their own eat the entire VRAM budget of two RTX PRO 6000s. vLLM-Moet answers with three tiers. An NVMe pack file holds the 2-bit base, a 57 GiB per rank pinned arena keeps the hot slice in host RAM, and the GPU holds a 46 GiB per rank expert cache rather than a full copy of the weights. A small FP4 pool sits on top, filled by the confidence gate, and the expert stores end up needing about 136 GiB of host RAM instead of the 568 GiB they would occupy pinned in full. The single RTX 5090 path for DeepSeek-V4-Flash runs the same shape at smaller scale, a 14 GiB pool plus NVMe stores and roughly 30 GiB of host RAM. The window is set by hand with a max model length of 131072 and 8 GiB per rank of KV, which measures out at 157K tokens of KV, and needle retrieval passes 4 out of 4 at 36K, 86K and 121K prompt tokens on fp8 KV with tol=0.

## A cache miss costs a replay, and the packs survive a reboot

A miss in the expert cache is not an exception path. It is a batched fetch followed by a bit-identical graph replay, and that is the mechanism holding a two-card GLM-5.2 window together. The pack files double as a persistent quantization cache, so a second boot reuses the quantization instead of redoing it. The difference shows up in wall time. Booting from existing quantization packs takes about 7 minutes, against about 11 minutes for a full re-quantizing load, and that load also stages roughly 405 GiB of transients while it runs. Two costs follow from the same design. A miss is a real latency event, and under batching its cost is spread across whichever streams touched the missing expert. The packs are also disk-resident state the server depends on, so a healthy NVMe device and enough pinned RAM stop being optional rather than a tuning choice. One thing stays undocumented: the section explaining FP4 recovery stops mid-sentence partway through its argument about decode bandwidth, so the confidence gate threshold and the delta cache sizing are not spelled out in the text quoted here.

## Bare 2-bit answers confidently and wrongly, and the gate costs twice

Compression this aggressive has a failure mode you notice before you benchmark anything. On bare 2-bit experts the model answers confidently and wrongly: ask it for the capital of Poland and Krakow comes back, and Polish text comes back garbled. The FP4 tier is what repairs those answers, and the confidence gate decides where the repair is spent. The price is published. On 4x RTX PRO 6000 with tensor parallelism 4 and MTP k=2, GLM-5.2 runs at 105 tok/s with a served 256K window; add the FP4 delta cache and the confidence gate and the same box falls to 83 to 85 tok/s while the served window falls to 128K. The reason is bandwidth rather than overhead: decode is HBM-bound, and an FP4 read moves twice the bytes of a 2-bit read. So the precision repair is paid for twice, once in tokens per second and once in how much context you can hold at the same time. The arithmetic and retrieval probes stay clean at both speeds, with strict tol=0 decode at 28.3 tok/s on the two-card layout and 31.7 tok/s once miss tolerance is set to 8.

## MTP speculative decoding roughly doubles pipeline-parallel throughput, bit for bit

Multi-token prediction does more work here than it usually gets credit for. With k=2 drafts, acceptance on 4x RTX PRO 6000 runs between 2.3 and 2.8 accepted tokens per step for GLM-5.2, and around 2.6 across the DeepSeek-V4-Flash configurations. Those numbers matter because they are measured against 2.73 from FP4 experts, the bar that makes the 2-bit path defensible: compressing the experts does not buy speed by surrendering acceptance. The more interesting half sits under pipeline parallelism, where most speculative decoders get messy. vLLM-Moet propagates drafts and shares the drafter embedding across ranks, and DeepSeek-V4-Flash on 4x RTX 5090 in a PP4 layout reaches 184 tok/s against 93 tok/s without it, about a factor of two. Greedy decode under pipeline parallelism is also bit-deterministic: 6 out of 6 identical runs, with and without MTP enabled. That matters more than the throughput number when you compare outputs across a fleet, because a serving stack that changes its answer when you change the parallel layout cannot serve as a fixed reference.

## NVFP4 KV cache buys 38% more pool, and still misses GLM's 1M window

The KV cache is where a frontier MoE model actually runs out of room, and vLLM-Moet adds a second packing format for it. The nvfp4 cache dtype packs 352 bytes per token against 656 bytes for fp8_ds_mla, which is a 38% larger pool at equal settings: 415K tokens of KV become 571K, at decode parity. Alternatively the freed VRAM can go to the FP4 expert pool instead, which is what the standing 4-card GLM-5.2 configuration does, running a 19.6 GiB per GPU pool next to a 175K-token KV. The numbers still fall short of the headline context figures. GLM-5.2's nominal 1M window is KV-bound on four cards, with needle retrieval passing to 126K on the nvfp4 cache and to 276K on fp8, where a 331K window fits at utilization 0.95. DeepSeek-V4-Flash at 159B has more room: needle retrieval passes at 453K on one RTX PRO 6000 with 947K tokens of KV measured, but only at 29.7K on a single RTX 5090 carrying a 131K-token KV. Same model, same weights, different card, a large gap in reachable context.

## Four RTX 5090s match two PRO 6000s once streams are batched

Per-stream numbers understate what these boxes do when several requests share them. Aggregate decode throughput for DeepSeek-V4-Flash at rising concurrency on a single RTX PRO 6000 goes 156, 290, 493, 659 and 933 tok/s across 1, 4, 8, 16 and 32 streams, with 29 tok/s per stream at 32. Four RTX 5090s under tensor parallelism 4 go 198, 460, 762, 1006 and 1560 tok/s, with 49 tok/s per stream at 32, so four consumer cards match two PRO 6000s on decode. Prefill behaves differently. GLM-5.2 on 4x RTX PRO 6000 sits near 2.5k tok/s, while DeepSeek-V4-Flash on an uncached 8k-token prompt measures 5340 tok/s on one PRO 6000, 5790 on two, 6100 on 4x 5090 and roughly 400 to 540 on a single 5090 where the NVMe tier sits in the path. Tool calling and reasoning are wired for agents rather than for benchmarks, with glm47 and glm45 parsers, and the endpoint drives coding agents such as opencode directly. Methodology for all of it lives in docs/v024-port.md.

## Conclusion

Judge vLLM-Moet by the cards you already own rather than by the model names in its README. It earns its place on a two-card or single-5090 box where the official checkpoints do not fit at all, and it loses on a four-card host with working stock vLLM, where the FP4 gate costs a fifth of the throughput and a half of the context window. Check three things before committing: that your own vLLM v0.24.0 image is broken on SM120 the way this one is, that you can host the pinned arena and roughly 140 GiB of RAM for the two-card GLM-5.2 path, and that you can work with a codebase whose serving base is a generated patch and which publishes no GitHub releases to pin to.

## FAQ

### What does vLLM-Moet actually do to vLLM?

It applies a generated runtime patch to the official vLLM v0.24.0 release and adds hand-written SM120 SASS kernels, because that release is broken as shipped on SM120 Blackwell cards. On that base it serves GLM-5.2, DeepSeek-V4-Flash and Kimi-K2.7-Code.

### Which GPUs does vLLM-Moet target?

The kernels are written for SM120, which covers RTX PRO 6000 and RTX 5090 cards. Published configurations run from one RTX 5090 at about 31 tok/s for DeepSeek-V4-Flash up to 4x RTX PRO 6000 at 105 tok/s for GLM-5.2.

### Do I need an NVMe drive and pinned host RAM for vLLM-Moet?

Only when the model overflows VRAM. The two-card GLM-5.2 layout uses an NVMe pack file, a 57 GiB per rank pinned arena and roughly 140 GiB of host RAM, while the all-VRAM setups on one or four cards report no host expert store at all.

### How much precision does vLLM-Moet give up?

Only the routed experts are compressed to 2 bits, on a sign-symmetric codebook of -4, -1, 1 and 4, and an FP4 delta cache with a confidence gate repairs what bare 2-bit gets wrong. On 4x RTX PRO 6000 that repair costs throughput, taking GLM-5.2 from 105 tok/s to 83 to 85 tok/s.

### Does vLLM-Moet publish releases or results for other models?

The repository publishes no GitHub releases, so there is no tagged version to pin to beyond the Dockerfiles it ships. DeepSeek-V4.1-Flash at 552B and Qwen3.8-Flash-Next-FP8 at 177B-A6B are reported on 4x RTX PRO 6000 from official checkpoints and official vLLM images.

## Sources

- [Issues](https://github.com/kacper-daftcode/vLLM-Moet/issues)
- [kacper-daftcode/vLLM-Moet on GitHub](https://github.com/kacper-daftcode/vLLM-Moet)
- [License: Apache-2.0](https://github.com/kacper-daftcode/vLLM-Moet/blob/main/LICENSE)
- [README](https://github.com/kacper-daftcode/vLLM-Moet/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/kacper-daftcode-vllm-moet
