Model or dataset
lucienhuangfu/eLLM avatar
lucienhuangfu/eLLM

eLLM, a CPU inference engine that preallocates its KV cache and hopes you never fill it

eLLM: Run Long-Horizon Inference Faster on CPUs Than on GPUs

595 stars57 forksRustAGPL-3.0

At a glance

What is it?
eLLM 0.0.3 trades compute for DDR capacity to close the bandwidth gap against GPU memory. Here is what the static KV cache buys, what it costs, and how it is checked for correctness.
Who is it for?
eLLM fits a specific shape of work: an agent that keeps the same context across hundreds of turns, runs on hardware you already own, and would rather pay for DDR than for a GPU. Session continuity and full-prompt prefill are the two mechanisms doing that work, and the vLLM-compatible surface means it drops into existing tooling.
Can I use it commercially?
Yes, with strict conditions. AGPL-3.0 is a network copyleft licence: if people use a modified version over a network, for example as a hosted service, you must offer them its source code under the same licence.
Is it still maintained?
Yes. The repository last received commits 2 days ago.
What is it written in?
Mainly Rust, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The bet is storage over compute, aimed at multi-round agents

The whole thesis is one sentence: trade storage for computation. A CPU server has DDR memory measured in tens or hundreds of gigabytes while a GPU has high-bandwidth memory measured in tens, and that gap in bandwidth is an order of magnitude. Rather than fight it, eLLM spends the capacity. The target scenario is named explicitly as multi-round execution plus long-term state maintenance plus low-latency interaction, which is the profile of an agent that keeps goals and reasoning consistent across a long run. The four use cases follow from that: a computer-use agent that loads skills on demand, a code copilot working a large cross-module repository, retrieval-augmented generation that injects knowledge mid-task, and deep research lasting hours or days rather than one long prompt. All four share the assumption that the same context is used again.

The KV cache is a preallocated tensor, not a paged one

This is the decision everything else follows from. Instead of paged block management, eLLM preallocates a fixed-shape tensor for the KV cache and locates entries directly by tensor coordinate when reading and writing. Reads run contiguously along the sequence dimension. The claimed wins are the ones you would expect from removing an indirection: less metadata maintenance, less address mapping, less dynamic allocation, and fewer TLB and cache misses. The companion mechanism is massive-dimensional tensors, which reserve a sequence dimension large enough that the KV tensor is effectively unbounded and a full prefill never has to be repeated. The cost is implied by the design and stated in the launch command rather than in prose: you choose the sequence length and the memory is committed at startup.

An elastic static graph with a dimension-first layout

The second mechanism is about the graph rather than the cache. eLLM builds one globally unique static computation graph and lays tensors out dimension-first, so elements at the same logical coordinates map stably to the same memory address. The consequence is that the same graph serves different input lengths without being rebuilt, which is what lets an incremental prefill append to an existing sequence instead of reconstructing the plan. Put together with the session cache, which holds KV state across multi-turn interactions and prefills only the new input with no recomputation of history, you get the property the project calls state continuity at the mechanism level rather than at the application level. A prompt that grows over a hundred turns costs the compute for one turn plus the deltas.

Prefill is claimed at two orders of magnitude over other CPU frameworks

The performance claims are the project's own and should be read as such. Prefill is said to improve on existing CPU inference frameworks by roughly two orders of magnitude, and the mechanism behind the number is specific: a full single-pass prefill over the entire long prompt, with no chunking and no repeated parameter loading, followed by incremental prefills on new input only. Decode is argued differently. It runs at a smaller batch size, which activates fewer parameters and gives each request a larger share of memory bandwidth, so throughput per request can exceed a GPU. The stated trade is explicit about the cost: single-instance concurrency is lower than a GPU solution, and the argument that real QPS is higher rests on end-to-end latency being smaller rather than on the server serving more simultaneous callers.

Correctness is the pitch, and there are eight binaries that check it

The beta release is described as having complete core features with inference results fully aligned with SGLang CPU, and the repository is organised around proving that. An `alignment/` directory holds eight separate binaries: a rope alignment test, a silu-and-multiply alignment test, a Qwen3 tokenizer alignment binary, an RMS map alignment binary, one-token, multi-token, and multi-batch alignment binaries, and a multi-batch test. Each is declared as its own `[[bin]]` in the manifest rather than living behind a test flag, so you can run one against a reference and diff the tensors. The dependency list supports the story: `safetensors` for weights, `tiktoken-rs` for tokenisation, and `criterion` in the development dependencies for benchmarking. For a framework whose claim is that a CPU produces the same answer as a GPU, that is the right place to spend the effort.

The binary takes chunk size, sequence length, and batch size at launch

There is no configuration file in the documented path. Capacity is three flags:

bash
git clone https://github.com/lucienhuangfu/eLLM.git
cd eLLM
# Copy the downloaded model to models/Qwen3-Coder-30B-A3B-Instruct
cargo build --release --bin main
./target/release/main \
  --model-path models/Qwen3-Coder-30B-A3B-Instruct \
  --chunk-size 50000 \
  --sequence-length 50000 \
  --batch-size 1

The example runs at roughly fifty thousand tokens of capacity with a single request slot, which tells you the intended shape is one long-running session rather than a busy endpoint. You are expected to fetch Qwen3-Coder-30B-A3B-Instruct first and place it under `models/`, and the first startup is expected to be slow while weights and the computation graph initialise. Hardware floors are stated up front: a CPU with AVX-512 FP16 support and 128 GB of memory or more.

Fat LTO, one codegen unit, and no GPU crate in sight

The release profile is tuned to the maximum. Fat link-time optimisation, a single codegen unit, `panic = "abort"`, debug assertions off, overflow checks off, symbols stripped, unwind tables off, and incremental compilation off. That combination maximises steady-state throughput and costs you a slow build, usable stack traces, and arithmetic checks in release. The dependency list is the clearest statement of what this is: `axum` for the HTTP surface, `tokio` for async, `core_affinity` for pinning threads, `memmap2` for mapping weights, `minijinja` for templating, `raw-cpuid` as a build dependency for detecting CPU features at compile time, and no CUDA or accelerator library anywhere. The library target also builds as a `cdylib` alongside the `rlib`, so the engine can be linked into something else as well as run as a binary.

The manifest says 0.1.0 and the newest tag is v0.0.3

A small mismatch worth knowing: the released version is v0.0.3, from 2026-09-04, while the package manifest declares 0.1.0. The tag history is short and dated: v0.0.1 open-sourced the project on 2025-12-20, v0.0.2 was the alpha on 2026-04-06, and v0.0.3 was the beta in September. The branch was last pushed on 2026-09-30, and the project says code goes to the main branch monthly, which those dates fit. Model support is one line: the Qwen3 series is supported and Qwen3.8 is in development, so nothing else runs today. The licence is AGPL-3.0, which matters if you intend to modify it and serve the result, and the README is bilingual with a Chinese version alongside the English one.

Editorial conclusion

eLLM fits a specific shape of work: an agent that keeps the same context across hundreds of turns, runs on hardware you already own, and would rather pay for DDR than for a GPU. Session continuity and full-prompt prefill are the two mechanisms doing that work, and the vLLM-compatible surface means it drops into existing tooling. It is a poor fit for short prompts, for a bursty workload that needs many concurrent slots, or for a non-Qwen3 model today. Before you commit, read the batch size and sequence length in the launch command as a capacity budget, and confirm your CPU has AVX-512 FP16, because there is no fallback path described for one without it.

Frequently asked questions

What hardware does eLLM need?

A CPU with AVX-512 FP16 support and 128 GB of memory or more, running Linux on x86-64. Intel Xeon and AMD EPYC are the named processors with a Xeon Gen4 or newer recommendation, and the memory is expected to be sized to the model. Rust is installed with rustup, with the toolchain file already pinning nightly.

How is eLLM different from a GPU inference server?

It spends memory instead of bandwidth. The KV cache is a preallocated fixed-shape tensor rather than a paged one, prefills run as a single pass over the whole prompt with no chunking, and decode runs at a smaller batch so each request gets more of the available bandwidth. The stated trade is lower single-instance concurrency in exchange for lower end-to-end latency.

Which models does eLLM support?

The Qwen3 series is supported, and Qwen3.8 is listed as in development. The documented walkthrough uses Qwen3-Coder-30B-A3B-Instruct, downloaded from Hugging Face and placed under the models directory before building and launching the binary.

How does eLLM verify its results match a GPU?

The beta release is described as fully aligned with SGLang CPU, and the repository ships eight separate alignment binaries covering rope, the silu-and-multiply step, the Qwen3 tokenizer, the RMS map, and single-token, multi-token, and multi-batch cases. Each is its own binary target rather than a flag, so you can run one and compare tensors against a reference.

Is eLLM compatible with vLLM tooling?

Yes, the API is described as vLLM compatible so it plugs into the existing ecosystem. Inference results are stated to stay consistent with GPUs, which is the other compatibility claim, and the licence is AGPL-3.0.

Official sources

  1. Issues
  2. License: AGPL-3.0
  3. lucienhuangfu/eLLM on GitHub
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/lucienhuangfu-ellm.svg)](https://hysenlabs.com/projects/lucienhuangfu-ellm)