Luminal: an inference compiler built on fifteen ops and e-graph search
Inference at the speed of light.
At a glance
- What is it?
- The Rust project behind luminal-ai/luminal reduces a neural network to fifteen primitives, searches optimization decisions with egglog instead of hand-written rules, and reaches PyTorch through a compiler backend rather than a library.
- Who is it for?
- Luminal is most convincing as a statement about compiler design rather than as a drop-in replacement for your current inference stack. Reducing a whole network to fifteen primitives, refusing to destroy information during lowering, and letting e-graph search make the decisions are coherent choices that add up, and the Rust core small enough to read in an afternoon is what makes them verifiable rather than merely stated.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 20 days ago.
- What is it written in?
- Mainly Rust, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 20, 2026, and from our analysis. They are not legal advice.
Editorial analysis
Fifteen operations, and the argument for stopping there
Luminal is described as a high-performance general-purpose inference compiler, and its central architectural claim is a reduction. Everything in the language boils down to fifteen primitive ops, listed explicitly in the README as unary Log2, Exp2, Sin, Sqrt and Recip; binary Add, Mul, Mod and LessThan; and the remaining SumReduce, MaxReduce, Iota, Gather, Scatter and Cast. The claim that follows is that this set is enough to support transformers, convnets and nearly every popular model in the world.
That is a small surface by design, and the stated goal behind it is legibility: the core of Luminal should be understandable in an afternoon. The README ties this to a philosophy of search over heuristics, arguing that the best heuristic is no heuristic. Rather than encoding when fusion is profitable, the compiler tries possible decisions and lets the search find them.
The implementation detail that makes the search possible is in the dependency list rather than the prose. `Cargo.toml` pulls in `egglog`, `egglog-ast`, `egglog-reports` and `egraph-serialize`, all pinned to a specific revision of the egraphs-good repository rather than a published version. Equality saturation over an e-graph is what allows the compiler to consider many rewrites simultaneously instead of committing to each one in sequence, which is the direct answer to the failure mode the README attributes to other stacks.
The project is written in Rust, was pushed to on 2026-09-20, is not archived, and has 2984 stars and 232 forks with 46 open issues. Its stated description is simply inference at the speed of light, with a homepage at luminal.com and separate documentation at docs.luminalai.com.
Why destructive rewrite rules are treated as the enemy
The ideology section is the most direct writing in the README, and it explains the whole design in one argument. Most deep learning libraries are eager-first, meaning each operator call acts on data immediately, so `x + y` in PyTorch performs the addition right there. The README grants that this is great for debugging because it matches developer expectation, then argues it is bad for the machine, in the same way nobody writes assembly by hand.
The target of the criticism is specific rather than general. XLA, torch.compile, TVM and similar stacks are said to suffer from complexity explosion: they are built from a very large set of destructive, one-direction rewrite rules that lower a graph from a high-level representation to machine code. Because those rules are destructive, they can only fire when there is certainty of a performance benefit, which forces them to become complex, special-cased and numerous. Add more hardware backends, model architectures and dtypes, and the claim is that they often produce suboptimal code and require a DSL such as Pallas or Triton to recover performance.
Luminal's answer is ahead-of-time compilation as a core tenet: push everything to compile time and leave nothing to run time. Writing `x + y` records an operation in a directed acyclic computation graph and performs no computation, and the README addresses the obvious objection directly by calling this lazy execution and then noting that everything in Luminal is done this way. All networks are built as static graphs, compiled, and executed later.
Because the whole network is visible, the README argues devices, datatypes and even autograd can be modeled ahead of time and optimized by the compiler. The listed consequences are aggressive kernel fusion, shape-specific kernels compiled at runtime, low-precision dtypes such as mxfp4, nvfp4 and fp8, and multi-device parallelism topologies searched ahead of time.
Symbolic shapes as a way to stay static without being rigid
A fully static graph would be unworkable for real models, and the README acknowledges this rather than pretending otherwise, conceding that dynamism is necessary. The response is to model dynamic shapes natively as symbolic dimensions and support arbitrary symbolic dimensions including complex expressions.
The examples given are shapes like `(s, 4096)`, `(b, h, w + 3)`. That notation is the point: sequence length, batch size, head count and even an expression combining width with a constant are all representable. The stated benefit is that this rich representation gives the compiler full visibility into shapes while still allowing aggressive specialization, which is what lets generic source code compile into kernels specific to a particular architecture.
This is also where the connection to PyTorch becomes concrete. Luminal integrates directly with PyTorch as a compiler backend, so you call `torch.compile(model, backend=luminal_cuda)` to compile PyTorch models through it. The project also exposes a tensor API in Rust. Validation is claimed through tests written against equivalent PyTorch implementations, and the README links to an open issue tracking the improvements still needed there, which is a more credible signal than an unqualified correctness claim.
Correctness for that backend matters more than it would elsewhere, because a compiler backend that silently diverges from PyTorch numerics produces wrong models rather than an error. The existence of a linked tracking issue is worth noting precisely because it means the authors consider this unfinished.
The performance claim is narrow and should be read narrowly: the README says Luminal can run Q8 Llama 3 8B at about 80 percent of theoretical maximum performance on an H100, with the stated goal of becoming the fastest framework for any model on any device. The first half is a measurement with a named hardware target; the second is an ambition.
Workspace layout, examples and what the versions imply
`Cargo.toml` describes a workspace with edition 2024 and a rust-version floor of 1.85. Members include every example directory plus a set of crates: `luminal_nn`, `luminal_cuda_lite`, `luminal_metal`, `luminal_tracing`, `luminal_bench`, the Rust side of `luminal_python`, and `luminal_training`. That spread tells you the backends are separate crates, with Metal handled by its own and CUDA by a deliberately lite-named crate.
The examples directory is the clearest statement of intended scope. It holds `flux2`, `gemma`, `gemma4_moe`, `llama`, `paged_llama`, `qwen`, `qwen3_moe`, `simple`, `visualization`, `whisper` and `yolo_v11`. That covers autoregressive text models, mixture-of-experts variants, an image model, speech, object detection and a deliberately small entry point, and the presence of a paged llama example indicates attention to serving concerns rather than only training-shaped research.
Running one is a two-line affair, and the README gives Llama 3 8B on CUDA as the example:
cd ./examples/llama
cargo run --releaseNow the versioning, which is where the repository's own facts disagree in a useful way. The manifest says version 0.2.0, and the only functional release tags are 0.1 from August 2023 and 0.2 from March 2024. But the repository was pushed to on 2026-09-20 and the examples list architectures that did not exist in 2024. Read plainly, the tags are far behind main, so anyone pinning to a version is pinning to code more than two years older than what the repository currently contains.
The release list also contains something that is not a release at all: a `yolo-v11n` tag from 2026-05-04 whose body is a set of fused safetensors for the yolo_v11 example, generated by a script in that example, with a sha256 checksum. Model weights are being versioned through the same release mechanism as code, which is convenient for reproducibility and blurs what a release tag means here.
Licensing, the test story and what to trust
Two facts about licensing sit side by side and are worth separating. GitHub reports the repository licence as Apache-2.0. The crate manifest declares `license = "MIT OR Apache-2.0"`, which is the standard Rust dual-licence expression, and the tree contains both `LICENSE-MIT` and `LICENSE-APACHE`. The metadata field and the manifest do not agree in form, and the dual-licence declaration is the more specific statement, so read both licence files rather than relying on the label.
The Rust dependency choices add context. Dev dependencies include `candle-core` and `candle-nn` at 0.9.2, plus `proptest` for property testing, which is a sensible pairing for a compiler: property tests generate inputs and check invariants, while candle provides an independent implementation to compare against. There is also a patch section redirecting `candle-kernels` to a pinned revision of the candle repository, again a commit pin rather than a release.
Supporting infrastructure suggests a project with more discipline than the version numbers imply. The tree carries a `.pre-commit-config.yaml`, a `clippy.toml`, a `ci/` directory, a `.devcontainer/`, and both an `AGENTS.md` and a `spec.md` at the root. Documentation is not confined to the README: there is a `docs/` directory, and the logo and screenshots are stored under it.
So what should a reader take away? The architecture argument is coherent and the search-over-heuristics commitment is real, evidenced by the e-graph dependencies rather than by adjectives. The main reservations are the version numbering that lags the code, the correctness work the authors themselves flag as incomplete, and the fact that a single quoted performance number on one GPU model is a thin basis for planning. Read `spec.md` and the primitive op list first, then run the llama example and look at what the compiler produced.
Editorial conclusion
Luminal is most convincing as a statement about compiler design rather than as a drop-in replacement for your current inference stack. Reducing a whole network to fifteen primitives, refusing to destroy information during lowering, and letting e-graph search make the decisions are coherent choices that add up, and the Rust core small enough to read in an afternoon is what makes them verifiable rather than merely stated. Two practical warnings. The crate is still versioned 0.2.0 with the matching tag dated March 2024, while the repository itself was pushed to on 2026-09-20 and the examples directory lists recent architectures, so the tags badly understate how much has landed on main. And the licence is worth reading directly, because GitHub reports Apache-2.0 while the manifest declares MIT OR Apache-2.0 and the tree carries both licence files. Start by running the Llama example on CUDA, then read the primitive op list and decide whether the eager-versus-static trade is one your team can afford.
Frequently asked questions
What is Luminal used for?
It is a general-purpose inference compiler written in Rust. You build a network as a static computation graph, compile it ahead of time, and execute it, with the compiler searching optimization decisions through an egglog-based e-graph rather than relying on hand-written rewrite heuristics.
Can I use Luminal with PyTorch models?
Yes. Luminal integrates with PyTorch as a compiler backend, so you can pass backend=luminal_cuda to torch.compile to compile a PyTorch model through it. The README also states correctness is checked against equivalent PyTorch implementations, and it links an open issue tracking improvements still needed on that front.
Which model architectures does Luminal support?
The examples directory covers Llama, paged Llama, Qwen, Qwen3 MoE, Gemma, Gemma4 MoE, Flux2, Whisper, YOLO v11 and a simple example. The stated design claim is that fifteen primitive operations are enough for transformers, convnets and nearly every popular model.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/luminal-ai-luminal)