XGrammar: constrained decoding that makes malformed JSON impossible
Fast, Flexible and Portable Structured Generation
At a glance
- What is it?
- A C++ library that sits inside a model's sampling loop and rejects any token that would break your schema, so structured output is correct by construction rather than by retrying. It is the default backend for vLLM, SGLang and TensorRT-LLM, and the release notes are a better guide to its current state than the README.
- Who is it for?
- XGrammar solves a problem that post-hoc validation cannot: a JSON parser can tell you the output was wrong, and only a decoding-time constraint makes it right. That is why it became the default structured generation backend for most major inference engines, and the near-zero overhead claim in the README is the number to check against your own workload before committing to it.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository received new commits within the last day.
- What is it written in?
- Mainly C++, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 7, 2026, and from our analysis. They are not legal advice.
Editorial analysis
Installing from PyPI and importing the module
The install is one line, and there is a second one for Apple Silicon:
pip install xgrammarpip install "xgrammar[metal]"The extra exists for MPS support on Apple hardware. On any other platform the base package is what you want, and the import is short enough to sit at the top of any file you are testing:
import xgrammar as xgrEverything else is in the documentation at xgrammar.mlc.ai/docs, with separate pages for installation and a quick start. That is the handoff point to be aware of: the README shows you how to get the module but nothing about how to attach a grammar to a sampler, so the quick start is a required read rather than an optional one.
The dependency list in `pyproject.toml` is worth a look before you install, because it explains the shape of the project. `torch>=1.10.0` is a hard dependency, alongside `apache-tvm-ffi`, `pydantic`, `numpy` and `transformers>=4.38.0`. A Triton dependency is declared conditionally on Linux x86-64. Python support is declared as `>=3.8, <4`, and the package classifies itself as Beta.
Those two facts together mean this is not a dependency-light library. Bringing XGrammar in means pulling in a specific region of the PyTorch and Transformers stack, which is fine if you already run inference and occasionally awkward if you wanted a parser.
How the token mask is built before each step
The mechanism is constrained decoding, and the promise is that output is 100% structurally correct rather than usually correct. Instead of generating text and checking it afterwards, XGrammar computes, at every sampling step, the set of tokens that are legal given the grammar position, and hands that as a mask to the sampler. Any token that would break the structure never gets a chance to be picked.
The grammar types supported are general context-free grammars, which is the important technical choice rather than a narrow JSON validator. JSON, regular expressions and custom context-free grammars all reduce to that framework, so the same engine handles a strict response schema, a loose pattern match and a bespoke DSL.
Performance comes from how that per-step set is computed. The v0.2.6 notes name the optimisations: direct JSON Schema AST construction, shared FSM construction across the grammar, in-place optimizer passes, and cached parser-state properties. The README's headline claim is near-zero overhead in JSON generation.
The release notes are where this project explains itself best, and the v0.2.7 entry shows the current edge of the work. It added the DeepSeek V4.1 structural tag and the `deepseek_v4_1_xml` schema style for that model's spaced tag format, exposed through Python, C++ and TypeScript, aligned with the official encoder and tokenizer. Two specific bugs fixed alongside it are instructive about the hard cases: correlating each parameter's `string="true|false"` attribute with its value type across unions, mixed enums, references and dynamic properties, and memoizing `$ref` parameter rules so recursive references terminate rather than recursing forever.
That recursion detail is the real signal. Supporting tool-call formats means handling schemas that refer to themselves, and getting that wrong is a stack overflow rather than a wrong answer.
One library, four language bindings, several backends
The deployment claim is about reach rather than features. XGrammar targets Linux, macOS and Windows; CPU, NVIDIA GPU, AMD GPU, Apple Silicon and TPU; and exposes Python, C++, JavaScript and Swift APIs. The repository tree backs the C++ and Swift claims with `cpp/`, `include/` and a `Package.swift` at the root, while `python/` holds the Python binding and `web/` suggests the JavaScript side.
The `3rdparty/` directory and `.gitmodules` file indicate vendored dependencies managed as submodules, and `cmake/` plus `CMakeLists.txt` show a conventional CMake build rather than a pure pip package. There is also `scripts/`, `tests/` and a `site/` directory for the documentation build.
Language coverage is uneven in practice, and it is worth noticing that the README documents only the Python path. The JavaScript and Swift APIs are listed as supported without a documented quick start, so an application team starting from either would be reading the API headers.
Third-party bindings exist for Rust, maintained as community bindings rather than in this repository. That distinction matters if you are assessing how much of the API surface is stable, since a community binding tracks a C++ API that can still move between 0.x versions.
The project is Apache-2.0 licensed, confirmed both by GitHub and by the LICENSE file at the root, with a separate NOTICE file alongside it.
Being the default backend is both the strength and the constraint
The adoption list is the strongest signal about the library's quality. XGrammar is the default structured generation backend for vLLM, SGLang, TensorRT-LLM and MLC-LLM, and the README names integrations with OpenVINO GenAI and Modular's MAX, plus adoption inside Mirai. Company logos at the bottom of the README include xAI, DeepSeek, NVIDIA and Databricks.
There is a specific reason for the vLLM integration that matters for how you think about it. vLLM is the most widely used open inference server, so XGrammar's structured output reaches anyone running vLLM without choosing it. That makes XGrammar less a library you adopt than an infrastructure component you inherit, which cuts both ways. You get it for free in vLLM, and you also inherit its version, its constraints and its release cadence whether or not you read anything about it.
The second thing to weigh is what a 100% structural correctness guarantee does and does not cover. XGrammar guarantees the output parses as your grammar. It does not guarantee the values are true, appropriate, or that a required field is semantically filled in rather than structurally present. If your schema allows a string in a field that should be an identifier, the grammar will happily enforce that it is a string.
For most use cases that separation is comfortable, because the schema is your contract and validation of meaning happens after parsing. It matters if you are about to treat guaranteed-parseable output as guaranteed-correct output.
Reading the release notes instead of the version number
The versioning tells you where this project is. The newest release is `v0.2.7` from 2026-09-15, after `v0.2.6` on 2026-09-09 and `v0.2.5.post1` on 2026-09-03, and the package metadata describes the project as Beta.
The `post1` suffix has a story worth knowing. That release was packaging only: it used the exact v0.2.5 source and rebuilt the full wheel matrix after the original PyPI upload was interrupted by a project storage limit. Anyone needing those wheels should pin `xgrammar==0.2.5.post1`. That is the kind of incident a small infrastructure project hits and handles in the open.
The v0.2.6 entry is the most substantive. Beyond the compilation optimisations, it hardened input validation and worker-thread error handling so invalid inputs raise exceptions instead of crashing the process, preserved tokenizer JSON round trips including binary vocabulary and token-grammar compilation, extended EBNF and Lark grammars with capture, lazy matching, suffix conditions and token budgets, added an NPU token-bitmask backend and Windows ARM64 wheels, and fixed several JSON and grammar edge cases.
Adding a MiniMax M3 tool-call tag and a `parallel_tool_calls` parameter in v0.2.7 tells you where the maintenance effort goes: keeping pace with new model-specific output formats. That is a maintenance obligation for a user, not just a maintainer, because your schema needs to know the tag dialect your model emits.
The README also documents a release policy: releases happen when features or fixes are ready rather than on a calendar, prereleases are marked separately, a maintainer selects a reviewed commit from `main`, and the package workflow builds and publishes when the GitHub release appears. Versions are generated from Git tags by `setuptools-scm`, so a build from a shallow clone without tags will not produce the right version.
What is settled, what is not, and where the docs take over
Three things are genuinely settled by the repository. The license is Apache-2.0 with a NOTICE file. The integration story is documented, dated and specific, with each engine named and each month given. And the mechanism is described precisely enough to reason about: a grammar compiles to something that produces a per-step token mask, which is why the overhead claim is plausible in the first place.
The README is otherwise a landing page. There is a technical report at arxiv.org/abs/2411.15100, a documentation site, and a blog post on achieving efficient, flexible and portable structured generation. The project has also published XGrammar-2, described in a May 2026 blog post as fast, customizable structured generation, which suggests a second-generation design exists that this repository's release notes only touch on indirectly.
So the documentation boundary is clear. The README tells you what the library is for and who uses it. The release notes tell you what is actually fixed and what new model formats are supported, which is the information you need before an upgrade. The technical report and the quick start tell you how it works.
If you are evaluating this against an alternative, the honest comparison is against regex-constrained decoding done by the inference server itself, and against libraries that parse and retry. XGrammar's argument is that structural correctness by construction plus grammar-level generality beats post-hoc repair, at a cost of a compiled grammar engine and a dependency on Torch. Whether near-zero overhead holds for your grammar complexity is the measurement to run, since the claim is stated for JSON generation and an expensive custom grammar is a different case.
Editorial conclusion
XGrammar solves a problem that post-hoc validation cannot: a JSON parser can tell you the output was wrong, and only a decoding-time constraint makes it right. That is why it became the default structured generation backend for most major inference engines, and the near-zero overhead claim in the README is the number to check against your own workload before committing to it. Two practical notes. The project is on version `0.2.7`, a beta per its own package metadata, with a fixed release policy and no calendar schedule, so pin an exact version. And model-specific structure is added per release, which means support for a new model's tagged output format means waiting for or building a grammar for it. Start with `pip install xgrammar`, run the quick start against one schema, then read the technical report before wiring it into a serving path where a generation stall would be your problem.
Frequently asked questions
How does XGrammar work?
It uses constrained decoding. At every sampling step XGrammar computes the set of tokens that are legal under your grammar at the current position and passes that to the model as a mask, so a token that would break the structure is never sampled and the output is structurally correct by construction. Grammars compile into a finite state machine, and the v0.2.6 notes name the optimisations as shared FSM construction, in-place optimizer passes and cached parser-state properties.
How do I install XGrammar and use it on Apple Silicon?
Install the base package with `pip install xgrammar`, then use `pip install "xgrammar[metal]"` for MPS support on Apple Silicon. The import is `import xgrammar as xgr`. The project supports Python `>=3.8, <4` and depends on `torch>=1.10.0`, `transformers>=4.38.0` and `apache-tvm-ffi`.
Which inference engines use XGrammar for structured output?
It is described as the default structured generation backend for vLLM, SGLang, TensorRT-LLM and MLC-LLM, with integrations into OpenVINO GenAI and Modular's MAX, and adoption inside Mirai. Practical upshot for many people is that using XGrammar means using vLLM, since it arrives with the server rather than as a separate choice.
Does XGrammar guarantee my structured output is correct?
It guarantees the output satisfies your grammar, so JSON, a regex or a custom context-free grammar will parse. It does not check that values are true or appropriate, so a field your schema types as a free string stays a string regardless of whether the content makes sense.
What license is XGrammar released under and what version is current?
Apache-2.0, with a LICENSE file and a NOTICE file at the repository root, which GitHub also reports. The current release is v0.2.7 from 2026-09-15, and the package metadata classifies the project as Beta, so pinning an exact version rather than tracking the latest release is the sensible habit.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/mlc-ai-xgrammar)