Guidance: constraining LLM output with regex, CFGs and Python control flow
A guidance language for controlling large language models.
At a glance
- What is it?
- Guidance is a Python library that interleaves generation with control flow and constrains model output with regex and context free grammars. It is for engineers who need structured output from an LLM rather than a prompt template that mostly works.
- Who is it for?
- Adopt Guidance when your output has a shape you can write down as a regex, a choice list or a grammar, and when you want that shape enforced during generation rather than checked afterwards. Do not adopt it if you only need free-form chat text, or if your backend is not among the supported ones, since the README states that full constraint support depends on the backend.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 132 days ago.
- What is it written in?
- Mainly Jupyter Notebook, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 20, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What Guidance solves that a prompt template does not
A prompt template asks the model for JSON and hopes. Guidance instead restricts what the model is allowed to emit while it emits it. The README frames the project as "an efficient programming paradigm for steering language models," and the claim attached to that framing is lower latency and cost compared with conventional prompting or fine-tuning. The audience is narrow and specific: Python developers who already have a model backend and who need output that matches a known shape, such as a digit string, one letter from a fixed set, or a string parseable against a context free grammar. If your application just needs conversational text, the constraint machinery buys you nothing and adds a grammar layer between you and the model.
How generation and control flow interleave
The mechanism is a model object that accumulates text. Calls like gen() and select() are appended to that object, and the object carries the state of the conversation so far. The README states that model objects are immutable, so `lm = phi_lm` produces a copy rather than a mutation of the original, which is why the examples reassign `lm` before each new exchange. Role context is expressed with context managers: `with system():`, `with user():` and `with assistant():` wrap the segments you append. Constrained generation is a property of the gen() call itself. `gen("lm_age", regex=r"\d+", temperature=0.8)` restricts that span to digits. `select(["A", "B", "C", "D"], name="model_selection")` restricts it to one of the listed strings. Named output is retrieved by key afterwards, as in `lm['lm_response']`. The underlying constraint packages are declared in pyproject.toml as `guidance-stitch==0.1.5` and `llguidance==1.6.1`, both pinned to exact versions, which tells you the constraint layer is treated as a dependency with a stable interface rather than something in flux.
Installing Guidance and running a first constrained call
The README gives a single install line and notes that Guidance supports several backends including Transformers, llama.cpp and OpenAI. If you already have the backend your model needs, the base install is enough.
pip install guidanceBackend support is split into optional extras in pyproject.toml. The extras are named `openai`, `azureai`, `llamacpp`, `transformers` and `onnxruntime-genai`, with an `all` extra that pulls in the first four. Installing the base package without the matching extra leaves you without a usable model class. The README's first example imports `Transformers` and loads `microsoft/Phi-4-mini-instruct`, then appends a system message, a user message and a generated assistant reply:
from guidance import system, user, assistant, gen
from guidance.models import Transformers
phi_lm = Transformers("microsoft/Phi-4-mini-instruct")
lm = phi_lm
with system():
lm += "You are a helpful assistant"
with user():
lm += "Hello. What is your name?"
with assistant():
lm += gen(max_tokens=20)
print(lm)Run at the command line, the README says this prints the full rendered conversation including the role markers, with the assistant turn filled by the model. In a Jupyter notebook the same code renders a widget instead, which the README shows as a screenshot. The second thing worth running early is the local grammar check, because it exercises the constraint system without touching a model API: `from guidance.models import Mock` is used with `grammar.match(...)` to assert that a string conforms and that an invalid string does not. That path is the fastest way to find out whether your regex actually describes what you think it does.
Writing your own Guidance functions with the decorator
The `@guidance` decorator turns a plain Python function into a reusable step that composes like gen() or select(). The README's multiple choice example defines `zero_shot_multiple_choice(language_model, question, choices)`, appends the question and lettered choices inside a `with user():` block, then appends a `select()` over the letter range inside `with assistant():`, and returns the model. The `language_model` parameter is filled in automatically when the function is called as `lm + zero_shot_multiple_choice(...)`. The return value is a new model object, which is why the example assigns it to `lm_temp` and reads the answer back with `lm_temp["string_choice"]`. This is the part of the design that earns its keep: the constraint lives in a function you can unit test, call in a loop, and reuse across question sets, instead of being buried in a prompt string. The trade-off is that the function signature is coupled to the model object, so the same helper is not directly reusable with a different library's client.
Where Guidance is the wrong tool
The README is explicit that the constraint system can enforce any context free grammar "so long as the backend LLM has full support for Guidance." That qualifier is the main limitation. Constraint enforcement is not uniform across the supported backends, and the README does not provide a per-backend compatibility table, so you cannot tell from the install section alone whether the backend you want gives you full grammar enforcement or a reduced version. The README also does not document rollback, retry semantics or what happens when a constrained generation cannot be satisfied, so failure handling around an impossible grammar is something you would have to determine yourself. Guidance is also the wrong choice when your output is prose. Wrapping a free-form answer in gen() with no constraint gives you the same result as any other client, plus an extra abstraction. And if your team is not writing Python, the whole interface is unavailable: the package declares `requires-python = ">=3.10"` and the examples are Python and notebook code.
How it differs from structured-output features in model client libraries
OpenAI's Python client, installed here through the `openai` extra, offers its own structured output mechanisms tied to that vendor's API. The difference in approach is where the constraint lives. Vendor structured output is a request parameter interpreted server-side, and it works with the models that vendor serves. Guidance applies constraints client-side through its own grammar engine, which is why a single `gen(regex=...)` call can in principle be pointed at Transformers, llama.cpp, OpenAI or Azure AI backends without rewriting the constraint. The cost of that portability is the backend-support caveat above: you get one constraint vocabulary, but its enforcement depends on the backend you attach it to. If you are committed to one vendor and its native structured output covers your shape, the vendor path has fewer moving parts. Guidance makes sense when you want the same grammar to survive a backend change.
Maintenance, licensing and upgrade cost
The repository is not archived, and the last push was on 2026-05-21. The most recent release listed is 0.3.2 from 2026-03-18, preceded by 0.3.1 in February 2026 and 0.3.0 in September 2025. That cadence suggests a project that ships a few times a year rather than continuously, so pinning a version and reading release notes before upgrading is reasonable. The licence is MIT, declared in pyproject.toml as `license = {file = "LICENSE.md"}`, which permits commercial and closed-source use; that is a statement about the licence text, not legal advice, and you should read LICENSE.md yourself if redistribution terms matter to you. The upgrade cost is concentrated in two pinned dependencies, `guidance-stitch==0.1.5` and `llguidance==1.6.1`. Because both are exact pins rather than ranges, a Guidance upgrade that moves either one is the change most likely to affect constraint behavior, and it is the first thing to check in a diff.
Editorial conclusion
Adopt Guidance when your output has a shape you can write down as a regex, a choice list or a grammar, and when you want that shape enforced during generation rather than checked afterwards. Do not adopt it if you only need free-form chat text, or if your backend is not among the supported ones, since the README states that full constraint support depends on the backend. Before committing, verify two things on your own machine: which of the optional extras matches your model backend, and whether your chosen backend honors the constraint system, because the README says CFG support holds only so long as the backend has full support for Guidance.
Frequently asked questions
How do I install Guidance?
The README gives `pip install guidance` for the base package. Backend support is split into optional extras named openai, azureai, llamacpp, transformers and onnxruntime-genai in pyproject.toml, so you install the extra matching your model backend.
Which model backends does Guidance support?
The README names Transformers, llama.cpp, OpenAI and Azure AI, and the optional dependencies in pyproject.toml list openai, azureai, llamacpp, transformers and onnxruntime-genai. The README states that full context free grammar support depends on the backend having full support for Guidance.
Can I test Guidance grammars without calling a model API?
Yes. The README shows validating strings directly with `grammar.match(...)` and running the same grammar against a local `Mock` model constructed from bytes, which avoids any API call.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/guidance-ai-guidance)