Open-source project
lalrpop/lalrpop avatar
lalrpop/lalrpop

LALRPOP: an LR(1) parser generator for Rust that ships its own macros

LR(1) parser generator for Rust

3,505 stars313 forksRustApache-2.0

At a glance

What is it?
A parser generator whose real selling point is ergonomics, with grammar macros, type inference and compact defaults, bootstrapped on its own grammar and used by RustPython and Solang.
Who is it for?
LALRPOP's argument is that the hard part of a parser generator is not the table construction, it is what you have to write to get a grammar you can maintain. That claim is carried by the macro system rather than the algorithm, and the project backs it up by bootstrapping on its own grammar and by shipping in RustPython, Solang and Gluon.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 15 days ago.
What is it written in?
Mainly Rust, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 7, 2026, and from our analysis. They are not legal advice.

Editorial analysis

Usability named as the primary goal

The README states the goal before describing any feature: LALRPOP is a Rust parser generator framework with usability as its primary goal, meaning you should be able to write compact, DRY, readable grammars. That framing matters because parser generators are exactly the class of tool where ergonomics get sacrificed for table construction convenience.

The feature list that follows is a list of things that make a grammar shorter. Macros let you extract common parts of a grammar, which the README explains goes beyond simple repetition. You can write `Id*` for a plain sequence, or define `Comma<Id>` for a comma-separated list of identifiers, and use it anywhere.

Macros can also produce subsets. The example given is defining `Expr<"all">` for the full range of expressions and `Expr<"if">` for the subset that can appear inside an if expression. That is a genuinely useful idea, because in real languages the context-valid subset of a construct is not a separate rule, it is the same rule with a restriction, and expressing that restriction as a macro parameter keeps one definition instead of two.

The remaining three items are about reducing what you must write. Builtin support for operators like `*` and `?` covers the common repetition cases. Compact defaults mean you avoid writing action code much of the time. Type inference means you can often omit the types of nonterminals. Together they are an argument that a grammar should read like the language it describes.

Why the name says LALR and the default is LR(1)

The most important sentence in the README is an aside. Despite its name, LALRPOP in fact uses LR(1) by default, though you can opt for LALR(1).

This is not a cosmetic naming issue. The distinction is about which grammars a generator can accept. Canonical LR(1) and LALR(1) construct different tables. LALR merges states that LALR(1) parsers derive from the same LR(0) core, which produces smaller tables and is cheaper to run, but it also introduces reduce-reduce conflicts into grammars that LR(1) would accept cleanly. A language grammar with even a small amount of dangling structure can trip LALR and pass LR(1).

So if you arrive expecting a LALR implementation because of the name, you will find the tool is more permissive than you planned for, which is the better default for a language project. If you arrive expecting full canonical LR(1) with no merge at all, LR(1) is still only one family of the deterministic bottom-up parsers, and the README says so.

The second half of that sentence is a statement about where the author wants the project to go: there is a hope of eventually moving to something general that can handle all context-free grammars, with GLL, GLR and LL(\*), named as candidates. Read that as an acknowledgment that a deterministic LR family will not cover every grammar you might want to parse. For most configuration and query languages, LR(1) is plenty. For a general programming language with unusual features, it may not be.

Bootstrapped on its own grammar

The first item under Example Uses is that LALRPOP is itself implemented in LALRPOP, with a link to the file that does it.

The path in the repository is `lalrpop/src/parser/lrgrammar.lalrpop`, a file that describes the grammar of the LALRPOP grammar language itself, written in LALRPOP's syntax. The file name is a small pun: the grammar of a tool that generates parsers is written in that tool, and the parser generator's own parser is generated by the parser generator.

Self-hosting is a real test rather than a party trick, and it is worth understanding why it is here. It means any change to the grammar syntax has to be expressible in the previous version of the syntax, which catches the class of change that would otherwise break a bootstrapping step. It also means the grammar file is the definitive specification of what the syntax accepts, more authoritative than prose documentation.

There is a repository script for keeping it straight: `update_lrgrammar.sh` sits in the root next to `version.sh`, which are the sort of files a self-hosted project needs and everyone else does not.

The tree also shows `CONTRIBUTING.md` and a specific note in the README about it. If you intend to change LALRPOP's own grammar, the README says you really should read that file first, which is unusually strong phrasing for a contribution guide and reflects the bootstrapping complexity.

Who else runs it in production

The example uses list names three projects outside the repository, and all three are real users rather than showcase entries.

Gluon is a statically typed functional programming language, with its grammar linked at `parser/src/grammar.lalrpop`. RustPython is Python 3.5 and later rewritten in Rust, with its grammar at `parser/src/python.lalrpop`, which is a considerably more demanding test than a small configuration language since Python's grammar has unusual indentation rules and a large number of special cases. Solang is Ethereum's Solidity compiler rewritten in Rust, with its grammar at `solidity-parser/src/solidity.lalrpop`.

Two things follow from that list. First, the range from a functional language to Solidity to Python suggests the grammar features are general enough for real languages, and RustPython is the strongest evidence, since Python's tokenizer and indentation handling sit next to the grammar rather than inside it. Second, the LALR question above has an empirical answer that favors the LR(1) default: if a project as grammatically awkward as Python's can be handled without falling back to a GLR parser, the permissive default is doing its job.

There is also a `lalrpop-util/` workspace member and a `lalrpop-test/` member, and the test workspace is not trivial. Looking at the Cargo workspace members list, there are separate example crates for a calculator, a Nabol language, a Pascal grammar, whitespace handling, a lexer, lexer modes, and a C-like grammar, alongside the two main crates. That is a test matrix rather than a single smoke test, and it is a better signal of parser generator health than any single example.

Version, edition, and the documentation layout

The workspace `Cargo.toml` gives the hard numbers. The version is 0.23.1, the Rust edition is 2024, and the declared minimum is rust-version 1.86, with a comment explaining the policy: it is a soft limit, they prefer to avoid the latest two stable versions, and the comment notes that `test.yaml` has to be updated alongside it. The license field is dual, `Apache-2.0 OR MIT`, matching the `LICENSE-APACHE` and `LICENSE-MIT` files at the root.

The documentation is split across two systems, which is worth knowing. There is a `doc/` directory for the book, with the workspace members `doc/calculator`, `doc/nobol`, `doc/pascal/lalrpop`, `doc/whitespace`, `doc/lexer`, `doc/lexer-modes` and `doc/cfg` all living inside it. Those are runnable examples, and the fact that lexer modes and a C-like grammar are separate members suggests the book teaches lexing as a distinct concern rather than assuming you already have one.

There is also a `.mdl_style.rb` and `.mdlrc` at the root for markdown linting, and a `RELEASES.md` alongside a `.clog.toml`, so the changelog is generated from conventional commits. Both suggest a project with release engineering in place rather than ad hoc tagging.

The README itself routes you through the hosted book at lalrpop.github.io, and points to four specific entry points: the tutorial for basics, the quick start guide for adding the build dependency to `Cargo.toml`, a cheat sheet for returning users, and an advanced setup chapter for configuring the preprocessing step. If you are new, the tutorial plus the cheat sheet is the efficient path. If you find yourself wanting to write a macro to run at parse table construction time, that is the chapter you want.

What the README does not promise

Two absences are worth naming, because both are things you might otherwise assume from the feature list.

The first is error recovery. The README's feature list mentions nice error messages when the parser constructor fails, and that is a real and different thing from recovering from a syntax error in the input. A constructor failure means your grammar has a conflict that LALRPOP detected while building. Error recovery in the input means the parser can resynchronize and keep going, which is a separate feature with separate machinery. The related searches around this project do include queries about error recovery, so it is on people's minds, but the README does not list it as a feature, which means you should confirm it in the book rather than assume it.

The second is precedence and associativity handling, which also shows up in the related searches. Nothing in the README's feature list mentions operator precedence declarations. What the list does mention is compact defaults, which is a related but not identical mechanism, since you can often express precedence by structuring the grammar rather than declaring it. If you are porting a yacc grammar, that difference will matter and it is a good reason to read the tutorial rather than translate mechanically.

Neither absence is a criticism of the project, both are places where the README is a summary rather than a manual. The book is where those questions get answered, and the fact that the project lists a cheat sheet and an advanced setup chapter is a sign that the manual is treated as the real product here.

Editorial conclusion

LALRPOP's argument is that the hard part of a parser generator is not the table construction, it is what you have to write to get a grammar you can maintain. That claim is carried by the macro system rather than the algorithm, and the project backs it up by bootstrapping on its own grammar and by shipping in RustPython, Solang and Gluon. Two things to know before you start. The name is misleading, since LALRPOP uses LR(1) by default and only offers LALR(1) as an option, which matters if you arrive expecting canonical LALR. And the project openly admits its algorithm family does not cover all context-free grammars, naming GLL, GLR and LL(*) as the general algorithms it would rather move toward. Read the tutorial chapter for the setup path, keep the cheat sheet open while writing a real grammar, and read the advanced setup chapter before reaching for preprocessing.

Frequently asked questions

What does LALR stand for?

That question is about the parser theory rather than this tool, and the README's answer to it is a surprise worth knowing: despite the name, LALRPOP uses LR(1) by default and only offers LALR(1) as an opt-in. The distinction matters because LALR merges parser states and can introduce reduce-reduce conflicts that LR(1) accepts, so the default is more permissive than the name suggests.

What is the difference between LR, SLR, CLR, and LALR parsers?

It is a family of deterministic bottom-up parsers that differ in how much lookahead they use to decide a reduction, which affects table size and which grammars they accept. LALR merges states with the same core, giving smaller tables and more conflicts. SLR merges further and conflicts most. LALRPOP sits at LR(1) by default, with LALR(1) available if you want smaller tables and your grammar can take it.

Which is the most powerful parser?

That is a theory question with no single answer, and the README sidesteps it by naming the general algorithms instead. Deterministic LR family parsers are fast but reject some grammars, while GLL, GLR and LL(\*), which the README lists as things the author would like to move toward, accept all context-free grammars at the cost of more machinery. LALRPOP is an LR(1) generator today.

How do I define a comma-separated list in a LALRPOP grammar?

With a macro. The README's example is `Comma<Id>` for a comma-separated list of identifiers, which the generator expands wherever you use it. That is the macro system doing its job, and it goes beyond the built-in `*` and `?` operators for simple repetition. The same mechanism can parameterize a nonterminal to expose a subset, such as `Expr<"if">`.

Which projects use LALRPOP?

Three are named with grammar file links. LALRPOP is implemented in LALRPOP, using its own grammar language. Gluon, a statically typed functional language, uses it. Solang, the Ethereum Solidity compiler in Rust, uses it. And RustPython, Python 3.5 and later rewritten in Rust, uses it, which is the most demanding of the set.

Official sources

  1. Issues
  2. lalrpop/lalrpop on GitHub
  3. License: Apache-2.0
  4. Project website
  5. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/lalrpop-lalrpop.svg)](https://hysenlabs.com/projects/lalrpop-lalrpop)