Library / SDK
lark-parser/lark avatar
lark-parser/lark

Lark: a Python parsing toolkit with Earley, LALR(1) and CYK under one grammar format

Lark is a parsing toolkit for Python, built with a focus on ergonomics, performance and modularity.

5,999 stars533 forksPythonMIT

At a glance

What is it?
Lark parses any context-free grammar from an EBNF definition and builds the parse tree for you. It is the right pick when your grammar is ambiguous or still moving, and the wrong pick when you need a fixed, dependency-free artifact in another language.
Who is it for?
Adopt Lark if you are writing a DSL, a config format or a query language in Python and want the tree handed to you from an EBNF grammar, especially if the grammar is ambiguous and still changing. Do not adopt it if you need a parser artifact that runs outside Python or without the library at runtime: only LALR(1) grammars can be turned into a stand-alone parser, and that path is documented in docs/tools.md rather than in the README.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 37 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 3, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The problem Lark solves: a grammar file instead of hand-written parsing code

Writing a parser by hand means writing a lexer, a state machine and tree construction code, then keeping all three in sync every time the language changes. Lark inverts that: you write the grammar in EBNF, and the library produces an annotated parse tree from the grammar and the input. The README states that no construction code is required, and the Hello World example shows the scale of it, a few lines of grammar producing a Tree of Tokens. The audience is split in two. The README addresses beginners directly, saying Lark is friendly for experimentation and can parse any grammar you throw at it, ambiguous or not. It also addresses experts, offering a choice between Earley and LALR(1) plus several lexers, so power and speed can be traded off. That split matters more than it sounds. Most parsing libraries assume you already know which algorithm you want. Lark assumes you may not, and lets you start with the permissive one and move down later. The pyproject.toml classifies the package as Production/Stable, and the projects listed as users include Poetry, Vyper, PyQuil, Preql and Hypothesis, which is a reasonable signal that the tree-building approach survives contact with real languages. Poetry uses it for dependency and packaging metadata, Vyper for a smart contract language, Preql for a query language that compiles to SQL. Those are three different kinds of grammar, and none of them is a toy.

How Lark works: grammar in, parse tree out, with the algorithm as a setting

The mechanism is a grammar file or string compiled into a parser object, which is then called with input text. The Hello World example shows the whole flow, a Lark instance built from an inline grammar with a start rule, a %import of common.WORD from the standard terminal library, and a %ignore directive for spaces. Calling parse on the string returns Tree(start, [Token(WORD, 'Hello'), Token(WORD, 'World')]). Two details in that output are worth pausing on. The punctuation is absent from the tree, which the README says happens automatically because punctuation is filtered away. And the tree is annotated, meaning the rule names from the grammar become the tree node names, so the shape of your grammar is the shape of your data. Underneath, Lark implements Earley with SPPF, LALR(1), and CYK according to the pyproject.toml description, and the topic list on the repository includes cyk, earley and lalr. The repository layout reflects the same split: lark/ holds the package, examples/ holds runnable grammars including json_parser.py, calc.py, lark_grammar.py and a standalone directory, and tests/ holds the suite. Grammar composition is a first-class feature, so terminals and rules can be imported from other grammars, and grammars can be imported from Nearley.js. If you have ever copied a terminal definition between two grammars and watched them drift, that feature is the answer to your problem. The interactive parser is the other piece worth knowing about before you start: the README lists it for advanced parsing flows and debugging, which suggests it is the tool to reach for when a grammar fails on one input and you need to see where.

Installing Lark and parsing something real

The README gives one install command, with no dependencies to resolve.

bash
pip install lark --upgrade

After that, the smallest useful program is the README's own example. Put it in a file and run it.

python
from lark import Lark

l = Lark('''start: WORD "," WORD "!"

            %import common.WORD   // imports from terminal library
            %ignore " "           // Disregard spaces in text
         ''')

print( l.parse("Hello, World!") )

The output is a Tree object, printed as Tree(start, [Token(WORD, 'Hello'), Token(WORD, 'World')]). If you see that, the grammar compiled and the tree was built. Note that the grammar is passed as a plain string here, which is fine for experimentation; for anything longer you would keep it in a .lark file and point Lark at it. The repository ships a syntax highlighting definition for .lark files for Sublime Text, TextMate, vscode, IntelliJ and PyCharm, Vim and Atom, which tells you the project expects grammars to live in files rather than in string literals. For a fuller first exercise, examples/json_parser.py in the repository is a JSON parser, and the README links a JSON tutorial that walks through writing one. That tutorial is also where the project explains how its performance comparison was made, so it doubles as the honest accounting of the benchmark images in the README. If you would rather not set up a project at all, the README links an online IDE, which is the quickest way to test whether a grammar idea works before you commit to it.

The Earley and LALR(1) split is the real decision, and it is not free

Lark's headline claim is that it parses all context-free grammars, and that claim belongs to Earley. The LALR(1) parser is described as fast and light and competitive with PLY, but LALR(1) accepts a strictly smaller class of grammars. So the two features that make Lark attractive pull in opposite directions. If your grammar is ambiguous, or uses constructs that are not LALR(1), you stay on Earley and give up the speed advantage and the stand-alone parser, which the README says is available for LALR(1) grammars only. If you need the stand-alone parser, you must first get your grammar into LALR(1) shape, and that is a grammar rewrite, not a flag. The README does not document a rollback path for that rewrite, and it does not describe how to diagnose an LALR(1) conflict beyond what the documentation covers. That is the honest limitation of the toolkit: the permissive mode and the fast mode are not the same product, and moving between them is work you do on the grammar. The CYK parser appears in the pyproject.toml feature list as the option for highly ambiguous grammars, which suggests a third point on the same axis rather than a way off it. The release history backs this reading. Release 1.2.2 was a bugfix for 1.2.1 covering Earley issues with ambiguity, and 1.3.0 included an Earley fix. The parser that carries the broad grammar claim is also the one that keeps needing corrections.

Where Lark is the wrong tool

Lark is pure Python, and the README presents that as a portability feature, since it runs on every Python interpreter. The cost is that your parser is a Python program. If the thing consuming the parse tree is not Python, or if you want to ship a parser as a binary artifact with no Python runtime, Lark is the wrong shape unless your grammar is LALR(1) and you use the stand-alone parser generator. Even then, the README's clones section is instructive: Lark.js is described as a port of the stand-alone LALR(1) parser generator to JavaScript, and Lerche is an unofficial clone in Julia. Those exist because the main project stops at Python. The second wrong-tool case is simpler. If your input format is already covered by a standard library module or a well-tested format-specific parser, writing a grammar is unnecessary work. The third is performance-critical parsing of very large inputs in a non-Python service, where a generated parser in the host language will beat a tree-building library that runs in the interpreter. The README's own performance framing is careful here, saying first-rate performance considering that this is Python. Read that qualifier as the boundary it is. There is also a packaging angle: the pyproject.toml registers a PyInstaller hook directory entry point, which means the project expects some users to freeze their applications, and that path is worth testing early if it applies to you.

Lark compared with PLY, PyParsing and ANTLR

The README's feature table is the clearest statement of the differences. PLY uses LALR(1) with BNF grammars, does not build a tree, does not support ambiguity, cannot handle every CFG and does not track line and column. So the difference from Lark is not speed, it is that PLY hands you a grammar and expects you to write the tree construction and the position tracking yourself. PyParsing, Parsley and Parsimonious are PEG-based, and the table marks them as unable to handle every CFG, with a footnote explaining that PEGs cannot handle non-deterministic grammars. That is the substantive difference from Lark: PEG combinators are written as code, and a PEG cannot express an ambiguous grammar at all, so ambiguity has to be resolved by ordering alternatives. ANTLR is LL(*) with EBNF, builds a tree, tracks line and column, and the table marks it as possibly able to handle every CFG, but it does not generate a stand-alone parser and it is not a Python library in the same sense. If you are choosing between Lark and PLY, the question is whether you want to write tree construction code. If you are choosing between Lark and a PEG library, the question is whether your grammar is ambiguous or non-deterministic. Those are different questions, and Lark's answer is yes to both. The README also points at the Python Parsing Benchmarks repository for third-party numbers, which is the right place to look rather than the project's own comparison images.

Maintenance, upgrades and the MIT licence

The repository is not archived, and the last push was on 2026-08-27. Release 1.3.1, dated 2025-10-27, is described as a bugfix release, and its notes add that the source build now contains complete project data, which matters if you build from a source distribution rather than a wheel. Release 1.3.0, dated 2025-09-22, introduced text-slices, an Earley fix and various small improvements. Release 1.2.2, dated 2024-08-13, was a bugfix for 1.2.1 covering Earley issues with ambiguity. The pattern across those three releases is worth reading: the Earley parser, the one that makes the broad grammar claim possible, is where the fixes land. If you rely on ambiguous grammars, budget for the possibility that a point release changes Earley behaviour, and pin your version. The package requires Python 3.8 or newer, and the pyproject.toml notes that since version 1.2 only Python 3.8 and up are supported, so an old interpreter is a hard blocker rather than a warning. On licensing, Lark is MIT, and the pyproject.toml declares it both as a license text field and as an OSI Approved :: MIT License classifier. MIT is permissive, but the repository contains a .gitmodules file, which means there are submodules whose licences are separate from the main project. If you vendor the source, check those, and treat this as a question for your own counsel rather than something a review can settle.

What to try before you commit

The fastest way to find out whether Lark fits is to write the grammar badly first. Take a few real inputs, write the smallest EBNF that accepts them, and run it with the default parser. If Lark builds a tree you can live with, the ergonomics claim is confirmed for your case. Then try the same grammar with the LALR(1) parser. If it compiles, you have the fast path and the stand-alone option; if it does not, you have learned early that your grammar is not LALR(1) and that the stand-alone parser is off the table without a rewrite. The examples directory gives you working grammars to compare against, including a JSON parser, a calculator and a grammar for Lark's own grammar format. The online IDE linked from the README is the place to do this without setting up a project, and the cheatsheet is a single PDF if you would rather have the EBNF syntax on one page than in a browser tab. One last check that costs nothing: run pip install lark --upgrade in the interpreter version you actually deploy on, since the 3.8 floor is enforced by the package metadata rather than by a warning at import time.

Editorial conclusion

Adopt Lark if you are writing a DSL, a config format or a query language in Python and want the tree handed to you from an EBNF grammar, especially if the grammar is ambiguous and still changing. Do not adopt it if you need a parser artifact that runs outside Python or without the library at runtime: only LALR(1) grammars can be turned into a stand-alone parser, and that path is documented in docs/tools.md rather than in the README. Before committing, verify one thing yourself: take your real grammar, run it through both the Earley and the LALR(1) parser, and see whether the LALR(1) build accepts it. That single check decides whether you get the speed and the stand-alone option or stay on Earley.

Frequently asked questions

What is Lark used for?

Lark is a parsing toolkit for Python that turns an EBNF grammar and an input string into an annotated parse tree. The README lists Poetry, Vyper, PyQuil, Preql and Hypothesis among the projects using it, which places it mostly in DSLs, config formats and language tooling.

What does a parser do in Python?

In Lark's case the parser takes a grammar you wrote and an input string, applies the grammar to the input, and returns a Tree of Tokens rather than a raw string. The README's Hello World shows this directly: parsing "Hello, World!" returns Tree(start, [Token(WORD, 'Hello'), Token(WORD, 'World')]), with punctuation filtered out automatically.

How do I install Lark?

The README gives a single command, pip install lark --upgrade, and states that Lark has no dependencies. The package requires Python 3.8 or newer, and the pyproject.toml notes that since version 1.2 only Python 3.8 and up are supported.

What is the difference between Lark's Earley and LALR(1) parsers?

Earley can parse all context-free grammars and supports ambiguous grammars fully, while LALR(1) is described as fast and light and competitive with PLY but accepts a smaller class of grammars. Only LALR(1) grammars can be turned into a stand-alone parser, so the two modes are not interchangeable.

Does Lark work outside Python?

Lark itself is pure Python and runs on every Python interpreter, so the library is Python-only. The README lists clones for other languages: Lerche, an unofficial clone written entirely in Julia, and Lark.js, a port of the stand-alone LALR(1) parser generator to JavaScript.

What licence does Lark use?

Lark is MIT licensed, declared in pyproject.toml as both a license text field and an OSI Approved :: MIT License classifier. The repository also contains a .gitmodules file, so any submodules carry their own licences separately from the main project.

Official sources

  1. Issues
  2. lark-parser/lark on GitHub
  3. License: MIT
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/lark-parser-lark.svg)](https://hysenlabs.com/projects/lark-parser-lark)