Library / SDK
pyparsing/pyparsing avatar
pyparsing/pyparsing

pyparsing ships AI instructions as an importable module and points a badge at another project

Python library for creating PEG parsers

2,494 stars349 forksPythonMIT

At a glance

What is it?
A parsing expression grammar library for Python whose README is twenty years old in its framing and current in its details. What the file and the packaging configuration show is a project that installs documentation for assistants into the package itself, benchmarks itself against every tag in its history, and carries one badge whose link goes somewhere else entirely.
Who is it for?
pyparsing suits code that has a grammar someone can read in Python source and wants to keep that readable, rather than a pattern that a regular expression would also cover. Three things are worth knowing before you commit.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 16 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 2, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The Python versions badge links to a different project on the package index

The badge row at the top of the file is six substitutions: version, build status, coverage, licence, Python versions and a security score. Five of them point where you would expect. The sixth does not. The Python versions badge is generated from the correct project name, but its link target is a different package entirely:

code
https://pypi.org/project/python-liquid/

The image URL names pyparsing, the link underneath it names python-liquid, which is a template language package rather than a parsing library. So clicking the badge that tells you which Python versions are supported lands you on someone else's release history. Everything else in the row is a link to the right place: the build badge goes to a workflow called ci.yml, the coverage badge to a codecov page on the master branch, the version and licence badges to the package index page, and the security score to a vulnerability advisor entry. It is a single copy and paste slip in a row of five correct ones, and it is the kind of error that survives because the badge image renders correctly either way. The row is defined at the bottom of the file as a set of substitution definitions rather than as inline HTML, so all six badges share one place and one editing convention, which makes the odd one out easier to miss. The same definitions carry the target for the coverage badge as a branch, master, which is also the default branch of the repository and therefore the only branch the coverage number describes.

The AI instructions ship as an importable module, not just a file

There is a section headed AI Instructions, and it describes two ways to get at the same document after installing the package. One is from the command line:

bash
python -m pyparsing.ai.show_best_practices

The other is from inside Python, by importing the package and calling a function:

python
import pyparsing; pyparsing.show_best_practices()

Either way the payload is a Markdown file called `best_practices.md` that lives inside the installed distribution rather than in a documentation site. That is a design decision with a cost. Every install now carries text meant to be read by a machine, and the copy inside the wheel is the copy an assistant will find first, which means it has to be kept current with the library rather than with the website. The file also names the source location for humans, saying it can be pulled from the project repository at `pyparsing/ai/best_practices.md`. The packaging configuration names the same file differently, listing `ai/best_practices.md` in what the source distribution includes. Those are two different paths for one document, and nothing visible here says which is the right one.

Benchmarking runs across every tag with two hand-written launchers

The performance section is two lines long. It points at `tests/README.md` in the repository for usage instructions and details on the benchmark suite. The root tells you the rest. Two files sit there named `run_perf_all_tags.sh` and `run_perf_all_tags.bat`, a shell script and a Windows batch file with the same purpose and no third variant. Running the benchmark suite across every tag in the project's history means the suite has to remain executable against old revisions of the library as well as the current one, which constrains what the benchmark code may use. It also means the comparison spans the entire life of the project rather than a chosen window of it. The two launchers are unusual next to the rest of the tooling: there is a tox.ini, a conftest.py at the root and a .coveragerc, so the ordinary Python work is declared rather than scripted, while the benchmarking is invoked by two platform files committed to the repository. Nothing in the README says how to interpret the numbers the suite produces. That gap matters more here than it would elsewhere, because the file is otherwise unusually confident about its own history. An italic note records that the description of the library was first written in late 2003, and that the technique has since become widespread under the name Parsing Expression Grammars. A library that can say precisely when its README was written is the same one that should say what a benchmark result means.

One result type is read three ways, and the literals come back as tokens

The worked example builds a grammar for a greeting out of three calls and two operators:

python
from pyparsing import Word, alphas
greet = Word(alphas) + "," + Word(alphas) + "!"
hello = "Hello, World!"
print(hello, "->", greet.parse_string(hello))

and the printed result is:

code
Hello, World! -> ['Hello', ',', 'World', '!']

Two things are visible there. First, the punctuation is not consumed silently. Each literal in the expression, the comma and the exclamation mark, appears as its own element between the two words, so the returned list is longer than the number of meaningful fields and a caller indexing it has to account for the literals. Second, the README states that the result is of type `ParseResults` and that this one object can be read as a nested list, as a dictionary, or as an object with named attributes. The readability claim rests on the operators, with the file naming plus, or and caret as the three definitions that carry the grammar. None of the three access styles is illustrated here, and the naming that would make attributes readable comes from the expression, not from a label. The file does say what the library handles that a hand-written parser usually does not, naming extra or missing whitespace, quoted strings and embedded comments, and pointing out that the same expression accepts `Hello,World!` and `Hello , World !` without changes. That tolerance is the argument for a grammar-based approach over a regular expression, and it is stated in a sentence rather than demonstrated.

The examples directory is full of files an optional extra produces

The packaging metadata declares an optional dependency group called diagrams, containing a diagram library and a template engine. The tree shows what that group is for. A large share of the example files exist in pairs, one HTML rendering and one PNG rendering of the same thing: a test suite parser has both, an adventure game parser has both, a grammar for another parser library has both, and so on across the directory. Those files are generated, and the generating step needs the extra. Meanwhile the source distribution configuration includes `examples` and `docs` in full. So a source distribution ships a set of generated diagrams produced by tooling the installer may never install, alongside the parser library it is meant to document. That is not wrong for a source tarball, where carrying the rendered output makes the examples readable without running anything, but it does mean the size of the distribution and the size of the library are two different numbers. The same include list also pulls in the AI instructions file and the test suite. One thing the file does not mention is any optional group for the examples themselves, so a contributor who wants to regenerate the diagrams has to know the group exists from reading the metadata rather than from reading the documentation.

The version is read from the module, and a script at the root writes it

The build configuration declares version and description as dynamic fields, which means they are not literals in the metadata file but are taken from somewhere in the package when the distribution is assembled. The backend is flit, pinned to a version range, and the author is a single name with a gmail address. At the root sits `update_pyparsing_timestamp.py`, a maintenance script whose name says what it touches. Read together, the version string is stamped into the package by a script rather than edited by hand, which is the arrangement that lets a release carry both a version and a date. The releases agree with that. 3.3.1 was published 2025-12-23, 3.3.2 on 2026-01-21 and 3.3.3 on 2026-09-20, and the last recorded push carries the same timestamp as 3.3.3, so the tagged release and the tip of master are the same commit. One classifier deserves a second look for a library of this age: the metadata declares it typed, while nothing in the visible documentation mentions type hints at all. The same block enumerates interpreter versions from 3.9 through 3.15, and states the floor twice, once as requires-python and once as the first classifier, which is the arrangement that lets a packaging tool check them against each other.

A dest/ directory and a config file without a leading dot

The root listing mixes conventions more than a library of this age usually does. Most configuration files carry a leading dot, namely .coveragerc, .pre-commit-config.yaml, .gitignore and the .github directory. One does not: the Read the Docs configuration is `readthedocs.yaml` with no dot in front of it, sitting in a row of dotted files. Both spellings exist in the wild, so this may well work, but it is the one configuration in the root that breaks the pattern around it. The documentation files are inconsistent in a second way. The readme is `README.rst` and the code of conduct is `CODE_OF_CONDUCT.rst`, both reStructuredText, while the change log is a bare `CHANGES` with no extension at all and the licence is `LICENSE`. Both of those are named in links from the file itself. And then there is `dest/`, a directory whose name reads like a build output folder, sitting in the source root with nothing in the documentation explaining what it holds or who generates it.

One of the shipped examples is a grammar for a different parser

The file names five example parsers: a simple SQL parser, a simple CORBA IDL parser, a config file parser, a chemical formula parser and a four-function algebraic notation parser. The tree holds considerably more, including an adventure game engine, a boolean search parser, a BigQuery view parser, a bytecode disassembler and a grammar for another parser library. The last of those is the interesting one, because a file named for a competing generator sitting in this project's examples says something about how the author positions it, or at least about what a reader might expect to find there. The directory also contains files that are not parsers at all: two Delphi form definitions, a Windows setup file, and an init file that makes the directory importable. Alongside the Python examples there is a README of its own. So the examples directory is a grab bag of runnable samples, generated diagrams, and files borrowed from other projects, and the file describes only the four or five of them that fit on one line of prose. The documentation section points in three directions at once, and the first of them explains why the tree looks so much bigger than the prose. It says there are many examples in the online docstrings of the classes and methods, that these are compiled into the hosted documentation, that further resources are on the project wiki, and that an entire directory of examples sits in the repository. The docstrings are the primary documentation, which is why the README itself stays short.

Editorial conclusion

pyparsing suits code that has a grammar someone can read in Python source and wants to keep that readable, rather than a pattern that a regular expression would also cover. Three things are worth knowing before you commit. The examples in the tree are full of diagram files produced by an optional diagrams extra that a plain install does not pull in, so a source distribution is larger than the parser. The AI instructions are installed as importable code, which is a deliberate choice rather than a documentation convenience. And the Python versions badge in the README links to an unrelated project on the package index, so read the classifiers in the packaging metadata for the supported versions instead of trusting the badge.

Frequently asked questions

what is pyparsing

It is a Python module for creating and executing simple grammars, an alternative to the traditional lex and yacc approach and to regular expressions. It provides a library of classes that client code uses to construct the grammar directly in Python code.

how to install pyparsing

The file does not give an install command. It points at the package index through its version and licence badges, and the packaging metadata declares requires-python >=3.9 with a flit build backend.

what does pyparsing do

It handles parsing problems that are typically vexing in text parsers: extra or missing whitespace, quoted strings, and embedded comments. Results come back as a ParseResults object readable as a nested list, a dictionary, or an object with named attributes.

how to use pyparsing

Build the grammar from imported classes and operator definitions, then call parse_string on the input. The worked example uses Word and alphas joined with plus signs, and calls greet.parse_string(hello).

pyparsing vs regex

The file positions pyparsing against regular expressions and against lex and yacc, saying it is an alternative approach to creating and executing simple grammars. It notes that since the description was first written in late 2003 the technique became widespread under the name Parsing Expression Grammars.

Official sources

  1. Issues
  2. License: MIT
  3. pyparsing/pyparsing on GitHub
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/pyparsing-pyparsing.svg)](https://hysenlabs.com/projects/pyparsing-pyparsing)