Library / SDK
eliben/pycparser avatar
eliben/pycparser

pycparser: no dependencies, but you still need a preprocessor

:snake: Complete C99 parser in pure Python

3,573 stars653 forksPythonNOASSERTION

At a glance

What is it?
pycparser parses C in pure Python with no external Python dependencies, and that is both its main argument and its main trap: real code has to be preprocessed before the parser sees it, so a system compiler is required anyway. The repository is also unusually explicit about how its own abstract syntax tree is generated.
Who is it for?
pycparser fits a tool that needs to read C declaratively, such as an obfuscator, a specialised compiler front end or an interface generator, and it fits it better than a grammar tool would because the source is Python you can read and patch. Two things to check before you commit.
Can I use it commercially?
Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
Is it still maintained?
Yes. The repository received new commits within the last day.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 5, 2026, and from our analysis. They are not legal advice.

Editorial analysis

No dependencies, and a compiler you still have to install

The claim at the top is that this parser has no external dependencies except a Python interpreter, and that this makes it very simple to install and deploy. As a statement about Python packages, that is exactly right, and the manifest has no dependency list at all.

Then the section on using it explains why you cannot actually parse a C file without something else on the machine. C code has to be preprocessed by the C preprocessor before it can be compiled, and that step handles the include directives, the definitions, removes comments and does the other small tasks that make the source compilable. For all but the most trivial snippets, the parser has to receive preprocessed code, exactly as a compiler would.

The convenience is that the top-level file-parsing function will do it for you. If you import that function rather than the lower level interface, it invokes the preprocessor itself as long as it is on your PATH, or you can give it a path.

So the dependency-free property is real for installation and irrelevant for operation. A virtual environment on a machine with no compiler installs the package happily and then fails on the first real source file.

bash
pip install pycparser

The syntax tree is generated code, and the config file ships with it

This is the detail that most surprises people reading the source, and it is documented plainly.

The code for the parser's syntax tree nodes is not written by hand. It is generated from a configuration file by a generator script, and if you modify that configuration you are expected to re-run the generator, from either the repository root or the package directory.

The consequence is that the public shape of the syntax tree has a single point of definition in a data file. Anything that wants to know exactly which node types exist and what they contain is pointed at that same file, rather than at the generated module.

That is why the package data configuration in the manifest includes configuration files as package contents. The file is not just a build-time input; it ships with the installed package, which is what makes it possible for the generated code and any consumer that inspects it to agree about the tree.

The other half of the configuration story is on the parser class itself. The documentation says to read the constructor's docstring for details on the configuration and compatibility arguments, which is where the knobs for parsing behaviour live rather than in the configuration file.

Fake standard headers are the trick that makes type parsing work

Almost all real C includes headers from the standard library, and those headers are where a parser like this one would normally drown.

The document says that the parser can be made to parse the real standard headers from any C compiler, with some effort, and that it is much simpler to use the fake standard includes the project provides for C11. Those are standard-looking header files containing only the bare necessities required for valid parsing.

The reasoning that makes this work is the most useful paragraph in the readme. The parser does not care about the semantics of types at all. It only needs to know whether a token it has already seen was defined as a type. That single requirement is what allows it to parse C declaratively, and it is why a minimal fake header is sufficient: the declarations only have to name the types.

There is a second benefit, mentioned in the same paragraph. Because the fake headers are minimal rather than complete, they can significantly improve the performance of parsing large C files, which is the opposite of the intuition that fuller headers would be safer.

The document links a blog post on parsing C type declarations and fake headers for anyone who wants the longer explanation.

The fake headers are not shipped, and the package list still claims they are

There is a contradiction inside the readme, and it affects anyone installing the package.

One note states that the fake headers are not included in the installed package and are not installed by the package build, and it cites an issue number for the tracking of that gap. The same document, a few sections later, has a package contents section listing the files and directories you will find once you unpack the distribution, and that list names the fake standard include directory and describes it as minimal headers that should allow any C code to be parsed.

The manifest settles which statement is current. It declares exactly one package for installation, the parser's own directory, plus configuration files as package data. The directory holding the fake headers is not a package and is not listed, so it is not installed.

The practical consequence is real rather than pedantic. A pipeline that installs the wheel and expects to parse C against the shipped headers will not find them, and the failure appears at parse time as a missing include rather than at install time. If you rely on them, copy them from the repository or point the preprocessor at the version in your source tree.

Two tools, pinned in a makefile and run on demand

The build file is four targets and it tells you how the project is checked.

There is a lint tool and a type checker, and neither is invoked from an installed environment. Both are run through an on-demand runner at a version pinned as a makefile variable: one variable for the lint tool's version and one for the type checker's, each overridable from the command line.

The check target runs three commands in order: the linter in check mode, the type checker, and then the formatter. So formatting is part of the check rather than a separate step, and a contributor running make check gets linting, type checking and rewriting in one command.

The test target runs the standard library's test discovery. No third-party test runner appears in it, which for a project of this size and age is a deliberate simplification rather than an omission.

The lint configuration is also worth reading, because it ignores exactly two rules, and both of them are about importing names from a module without being able to prove where they came from. In a parser that loads generated code and fake headers dynamically, those are the two rules you would expect to switch off, and nothing else is.

Three packaging files, a committed editor config, and a to-do list

The root of the repository has a few things that say more about the project's habits than its feature set does.

There are three packaging files. The legacy installation script is a two-line shim that imports the packaging call and invokes it, and the readme labels it a legacy script and points at the modern metadata file for build information. That modern file is present and complete. A third configuration file also exists, which is a leftover from an era when the metadata lived there.

There is a manifest file for controlling what goes into a source distribution, and a lockfile for the dependency and tooling environment.

Then the contributor conveniences. A pinned interpreter version file for version managers, and an editor configuration file committed at the repository root, which is a small act of respect for contributors who did not choose the same editor as the maintainer.

The documentation files are mixed: a readme in reStructuredText, a contributors file, a security policy, a plain text to-do list, and a file of older changes kept alongside whatever the current history file is. A to-do list in a plain text file at the root is a project telling you it still uses that system, which is more honest than an empty tracker.

Sixteen examples are an inventory of the tree API

The examples directory is the best documentation in this project, and reading the file names tells you what the library can do to a syntax tree.

Two are about seeing: one dumps the tree and one explores it. Then the reverse operations. One constructs a tree from scratch rather than parsing it, one rewrites a tree, and one serialises it. Then the query operations: listing function calls, listing function definitions, and adding a parameter to a definition.

Then the converters. One turns C into C, which is a restatement and a round trip test in one file. One produces JSON. One produces a declaration string in the format C compilers emit, which is how you generate a header from a parsed source. One is about typed usage, which is the parser's notion of a type.

And two are about the preprocessor, one driving it directly and one using the compiler's own preprocessing mode instead. Those two exist because the preprocessor step is the part of the pipeline that is not this library's job.

The examples directory also carries its own readme, and the main readme notes that most realistic samples need preprocessing first, which is the same requirement stated in the operating instructions rather than in a footnote.

Version 3.00, a release_ tag prefix, and the C dialect it targets

A few version and compatibility facts, gathered in one place.

The manifest records the version as 3.00, with two decimals, and the readme's own title uses the same form. The release tags carry a prefix before the version and keep the two-decimal scheme, so the newest is a third major, preceded by a second-series release from 2025 and another from 2024. The newest release is dated in January 2026 and the last recorded push to the default branch is dated 2026-10-02.

The Python floor is 3.10 with no upper bound, and the classifiers list Python 3.10 through 3.15. The library is tested on Linux, macOS and Windows, with a continuous integration dashboard linked from the readme.

On the C side, the stated goal is the full C99 language as defined by the standard, and the grammar followed is the one in the standard's annex, followed very closely. Some features from the next revision of the language are supported, and patches for more are welcome.

Compiler extensions are the stated limit. Very few are supported out of the box, though the readme says it is fairly easy to configure the parser to handle code with a great many of them, and points at its FAQ for the details.

Editorial conclusion

pycparser fits a tool that needs to read C declaratively, such as an obfuscator, a specialised compiler front end or an interface generator, and it fits it better than a grammar tool would because the source is Python you can read and patch. Two things to check before you commit. The no-dependency claim is about Python packages, not about the system: anything beyond trivial snippets needs a C preprocessor on the machine, so a container that ships only a Python interpreter cannot parse real files. And the fake standard headers are not installed with the package, so a pipeline that assumes the wheel is self-contained will fail on the first header. Everything else is unusually well documented for a parser of this size, including the grammar source and the extension story.

Frequently asked questions

how to install pycparser

The recommended installation is pip install pycparser. The package has no external Python dependencies at all and is tested with modern Python versions on Linux, macOS and Windows. Parsing real C code still requires a C preprocessor on the system.

how to use pycparser

Give it preprocessed C. The top-level file-parsing function will invoke the preprocessor for you when it is on your PATH or when you provide a path, and the compiler's own preprocessing mode can be used instead. The examples directory covers dumping the tree, constructing and rewriting it, serialising it and listing functions.

Does pycparser need external dependencies?

It has no external Python dependencies, and the manifest declares none. However, C code must be preprocessed before parsing, so for anything beyond trivial snippets a preprocessor has to be available on the machine, either the standalone one or the compiler's own preprocessing mode, and Windows users are pointed at a Clang binary.

Which versions of C does pycparser support?

It aims to support the full C99 language as defined by the standard, following the grammar in that standard's annex closely, and some features from the next revision are supported with patches welcome. Very few compiler extensions work out of the box, although the readme says the parser can be set up to handle code with many of them.

Do the fake C standard headers ship with the installed pycparser package?

No. The readme states that they are not included in the installed package and not installed by the package build, and the manifest installs only the parser package. The package contents section of the same document still lists the directory, so that part of the readme is out of date.

Official sources

  1. eliben/pycparser on GitHub
  2. Issues
  3. README
  4. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/eliben-pycparser.svg)](https://hysenlabs.com/projects/eliben-pycparser)