antlr/grammars-v4: The Official Collection of Action-Free ANTLR v4 Grammars
Grammars written for ANTLR v4; expectation that the grammars are free of actions.
At a glance
- What is it?
- antlr/grammars-v4 is the canonical repository of ANTLR v4 grammars covering dozens of programming languages and file formats. The grammars are free of embedded actions, making them portable across ANTLR's supported target runtimes including Java, Python, C#, JavaScript, and Go.
- Who is it for?
- Teams that need a ready-made ANTLR v4 grammar for a common language will find this collection the first place to check. Those needing a production-grade parser for a high-stakes use case should verify the specific grammar's completeness independently, since the README makes no quality guarantees for individual grammars.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 4 days ago.
- What is it written in?
- Mainly ANTLR, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 27, 2026, and from our analysis. They are not legal advice.
Editorial analysis
Action-Free Grammars: The Central Design Constraint
ANTLR grammars can embed arbitrary code (called actions) directly into grammar rules. An action might transform a parse result, call a Java method, or assign a variable, and it is written in the target language (Java, Python, C#, etc.) of the generated parser. When a grammar embeds Java actions, it can only generate a Java parser; porting it to Python or C# requires rewriting all the embedded code.
The grammars in antlr/grammars-v4 are expected to be free of such actions. The repository description states this constraint explicitly: 'expectation that the grammars are free of actions.' An action-free grammar describes only the language's syntactic structure. ANTLR generates parsers and lexers from it for any supported target runtime without modification. This is the primary reason the collection is useful across different projects and languages.
From an engineering standpoint, the trade-off is clear. An action-free grammar is more general but less optimized for any specific use case. A grammar with embedded actions can produce exactly the data structure a particular application needs. A grammar without actions produces a parse tree in ANTLR's standard representation, and the application must transform that tree into whatever it actually needs using a visitor or listener. For exploratory use and tooling integration this is often acceptable. For a production parser where performance is critical or the output type is fixed, you may need to add actions yourself or fork the grammar.
The action-free constraint also makes grammars easier to reason about and test. The grammar file describes one thing: what sequences of tokens are syntactically valid. This separation of parsing from transformation is a sound design principle for maintaining grammar files over time.
Repository Organization: One Lowercase Directory Per Language
The repository organizes grammars by language and format. The README states the naming convention directly: 'The root directory name is the all-lowercase name of the language or file format parsed by the grammar. For example, java, cpp, csharp, c, etc.'
The top-level directory listing confirms the breadth of coverage. The collection includes grammars for programming languages (ada, awk, basic, c, cpp, csharp, java, asm, apex, angelscript, alloy, algol60, alef, and dozens more), file formats (bdf, bencoding, bibcode, bibtex, capnproto, clf), domain-specific languages (arden, aql, caql, amazon-states-language, amazon-states-language-intrinsic-functions), and specification languages (bnf, antlr itself, abnf, asl, asn). The directory listing in the repository starts with abb and continues through the full alphabet.
To find the grammar for a language you need, navigate to the directory with the language's all-lowercase name. The Java grammar is in java/, the C grammar in c/, the C++ grammar in cpp/, the C# grammar in csharp/. The consistent naming means no guessing or searching is required for well-known languages.
Inside each language directory, the grammar file or files use the .g4 extension. Some languages have multiple grammar files: one for the lexer and one for the parser, or separate grammars for different dialects or versions. The README does not document the internal structure of individual language directories; that is available in each directory's own files.
The repository also includes configuration directories (.config/, .github/), a CLAUDE.md file with instructions for AI-assisted contribution workflows, a GLOSSARY.md, a House_Rules.md with the code of conduct, and a performance HTML file that the README links to.
The Performance Table and How to Read It
The README links to a grammar performance table hosted as an HTML file in the repository itself, accessible at the htmlpreview.github.io preview URL for performance.html. The README notes the table is sortable by column header.
Performance data for parser generators matters because a grammar that is correct but extremely slow to parse is not always usable in practice. Grammar ambiguities can cause ANTLR's LL(*) parsing algorithm to explore exponentially many paths, and the performance table gives visibility into which grammars have this characteristic.
The README does not explain which metrics the table shows, what inputs were used to produce the benchmarks, or what an acceptable performance figure looks like. That context lives in the table itself. A developer evaluating a grammar for use in a latency-sensitive tool should consult the performance table for their target grammar before integrating it.
The performance table is linked from the main README but is otherwise not integrated into the repository's navigation. It is generated separately and committed as a static file. The exact cadence of updates to the performance data is not documented in the README.
How to Use a Grammar from This Collection
To use a grammar from antlr/grammars-v4, the general workflow is to navigate to the relevant language directory, copy the .g4 file or files, and use the ANTLR tool to generate a parser and lexer in the target language. The repository does not document this workflow in the README; the README points to the Wiki for FAQ and to the ANTLR project itself for tooling documentation.
ANTLR v4 is a parser generator that takes a grammar file and produces source code for a parser, lexer, and optional visitor and listener in the chosen target language. The generated code is then compiled into the application. This separation between grammar authoring and parser generation is the standard ANTLR workflow; the grammars in this repository fit into that workflow without modification as long as they remain action-free.
The CLAUDE.md file in the repository root suggests the project uses AI-assisted tooling for some contribution workflows. The contents of that file are not in the materials available for this review, but its presence at the repository root indicates it contains instructions relevant to contributors. The GLOSSARY.md is similarly present but its contents are not available here.
For contributors, the House_Rules.md establishes the code of conduct and the guidelines for submitting grammars. The README points to the grammars-v4 Wiki for the full FAQ, which covers questions about grammar quality standards, supported ANTLR versions, and contribution expectations.
Limitations: Quality Variability and What the Repository Does Not Guarantee
antlr/grammars-v4 collects grammars from many contributors over many years. The README makes no claim that all grammars are complete or production-ready. For any given language in the collection, the grammar may cover the common subset of the language well and handle edge cases poorly, or it may target a specific version of the language that has since evolved, or it may have known ambiguities that cause performance problems on certain inputs.
The performance table gives quantitative data for some grammars, but it does not indicate whether a grammar correctly handles all valid programs in the target language. Completeness is a different question from performance, and the repository does not publish completeness test results for individual grammars.
The action-free constraint is described as an 'expectation' rather than a hard requirement. The README does not specify what happens when a submitted grammar includes actions, only that the expectation is for grammars to be action-free. Some grammars in the collection may not fully meet this expectation.
For critical applications, for example building a production linter, a refactoring tool, or a security analysis pipeline, the appropriate approach is to validate the specific grammar against a test suite of programs in the target language before relying on it. The grammars in this collection are starting points, and many are well-tested, but the repository's scope (many languages, community contributors) means quality varies.
The repository is actively maintained, with the last push on 2026-09-26. Pull request merge builds and weekly builds are tracked via GitHub Actions, as shown by the build status badges at the top of the README.
Contributing: Code Style and the Review Process
The README directs contributors to open an issue before submitting a pull request for changes, particularly for significant ones. This is the same contribution pattern used in the companion wireguard-install project; it is a common practice for single-maintainer or small-team projects to manage the review queue and avoid duplicate work.
The full contribution guidelines are in House_Rules.md and the grammars-v4 Wiki rather than in the README. The README refers to both as the authoritative sources for conduct and contribution expectations. This keeps the README short and the detailed guidance in dedicated documents that can be updated without changing the primary repository page.
The build infrastructure uses GitHub Actions for both per-commit builds triggered by pull request merges and periodic weekly builds. The README shows two separate workflow badges: one for the 'Last PR merge build' and one for the 'Weekly .jar build' and 'Weekly dev build.' The weekly builds appear to test grammar compilation against the ANTLR .jar, providing a recurring check that grammars still compile correctly against the current ANTLR release.
The MIT licence applies to the repository. Individual grammars contributed over the years may have been written under different conditions, but the repository-level licence is MIT, which permits use, modification, and redistribution with attribution.
Editorial conclusion
Teams that need a ready-made ANTLR v4 grammar for a common language will find this collection the first place to check. Those needing a production-grade parser for a high-stakes use case should verify the specific grammar's completeness independently, since the README makes no quality guarantees for individual grammars. The MIT licence permits use and modification. The action-free design means grammars in this collection need no adaptation when the ANTLR target runtime changes. The repository was last pushed on 2026-09-26.
Frequently asked questions
Is ANTLR still used?
The antlr/grammars-v4 repository received its last push on 2026-09-26 and has active weekly build workflows running against the current ANTLR release, which indicates the tool and its grammar ecosystem are in use. ANTLR v4 is used in compilers, IDEs, linters, and analysis tools across many projects.
Can I use ANTLR with C#?
ANTLR v4 supports C# as a target runtime, and antlr/grammars-v4 includes a csharp/ directory containing a C# grammar. The action-free design of grammars in this collection means grammars written for other targets can also be used to generate C# parsers without modification.
How do I find a grammar for a specific language in antlr/grammars-v4?
Navigate to the repository root and look for a directory named with the all-lowercase name of the language or format. The README states this naming convention explicitly: the Java grammar is in java/, C++ in cpp/, C# in csharp/. If no directory exists for the target language, the README does not document a fallback.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/antlr-grammars-v4)