# apache/datafusion-sqlparser-rs: an extensible SQL lexer and parser for Rust

> The sqlparser crate turns SQL text into a typed Rust AST across many dialects, and deliberately stops before semantics. It fits engines, linters and analysis tools, not applications that need to know whether a query is valid.

**apache/datafusion-sqlparser-rs** — Extensible SQL Lexer and Parser for Rust

- Repository: https://github.com/apache/datafusion-sqlparser-rs
- Stars: 3,467 · Forks: 780
- Language: Rust
- License: Apache-2.0
- Published: 2026-09-23 · Updated: 2026-09-23 · Language: en
- Canonical page: https://hysenlabs.com/projects/apache-datafusion-sqlparser-rs

## What sqlparser solves, and who it is for

SQL text arrives as a string. Engines, formatters, linters, migration tools and lineage scanners all need the same first step: turn that string into a structure the rest of the program can walk. sqlparser does that step and nothing after it. The README describes the crate as "a lexer and parser for SQL that conforms with the ANSI/ISO SQL standard and other dialects", and says it is used as "a foundation for SQL query engines, vendor-specific parsers, and various SQL analysis".

The audience follows from that. If you are writing a query engine in Rust, or a tool that inspects queries rather than runs them, this crate is the parsing layer you would otherwise write yourself, badly, over several months. If you need to execute SQL, this is not the library. It has no planner, no catalog and no execution. The output is an AST, and the README is explicit that the crate "provides only a syntax parser, and tries to avoid applying any SQL semantics".

## The parse pipeline: dialect in, AST out

The mechanism is small enough to describe in one sentence: you pick a Dialect, hand the parser a string, and get back a Vec of Statement nodes. The README example uses GenericDialect, and notes that AnsiDialect or a custom dialect are alternatives. Dialects are where the vendor-specific behavior lives, so the choice is not cosmetic. A statement that parses under one dialect can fail under another.

The AST is a plain Rust enum tree. In the README's example, a SELECT becomes Query, then Select, then projection entries such as UnnamedExpr(Identifier("a")) and UnnamedExpr(Function(...)), a from list of TableWithJoins, and a selection holding a BinaryOp tree for a > b AND b < 100. That shape is the whole point: once your query is a tree of typed variants, matching on it is ordinary Rust.

Two design choices deserve attention. First, the crate "Preserves Syntax Round Trip": apart from comments, whitespace and keyword capitalization, parsing then printing should return the original SQL. The README shows assert_eq!(ast[0].to_string(), sql) for "SELECT 'hello'", and a pretty-printed form via the {:#} format specifier. Second, source locations come from the Spanned trait, which the README labels a work in progress: "many nodes report missing or inaccurate spans". If you are building diagnostics that point at a column in an editor, budget time for the gaps.

## Installing the crate and parsing your first statement

The crate is published as sqlparser on crates.io, and Cargo.toml gives version 0.63.0 with license Apache-2.0. Add it to a Rust project the usual way:

```bash
cargo add sqlparser
```

Then parse a statement. This is the README's own example, trimmed to the essential lines:

```rust
use sqlparser::dialect::GenericDialect;
use sqlparser::parser::Parser;

let sql = "SELECT a, b FROM table_1 WHERE a > b ORDER BY a DESC, b";
let dialect = GenericDialect {};
let ast = Parser::parse_sql(&dialect, sql).unwrap();
println!("AST: {:?}", ast);
```

Running this prints the AST as a debug string, starting with Query and descending into Select, projection, from and order_by. The repository also ships examples/parse_select.rs and examples/cli.rs, and Cargo.toml defines a json_example feature that enables JSON output in the cli example, which pulls in serde_json and serde.

Features are worth setting deliberately. The default set is std plus recursive-protection; the latter uses the recursive crate for stack overflow protection. The serde feature implements Serialize and Deserialize for all AST nodes, and visitor adds a Visitor that recursively walks the tree. If you disable recursive-protection, deeply nested expressions become your problem.

## It parses syntax, not meaning

The README states the limitation directly: CREATE TABLE(x int, x int) is accepted, even though most SQL engines reject it for the repeated column name x. That is not a bug report. It is the boundary the project chose. Semantic analysis "varies drastically between dialects and implementations", so the maintainers left it to consumers.

The practical consequence is that a successful parse proves very little. A statement can parse cleanly and still be nonsense for your database. If your product promises users that a query is valid, sqlparser gives you half the answer, and you own the other half: catalog lookups, type checks, name resolution.

Compliance claims are equally bounded. The README says the parser supports most SQL-92 syntax plus newer syntax that has been requested, and that the online SQL:2016 grammar guides what to accept. It then admits that "stating anything more specific about compliance is difficult", because no public test suite can assess it automatically, and advises anyone assessing fit to "experimentally verify whether it supports the subset of SQL that you need". Treat that as the honest state of the project rather than a marketing hedge.

## Dialect coverage and the round-trip guarantee in practice

Dialect support is the feature people most often misread. A Dialect in this crate is a configuration of the parser, not a compatibility certificate. The README's own phrasing is that the crate accepts queries specific databases would reject "even when using that Database's specific Dialect". So choosing a dialect narrows the grammar; it does not make the parser agree with that vendor's server.

The round-trip property is the more reliable one, and it is genuinely useful. If you are rewriting a query, adding a filter or redacting a column, you can parse, mutate the AST, print it back, and expect the result to differ from the input only in comments, whitespace and keyword case. The README also acknowledges that some cases collapse: "different SQL with seemingly similar semantics are represented with the same AST", and invites pull requests to distinguish them. That matters for tools that must reproduce input exactly, such as migration diffing, where two spellings that map to one AST will not survive the trip separately.

## Where sqlparser is the wrong tool

Skip it if you need a database. There is no execution, no schema, no type inference, and no query planning. Wrapping sqlparser does not get you closer to running SQL; it gets you a tree.

Skip it if you need strict validation. The syntax-only stance means your validation layer is yours to build, and the README points you at semantic analysis as something to do on top of the project, not inside it.

Be careful if you need precise error locations today. Spanned is described as a work in progress with missing or inaccurate spans, and the README links an issue for contributing improvements. Tooling that underlines the exact offending token may need more work than the trait's presence suggests.

Finally, if your host language is not Rust, the crate is a Rust library. There is no bundled CLI binary in the repository layout beyond the examples, and no documented bindings for other languages. Porting the AST across an FFI boundary is a project of its own.

## Alternatives and how their approach differs

The most direct comparison is with parser generators such as LALRPOP or pest. With those, you write a grammar and they generate a parser, which means you own the grammar, the AST types and every dialect variation. sqlparser hands you a maintained grammar and AST for SQL specifically, at the cost of accepting its node shapes. If your language is not SQL, a generator is the better fit; if it is SQL, you are trading control for a large head start.

On the Python side, sqlparse is a non-validating SQL parser that tokenizes and formats statements. Its model is text-oriented rather than a typed AST, so it suits formatting and splitting scripts better than structural analysis. sqlparser's typed enum tree is the difference that matters when you want to match on expression variants rather than on token groups.

Within Rust, the parent project Apache DataFusion consumes this crate as its SQL front end, which is the clearest signal of intended use: parsing as the entry point to a query engine, with planning and execution living elsewhere. If you want the whole engine, DataFusion is the layer above; if you want only the front end, this crate is it.

## Maintenance, upgrades and licence

The repository is not archived, and the last push was on 2026-09-23. It is an Apache Software Foundation project, with .asf.yaml, LICENSE.TXT and NOTICE.TXT at the top level, and the crate is licensed Apache-2.0. That licence is permissive and includes an explicit patent grant; it also requires preserving the NOTICE file and stating changes. This is a description of the terms, not legal advice, and the LICENSE.TXT in the repository is the text that governs.

Upgrade cost is the part teams underestimate. The crate is at 0.63.0, which under Cargo's rules means every minor bump is a breaking change, and the AST is the public surface. A version bump can alter enum variants your code matches on. The repository keeps a CHANGELOG.md and a changelog/ directory, so the diff between versions is documented, but the compiler is what will actually find your breakages. Pin the version and read the changelog before moving.

There is no separate runtime to operate. The dependency is compiled into your binary, and the only operational cost is binary size and build time, plus whatever the recursive-protection feature costs you in stack safety if you turn it off.

## Conclusion

Adopt sqlparser if you are building a Rust engine, formatter, linter or lineage tool and can accept a syntax-only parse. Do not adopt it if you need queries validated against a real schema, since it accepts statements most engines reject. Before committing, take your own dialect's trickiest statements and check the AST and the round trip, because the README says compliance claims beyond SQL-92 are hard to make and the Spanned trait is still a work in progress.

## FAQ

### What is apache/datafusion-sqlparser-rs?

It is a Rust crate named sqlparser that provides a lexer and parser for SQL conforming to ANSI/ISO SQL and other dialects. The README describes it as a foundation for SQL query engines, vendor-specific parsers and SQL analysis.

### How is sqlparser different from SQLGlot or sqlparse?

sqlparser produces a typed Rust AST and deliberately avoids semantic analysis, so it accepts statements engines would reject. sqlparse is a Python non-validating parser that works at the token and formatting level rather than exposing a typed expression tree.

### What does SQL parsing mean in sqlparser?

In this crate, parsing means turning a SQL string into a Vec of Statement AST nodes using a chosen Dialect, with no catalog, planning or execution involved. The README states the crate provides only a syntax parser and tries to avoid applying SQL semantics.

## Sources

- [apache/datafusion-sqlparser-rs on GitHub](https://github.com/apache/datafusion-sqlparser-rs)
- [Issues](https://github.com/apache/datafusion-sqlparser-rs/issues)
- [License: Apache-2.0](https://github.com/apache/datafusion-sqlparser-rs/blob/main/LICENSE)
- [README](https://github.com/apache/datafusion-sqlparser-rs/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/apache-datafusion-sqlparser-rs
