# sqlparse: a non-validating SQL parser for Python formatters, linters and editors

> sqlparse tokenizes SQL into a tree of statements, clauses and identifiers without validating it or committing to a dialect. It is a building block for tooling, not a database client, and its formatting output is a convenience rather than a guarantee.

**andialbrecht/sqlparse** — A non-validating SQL parser module for Python

- Repository: https://github.com/andialbrecht/sqlparse
- Stars: 4,019 · Forks: 754
- Language: Python
- License: BSD-3-Clause
- Published: 2026-09-23 · Updated: 2026-09-23 · Language: en
- Canonical page: https://hysenlabs.com/projects/andialbrecht-sqlparse

## The gap sqlparse fills: reading SQL without a database

Most Python code that touches SQL never needs to understand it. It builds a string, hands it to a driver, and moves on. The moment you want to do something to the SQL itself, such as split a migration file into individual statements, reindent a query for a log line, or find the table names in a query, you need a parser. Writing one is a project in itself, and the obvious alternative is to let the database parse it, which means a connection and a round trip.

sqlparse occupies the space in between. The README calls it a non-validating SQL parser, and that phrase carries the design. It tokenizes the text, groups the parts it recognizes into a tree, and accepts any input without checking whether the statement is legal. It also makes no assumptions about a particular SQL dialect, so vendor extensions and templated SQL parse too. That is why tools that must handle SQL they did not write, formatters, linters, editors, query analysis, tend to reach for it.

The audience is therefore narrow and specific: developers building SQL-facing tooling in Python, and anyone who needs to manipulate SQL text as data. If you want to execute queries, this is the wrong library. If you want to know whether a query is valid for PostgreSQL, this is also the wrong library, because validation is explicitly out of scope.

## How the token tree works: split(), format() and parse()

Three module-level functions cover most needs: split(), format() and parse(). They are three views of the same tokenizer.

split() cuts a script into statements. The README gives the example of "select * from foo; select * from bar;" returning a list of two strings, each with its trailing semicolon. That is a text-level operation: no tree is built and nothing is validated.

format() runs the tokenizer and then re-emits the tokens with layout and casing options applied. The README shows reindent=True and keyword_case="upper" turning a one-line select into a multi-line statement with SELECT, FROM and WHERE uppercased while the identifiers and the literal 1 stay as written. The same options are available on the command line, including strip_comments and indent_width.

parse() returns the tree. Calling parse() on a statement and indexing element zero gives a Statement object, and get_type() reports what kind of statement it is. Iterating statement.tokens yields the top-level parts. The README's output shows the distinction that matters: SELECT is a leaf token with a ttype, while IdentifierList and Where are groups. Tokens with a ttype are leaves; the others are groups you can descend into, which is how nested constructs like subqueries and CASE expressions are represented. Recursion is the traversal model, and the documentation's Analyzing SQL statements page covers the traversal API in more detail.

One consequence of the design is worth stating plainly. Because the parser does not validate and does not target a dialect, the tree reflects what the tokenizer recognized, not what the database would accept. Unknown syntax does not raise; it lands as generic tokens.

## Installing sqlparse and formatting a query from the command line

The install is a single pip command. The package declares requires-python >=3.10 in pyproject.toml, so an older interpreter will refuse the install rather than fail at runtime.

```bash
pip install sqlparse
```

Installing the package also puts a sqlformat command on your path; pyproject.toml maps the console script to sqlparse.__main__:main. It reads from stdin when the filename is -, writes to stdout by default, and rewrites files in place with --in-place. The README's example formats a file named query.sql with reindented output and uppercased keywords:

```bash
sqlformat --reindent --keywords upper query.sql
```

Run sqlformat --help for all formatting options. The same work can be done from Python, which is useful when the SQL arrives as a string rather than a file. The README's formatting example uses reindent and keyword_case:

```python
import sqlparse

print(sqlparse.format("select id,name from users where active=1",
                      reindent=True, keyword_case="upper"))
```

The expected output is a multi-line statement with SELECT, FROM and WHERE uppercased and the column list broken across lines. If you are wiring this into a repository, sqlparse ships a pre-commit hook. The README's configuration pins a revision and passes args; note its warning that the hook defaults to --in-place --reindent, and that when you override args you must keep --in-place or the hook writes to stdout and leaves your files unchanged.

```yaml
repos:
  - repo: https://github.com/andialbrecht/sqlparse
    rev: 0.5.5  # use the latest release
    hooks:
      - id: sqlformat
        args: [--in-place, --reindent, --keywords, upper]
```

## Where sqlparse stops: no validation, no dialect guarantees

The limitation is in the name. A non-validating parser cannot tell you that a statement is wrong. Feed it a query with a missing parenthesis or a column that does not exist and it will tokenize what it can and return a tree. For a formatter that is the right behavior. For anything that gates a deployment it is the wrong one.

The second constraint follows from dialect neutrality. Because sqlparse makes no assumptions about a particular SQL dialect, it cannot apply dialect-specific rules. A formatter that knows PostgreSQL and one that knows BigQuery produce different canonical output for the same text, and sqlparse has no setting for that distinction. Vendor extensions and templated SQL parse, but they parse as whatever the tokenizer can recognize, which means the tree for a heavily templated file may be shallower than you expect.

There is also a maintenance reality the README states directly: sqlparse is maintained in spare time, so it can take a while before an issue or pull request gets a reply. The repository's last push was on 2026-08-13, and the project is not archived. That is a healthy signal for a mature library, but it also means you should not plan around a fast turnaround on a bug that affects your formatting edge case.

Finally, formatting is not a contract. The README does not document idempotence guarantees for format() across all inputs, so if you run a formatter in a pre-commit hook, the safe assumption is that you should check the diff rather than trust that a second run produces identical output.

## sqlparse vs SQLGlot: tokenizer and formatter against a transpiler

The comparison people search for is sqlparse vs SQLGlot, and the difference is architectural rather than a matter of feature checklists. sqlparse is a tokenizer plus a tree plus a formatter. SQLGlot is a transpiler: it parses SQL into an abstract syntax tree with a defined semantic model, and can emit that tree back out as a different dialect. That requires knowing what the tokens mean, which is exactly the assumption sqlparse refuses to make.

The practical consequence: if your job is to convert a MySQL query into a PostgreSQL query, sqlparse will not do it and is not trying to. If your job is to split a file of statements, reindent them, or find identifiers in a query whose dialect you do not control, sqlparse's neutrality is an advantage, because it will not reject input for being from the wrong vendor.

There is a cost to that neutrality, and it is worth being honest about it. A parser that does not know the semantics of what it reads can only offer structural operations. Anything requiring meaning, type inference, column resolution, dialect rewriting, is outside the design. Choosing sqlparse is choosing a smaller, more predictable surface over a larger one that demands more of the input.

## Licence, packaging and the cost of keeping it current

sqlparse is licensed under the New BSD license, identified as BSD-3-Clause in the repository metadata. That is a permissive licence, and the README notes that parts of the code are based on Pygments, written by Georg Brandl and others, so the file headers are worth reading if you are redistributing a modified copy. This is a description of what the repository states, not legal advice; if your organisation has a licence review process, run the LICENSE file through it.

The packaging story is deliberately small. The build backend is hatchling, the version is read dynamically from sqlparse/__init__.py, and the only optional dependency groups are dev (build) and doc (sphinx, furo). There is no runtime dependency list, which is the main reason the library travels well inside other tools. The supported interpreters are CPython and PyPy on Python 3.10 through 3.14, and the Makefile runs the test suite against each of those five versions with uv.

Upgrade cost is low but not zero. Formatting output can shift between releases, so a version bump that changes layout will show up as a large diff in any repository using the pre-commit hook. The Makefile's benchmark target runs the scripts under benchmarks/, which is the place to look if you are tracking tokenizer performance across versions. Pin the revision in your pre-commit config, as the README's example does, and read the release notes before moving the pin.

## Conclusion

Adopt sqlparse when you need to split a script into statements, reformat SQL for display, or walk a token tree in Python without caring which dialect produced the text. Do not adopt it if you need semantic validation, dialect-aware rewriting, or a guarantee that formatting is idempotent for every vendor extension. Before committing, run sqlformat --help against the version you pin, check that your SQL survives a round trip through format(), and confirm the pre-commit hook args include --in-place.

## FAQ

### What does sqlparse do?

It tokenizes SQL text and groups the parts it recognizes into a tree of statements, clauses, identifiers and expressions. It is non-validating and makes no assumptions about a particular SQL dialect, so it is used as a building block for formatters, linters, editors and query analysis tools.

### How do I install sqlparse?

Installing the package with pip install sqlparse also provides the sqlformat command. The package requires Python 3.10 or later.

### How do I use sqlparse in Python?

Three module-level functions cover most needs: split() cuts a script into statements, format() re-emits tokens with layout and casing options, and parse() returns the token tree. The README's parse example calls get_type() on a statement and iterates statement.tokens, where tokens with a ttype are leaves and the others are groups you can descend into.

### What are the differences between SQLGlot and sqlparse?

sqlparse tokenizes and formats without validating or targeting a dialect. SQLGlot parses into an abstract syntax tree with a semantic model and can emit a different dialect, which requires knowing what the tokens mean. If you need dialect conversion, sqlparse is not the tool.

### Is there a Python library that can parse SQL?

sqlparse is one, described in its README as a non-validating SQL parser for Python. It accepts any input without validating it and makes no assumptions about a particular SQL dialect.

## Sources

- [andialbrecht/sqlparse on GitHub](https://github.com/andialbrecht/sqlparse)
- [Issues](https://github.com/andialbrecht/sqlparse/issues)
- [License: BSD-3-Clause](https://github.com/andialbrecht/sqlparse/blob/master/LICENSE)
- [README](https://github.com/andialbrecht/sqlparse/blob/master/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/andialbrecht-sqlparse
