Open-source project
chdb-io/chdb avatar
chdb-io/chdb

chDB: An In-Process ClickHouse Engine for Python, and Its Pandas-Compatible DataStore API

chDB is an in-process OLAP SQL Engine 🚀 powered by ClickHouse

2,906 stars134 forksPythonApache-2.0

At a glance

What is it?
chDB embeds the ClickHouse OLAP engine in the Python process, so you query Parquet, CSV and Arrow files with SQL and no server. Its DataStore layer goes further and tries to run pandas code through the same engine.
Who is it for?
Adopt chDB when you want ClickHouse SQL and ClickHouse format support inside a Python process, especially for local analysis of Parquet or CSV files that never reach a server. Do not adopt it expecting a general-purpose transactional store or a drop-in replacement for pandas across every method.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 7 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 25, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The gap chDB fills between pandas and a ClickHouse server

The usual way to get ClickHouse performance in a Python workflow is to run a ClickHouse server and connect to it. That means a deployment, a port, credentials, and a network hop for data that may already be sitting on the same machine as a Parquet file. chDB takes the other route: the README describes it as an in-process SQL OLAP Engine powered by ClickHouse, and states that you do not need to install ClickHouse. The engine runs inside your Python process.

The audience is narrow but real. Data scientists and analysts who already write ClickHouse SQL, and who want to query local files in Parquet, CSV, JSON, Arrow, ORC or the 60+ other formats the README lists, are the primary fit. So are Python developers who need analytical SQL over files without operating a database. The project is Apache-2.0 licensed, with Python as the primary language, and the pyproject.toml classifies it as Production/Stable for Python 3.9 through 3.14.

How the engine and the DataStore layer actually fit together

There are two distinct surfaces in this repository, and conflating them is the most common source of confusion.

The first is chDB itself: an embedded engine. The README notes that data copies from C++ to Python are minimized using the Python memoryview interface, which matters because the cost of moving a large result set across the language boundary is often what kills embedded analytics. Input and output support spans Parquet, CSV, JSON, Arrow, ORC and more, and the package claims Python DB API 2.0 support, so existing DB-API style code has a path in.

The second surface is DataStore, a pandas-compatible API living in the datastore/ directory. Its design is lazy: operations are recorded rather than executed, then compiled into SQL. The README describes dual-engine execution, where a QueryPlanner routes each segment of the operation chain to either chDB for SQL or pandas for operations the SQL path cannot express, with intermediate results cached at each step. That routing decision is the interesting engineering, and also the part most likely to surprise you. A chain such as select, filter, sort, head is a natural fit for SQL. A chain containing an arbitrary Python callable is not, and will fall back. The README does not document where that boundary sits for every method.

Installing chDB and running a first query on a Parquet file

The README gives a single install command for the core package. It states support for Python 3.9+ on macOS and Linux, on x86_64 and ARM64.

bash
pip install chdb

If you want Arrow Database Connectivity support, the README requires Python 3.10 or later and a separate extra:

bash
pip install "chdb[adbc]"

There is also a browser path with no install at all. The README points at wasm.chdb.io, where chDB-Wasm runs the engine compiled to WebAssembly, and the corresponding source lives in the chdb-io/chdb-wasm repository.

For a first real use, the DataStore API is the shortest route from a file to a result. The README shows loading a file and then chaining operations:

python
from datastore import DataStore

ds = DataStore.from_file("data.parquet")

result = (ds
    .select("product", "revenue", "date")
    .filter(ds.revenue > 1000)
    .sort("revenue", ascending=False)
    .head(10))

print(result)

You should see the top ten rows by revenue, filtered to revenue above 1000, with only the three named columns. The README notes that these operations are lazy, so nothing executes until the result is materialized by print or an equivalent call.

The same DataStore object can point at remote sources. The README lists S3 with anonymous access, MySQL, PostgreSQL, SQLite, MongoDB, ClickHouse, HDFS, Azure and GCS, using a URI form:

python
from datastore import DataStore

ds = DataStore.uri("s3://bucket/data.parquet?nosign=true")

The nosign=true parameter is the README's way of requesting anonymous S3 access. Do not assume it works against a private bucket.

The pandas compatibility claim, and where it stops

DataStore's pitch is that you change one import and keep writing pandas. The README's example is literally import datastore as pd, followed by DataFrame construction, boolean filtering and groupby mean, with output shown as ordinary pandas output.

The coverage numbers in the README are specific: 209 DataFrame methods, 56 Series.str accessor methods, 42+ Series.dt accessor methods, and 334 ClickHouse SQL functions. Those figures are the project's own accounting, and the README points to dev-docs/PANDAS_COMPATIBILITY.md for the full list and dev-docs/FUNCTIONS.md for the function reference.

Treat the counts as a map, not a guarantee. A method appearing in a compatibility guide does not tell you whether it executes in chDB or falls back to pandas, and the README does not publish that per-method mapping. The honest position is that DataStore is worth evaluating on your own operation chain, not on the aggregate number. The README also states the API is recommended, which is a positioning choice worth noting: the project is steering new users toward the pandas-shaped surface rather than the raw SQL one.

Constraints: platform coverage, lazy evaluation and the wrong use cases

The README is explicit that chDB supports Python 3.9+ on macOS and Linux, x86_64 and ARM64. Windows is not listed. If your team develops on Windows, that is a blocker before any code is written, and the browser build at wasm.chdb.io is not a substitute for a local install in a CI pipeline.

Lazy evaluation is the second constraint. It is a feature when it lets the planner push work into SQL, and a debugging problem when it does not. If a chain silently falls back to pandas, you may be materializing a full dataset into memory to run an operation you assumed was pushed down. The README's caching of intermediate results helps iterative exploration and can also hold memory you did not plan for. The README does not document cache eviction or a memory ceiling.

This is also the wrong tool for transactional work. There is nothing in the README describing row-level updates, multi-writer concurrency or ACID transactions in the OLTP sense. It is an OLAP engine. If you need a general-purpose embedded relational store for application state, this is not it. Likewise, if your data already lives in a well-run ClickHouse cluster, adding an in-process copy of the engine solves a problem you do not have.

chDB against DuckDB and against running ClickHouse itself

The comparison people actually search for is chDB versus DuckDB, and the difference is lineage rather than category. Both are embedded analytical engines you install with pip and call from Python. chDB's engine is ClickHouse, which means your SQL dialect, function set and format handling are ClickHouse's. DuckDB has its own dialect and its own extension model. If your team already writes ClickHouse SQL, chDB removes a translation step that DuckDB would reintroduce. If you have no ClickHouse background, that advantage disappears and you are choosing between two dialects on other grounds.

The second comparison is chDB versus a ClickHouse server. The README's framing is that chDB removes the need to install ClickHouse, which is the whole point: no separate process, no connection configuration, no server to keep alive for a laptop-scale analysis. The trade-off is that you give up what a server provides, including shared access for multiple clients and the operational tooling around a long-running instance. chDB is a library in your process; a server is infrastructure. Pick based on whether the data needs to outlive the script.

On the packaging side, chDB depends on chdb-core, pandas and pyarrow, with adbc-driver-manager added by the adbc extra. The pyproject.toml also references a chdb.durable extra for durable analytical objects whose state lives in S3-compatible object storage, with boto3 needed for the S3-compatible backend and separate opt-in extras for native GCS and Azure Blob. That extra is documented in the package metadata rather than the README body, so read pyproject.toml before assuming what a plain install gives you.

Maintenance, releases and what the Apache-2.0 licence lets you do

The repository is not archived, and the last push was on 2026-09-23. Recent releases are v4.4.0 on 2026-09-11, v4.3.0 on 2026-08-17 and v4.2.1 on 2026-07-13, so the release cadence over that window is roughly monthly. Note that the version in pyproject.toml reads 3.7.0 while the release tags read 4.x; the repository includes a VERSION-GUIDE.md, which is where that discrepancy is presumably explained, and a CHANGELOG.md for release notes.

Upgrade cost is concentrated in two places. The Python package pins chdb-core>=26.7.0, and the pyproject.toml comments state that the backup, restore and query-analysis ABI arrived in 26.7, which is why chdb-core is named in the durable extras as well. That means the engine version and the Python wrapper version move on separate tracks, and a chdb-core bump can change behaviour underneath a stable-looking chdb release. The second place is DataStore's pandas compatibility surface. If you depend on a specific method's behaviour, a release that changes how it is routed to the SQL engine is a behavioural change even though the method name is unchanged.

chDB is Apache-2.0, the same licence as ClickHouse itself. That permits commercial and closed-source use and modification, with the usual obligations around attribution and notice files. This is a description of the licence text, not legal advice; have counsel review it if you are embedding the engine in a distributed product.

Editorial conclusion

Adopt chDB when you want ClickHouse SQL and ClickHouse format support inside a Python process, especially for local analysis of Parquet or CSV files that never reach a server. Do not adopt it expecting a general-purpose transactional store or a drop-in replacement for pandas across every method. Verify first that your platform is covered (the README states Python 3.9+ on macOS and Linux, x86_64 and ARM64) and that the DataStore methods you rely on appear in the Pandas Compatibility Guide before you rewrite working pandas code.

Frequently asked questions

What is chDB and what is it used for?

chDB is an in-process SQL OLAP engine powered by ClickHouse, distributed as a Python package. It is used to run analytical SQL over files and data sources such as Parquet, CSV, JSON and Arrow without installing or connecting to a ClickHouse server.

How do I install chDB in Python?

The README gives pip install chdb, and states support for Python 3.9+ on macOS and Linux on x86_64 and ARM64. ADBC support needs Python 3.10 or later and installs with pip install "chdb[adbc]".

What is the difference between chDB and running ClickHouse as a server?

chDB runs the ClickHouse engine inside your Python process, so the README states there is no need to install ClickHouse and no separate server to connect to. A ClickHouse server is a separate long-running process that multiple clients can share, which chDB is not.

Does chDB work with pandas code?

The DataStore API is described in the README as pandas-compatible, and its example changes the import to import datastore as pd and then uses ordinary DataFrame, filter and groupby syntax. Operations are lazy and compiled into SQL, with a QueryPlanner routing segments between the chDB engine and pandas.

Which operating systems and Python versions does chDB support?

The README states that chDB currently supports Python 3.9+ on macOS and Linux, on both x86_64 and ARM64. Windows is not listed among the supported platforms.

Can I run chDB without installing anything?

Yes. The README points to wasm.chdb.io, where chDB-Wasm runs the complete ClickHouse engine compiled to WebAssembly in the browser, with the source in the chdb-io/chdb-wasm repository. That browser build is separate from the pip package.

Official sources

  1. chdb-io/chdb on GitHub
  2. License: Apache-2.0
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/chdb-io-chdb.svg)](https://hysenlabs.com/projects/chdb-io-chdb)