Ibis: One Python Dataframe API Across DuckDB, Postgres and Snowflake
the portable Python dataframe library
At a glance
- What is it?
- Ibis compiles lazy Python dataframe expressions into SQL for more than 20 backends. It suits pipelines that must run locally on DuckDB and then move to a warehouse, but it is not a pandas drop-in and the SQL it emits is the contract you depend on.
- Who is it for?
- Adopt Ibis if your analysis starts on a laptop and has to end in a warehouse, and you are willing to treat the generated SQL as part of your code review. Skip it if you need pandas-specific behaviour, in-place mutation, or per-row Python execution, since the expression API is lazy and columnar.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository received new commits within the last day.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The problem Ibis solves: one expression, many SQL engines
Analysts write pandas code against a local file, then rewrite it in SQL when the same logic has to run on a warehouse. Ibis exists to remove that second rewrite. The README describes it as "the portable Python dataframe library" and lists its goals directly: fast local dataframes via DuckDB by default, lazy dataframe expressions, an interactive mode for exploration, and the same dataframe API across more than 20 backends. The intended audience is developers and research users, which matches the classifiers in pyproject.toml (Intended Audience :: Developers and Intended Audience :: Science/Research).
The portability claim is narrower than it sounds. Ibis does not move your data between engines. It moves your expression. For most backends the README states that Ibis works by compiling dataframe expressions into SQL, so what travels is the query, and the engine you point it at executes it. That is why the backend list matters more than any benchmark: the list covers Apache DataFusion, Apache Druid, Apache Flink, Apache Impala, Apache PySpark, Athena, BigQuery, ClickHouse, Databricks, DuckDB, Exasol, Materialize, MySQL/MariaDB, Oracle, Polars, PostgreSQL, RisingWave, SingleStoreDB, SQL Server, SQLite and Snowflake, and the truncated README entry suggests the list continues.
How Ibis works: expressions in, SQL out
An Ibis table is a lazy expression, not a materialised frame. You chain methods such as group_by, agg and order_by, and nothing executes until you ask for a result. The README's getting started example builds a grouped count, then calls ibis.to_sql on it and prints a SELECT with an inner aggregate subquery, aliases t0 and t1, a COUNT(*) AS "count", a GROUP BY on positional columns 1 and 2, and an ORDER BY on the alias. Reading that generated SQL is the fastest way to understand what a chain of Ibis calls actually means.
The Python and SQL sides are not separate worlds. The README shows t.sql(...) accepting a raw SQL string that selects from the penguins table, returning an Ibis expression that can then be extended with ordinary Ibis methods such as order_by. So a team can keep a hand-written SQL fragment where it is clearer and continue in Python afterwards.
Dependencies are deliberately small: atpublic, parsy, python-dateutil, sqlglot, toolz, typing-extensions and tzdata. sqlglot is the SQL generation and translation layer, with a pinned exclusion of version 26.32.0. Optional backends and features live in extras rather than the base install, which is why the install command in the README names them explicitly.
Installing Ibis and running a first grouped count
The README's install command pulls the framework together with the DuckDB backend and the bundled example datasets. The extra is named duckdb, and examples is a second extra in the same bracket.
pip install 'ibis-framework[duckdb,examples]'The README points to the installation guide at ibis-project.org/install for other options. In the Python session below, interactive mode is switched on so expressions render as tables instead of staying lazy, ibis.examples.penguins.fetch() loads the example dataset, and the group_by call produces one row per species and island with a count. The output shown in the README has five rows, from Adelie on Biscoe with 44 up to Gentoo on Biscoe with 124.
import ibis
ibis.options.interactive = True
t = ibis.examples.penguins.fetch()
g = t.group_by("species", "island").agg(count=t.count()).order_by("count")
gTo see what will be sent to the engine, pass the expression to ibis.to_sql. The README's output wraps the aggregate in a subquery and orders by the count column. If that SQL is not what you would have written, the expression is the thing to change.
Where Ibis stops: lazy expressions and backend gaps
The most common mismatch is expecting pandas semantics. Ibis expressions are lazy and columnar; there is no in-place mutation of a frame, and the API is built around group_by, agg and similar verbs rather than the full pandas surface. Code that relies on per-row Python functions or on pandas indexing idioms does not translate directly.
Backend coverage is per operation, not per library. Compiling to SQL means each method has to have a SQL equivalent the target engine accepts, and the README documents more than 20 backends rather than promising that every operation is available on every one. If you are on an engine with unusual SQL, the generated query is where the difference shows up.
The repository's own development setup reflects this breadth. compose.yaml defines ClickHouse, MySQL (as a mariadb image) and SingleStoreDB services with their own ports and healthchecks, and the justfile has a sync recipe that takes a backend name and installs that backend's extra plus examples and geospatial. Testing across that matrix is real work, and it is also why a niche backend can lag the common ones.
Ibis compared with pandas and Polars
pandas executes operations in memory on a materialised frame, with an index and a large API that many notebooks depend on. Ibis builds an expression and hands the work to an engine, which is why the same chain can run against DuckDB locally and Snowflake remotely. If your data fits comfortably in memory and your code is already pandas, Ibis adds a translation layer you do not need.
Polars is the closer comparison, and the difference is in what the expression targets. Polars is its own execution engine and appears in the Ibis backend list as a target, so the two can be used together: Ibis for the expression, Polars as the engine that runs it. Choosing between them is really a question of whether you want one engine with its own optimiser, or one API that can point at whichever engine the data already lives in.
Against writing SQL by hand, the trade is explicit. Hand-written SQL gives you exact control over the query the warehouse sees. Ibis gives you composability, reuse across engines, and a Python surface, at the cost of reviewing generated SQL rather than authoring it.
Maintenance, licensing and upgrade cost
The repository is not archived, and the last push was on 2026-09-21, one day before this was written, so development activity is current. Releases are not frequent: 12.0.0 is dated 2026-02-07, 11.0.0 is dated 2025-10-15, and 10.8.0 is dated 2025-07-28. Roughly one major release every three to four months, with a long tail of patch releases between them.
That cadence matters because major versions can change the expression API or the generated SQL. The project publishes release notes at ibis-project.org/release_notes, and the repository uses conventional commits with a .releaserc.js file, which suggests changelog generation is automated. Pin the version in your own project and read the release notes before moving between majors.
The licence is Apache-2.0, declared in both pyproject.toml and LICENSE.txt, with license-files pointing at that file. That is a permissive licence, but the practical question is not the licence text. It is whether the SQL Ibis generates for your backend is something your organisation is willing to run in production. That is a review question, not a legal one.
Editorial conclusion
Adopt Ibis if your analysis starts on a laptop and has to end in a warehouse, and you are willing to treat the generated SQL as part of your code review. Skip it if you need pandas-specific behaviour, in-place mutation, or per-row Python execution, since the expression API is lazy and columnar. Before committing, check three things in 12.0.0: that the backend you target is listed in the README's backend list, that the operations you need compile to SQL you would accept by hand, and that the optional extra for that backend installs cleanly on Python 3.10 or newer.
Frequently asked questions
What is Ibis software?
Ibis is a portable Python dataframe library. According to its README, it provides lazy dataframe expressions, fast local dataframes via DuckDB by default, an interactive mode, and the same dataframe API for more than 20 backends.
What is the Ibis database?
Ibis is not a database. For most backends the README states that it works by compiling dataframe expressions into SQL, and the resulting query runs on the database or engine you connect to, such as DuckDB, PostgreSQL or Snowflake.
Who supports Ibis?
The project lists Ibis Maintainers as both authors and maintainers in pyproject.toml, with a [email protected] contact. The README also links a Zulip chat at ibis-project.zulipchat.com for project discussion.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/ibis-project-ibis)