# ConnectorX: loading SQL results into DataFrames from Rust and Python

> ConnectorX reads a SQL query straight into a Pandas, PyArrow, Polars, Dask or Modin DataFrame, with partitioning done in the database. The design is narrow on purpose, and that narrowness is what you are adopting.

**sfu-db/connector-x** — Fastest library to load data from DB to DataFrames in Rust and Python

- Repository: https://github.com/sfu-db/connector-x
- Website: https://sfu-db.github.io/connector-x
- Stars: 2,651 · Forks: 225
- Language: Rust
- License: MIT
- Published: 2026-09-28 · Updated: 2026-09-28 · Language: en
- Canonical page: https://hysenlabs.com/projects/sfu-db-connector-x

## The problem ConnectorX solves is the download step, not the query

Most Python data work splits into two halves: getting rows out of a database, and doing something with them afterwards. The second half has plenty of good tools. The first half is usually handled by whatever driver the ORM already pulled in, and that path copies the data several times before it becomes a DataFrame. ConnectorX targets only that first half.

The audience is narrow: people who pull a large result set from Postgres, MySQL, MariaDB, SQLite, Redshift, ClickHouse, SQL Server, Azure SQL Database, Oracle, BigQuery or Trino, and want it as a DataFrame. The README states the goal plainly, describing ConnectorX as enabling you to load data from databases into Python in the fastest and most memory efficient way. It is a loader, not a database toolkit. There is no connection pooling story, no session object, no migration tooling. If your query is small, the difference between this and a driver will not matter to you, and the extra dependency is not worth it.

## Schema first, count second, partition third: the actual data flow

The README documents the sequence, and it is worth reading before you point this at a production replica, because the library issues more than one statement per call.

On receiving a query, ConnectorX first fetches the schema of the result set. For some sources this means issuing a LIMIT 0 query, for example SELECT * FROM lineitem LIMIT 0. If partition_on is set, it then runs SELECT MIN($partition_on), MAX($partition_on) FROM (SELECT * FROM lineitem) to learn the range of the partition column, and splits the original query into ranges such as SELECT * FROM (SELECT * FROM lineitem) WHERE $partition_on > 0 AND $partition_on < 10000. It follows with a count query per partition, or SELECT COUNT(*) FROM (SELECT * FROM lineitem) when no partition is given. The schema and the count are used to allocate memory up front, and only then does the download run, one thread per partition, writing rows or columns in a streaming fashion depending on the database.

So a single read_sql call can become several round trips. The count query in particular runs over the full result set, which is a real cost on a large table and something to keep in mind if your query is expensive to execute. The payoff is that memory is sized before the data arrives, and the README describes the library as following a zero-copy principle with data copied exactly once from source to destination. The partitioning itself is done in the database, not in Python, which is why the parallelism works without holding multiple copies in the client.

## Installing ConnectorX and a first partitioned read

The README gives a single install command. It pulls a prebuilt wheel, so there is no Rust toolchain involved; the documentation links to a separate page for building a Python wheel from source if you need that.

```bash
pip install connectorx
```

The minimal call takes a connection string and a query, and returns a DataFrame.

```python
import connectorx as cx

cx.read_sql("postgresql://username:password@server:port/database", "SELECT * FROM lineitem")
```

To use more than one thread, name a partition column and a partition count. The README's example uses partition_on="l_orderkey" and partition_num=10, and states that the function splits the column evenly across the partitions, assigning one thread per partition.

```python
import connectorx as cx

cx.read_sql("postgresql://username:password@server:port/database", "SELECT * FROM lineitem", partition_on="l_orderkey", partition_num=10)
```

If the query runs, you get a DataFrame back and nothing else to configure. If it fails on the partition step, the cause is almost always the partition column: it must be numerical and cannot contain NULL, and the query must be a select-project-join-aggregate shape. The README states those constraints directly, and they are the two things most likely to trip up a first attempt.

## Partitioning rules are stricter than they first look

The parallel path is the selling point, and it is also the part most likely to be unusable on a real schema. Three conditions have to hold at once: the partition column must be numerical, it cannot contain NULL, and the query must be an SPJA query. A surrogate integer key satisfies all three. A natural key, a UUID, a timestamp with gaps, or anything nullable does not.

The reason is the min/max split. ConnectorX computes the range of the column and divides it evenly, so a column with a skewed distribution produces partitions of very uneven row counts even though the numeric ranges are equal. The README says evenly splitting the specified column, which is a statement about the value range, not about the number of rows. On a table where one key range holds most of the data, one thread will finish long after the others, and the wall-clock gain is smaller than the thread count suggests. That is a property of range partitioning, not a bug, but it is worth knowing before you attribute a disappointing speedup to the library.

Without partition_on the load is single-threaded, and the count query still runs. On a very large unpartitioned table that count is pure overhead on top of the download.

## Federated queries across two databases, marked experimental

The README calls federated query support experimental, and the wording is worth taking literally. You pass a dictionary of named connection strings instead of one, and reference each name as a schema prefix in the SQL.

```python
import connectorx as cx
db1 = "postgresql://username1:password1@server1:port1/database1"
db2 = "postgresql://username2:password2@server2:port2/database2"
cx.read_sql({"db1": db1, "db2": db2}, "SELECT * FROM db1.nation n, db2.region r where n.n_regionkey = r.r_regionkey")
```

The README states that joins from the same data source are pushed down by default, and points to Federation.md for setup and configuration. That pushdown rule is the important detail: a join between two tables in the same database stays in the database, while a join across db1 and db2 has to be executed somewhere in the client. The README does not document how the cross-source join is planned or what its memory behaviour is. Treat this as a feature to evaluate on your own data rather than one to design around.

## Where ConnectorX is the wrong tool

It is a read path. There is no write path in the README, no transaction control, and no cursor semantics. If your job involves loading a staging table, running several statements in one transaction, or streaming rows into application code as they arrive, this is not the library for that and you should not try to make it one.

The driver coverage is also a real boundary. ODBC is listed as work in progress, so a database reachable only through an ODBC driver is out. MariaDB is supported through the MySQL protocol, Redshift through the Postgres protocol, and Azure SQL Database through the MSSQL protocol, which means you inherit the behaviour of that protocol rather than a dedicated implementation.

The third boundary is the destination. The README lists Pandas, PyArrow, Modin, Dask and Polars, with Modin and Dask going through Pandas and Polars going through PyArrow. A destination outside that list is not available. And because the library issues its own schema and count queries, pointing it at a source where extra statements are expensive or restricted (a warehouse with per-query billing, a read-only replica with strict statement limits) deserves a second thought before you swap out an existing loader.

## ConnectorX against SQLAlchemy and pandas.read_sql

The comparison people actually make is with pandas.read_sql, usually sitting on top of SQLAlchemy. The difference is architectural rather than incremental. pandas.read_sql goes through the DBAPI: the driver fetches rows into Python objects, and pandas then builds arrays from them. SQLAlchemy adds a layer above the driver for dialect handling, connection management and query construction.

ConnectorX removes both layers from the read path. It talks the database protocol directly in Rust, and the README describes the result as data copied exactly once, from source to destination, with no intermediate Python objects. That is the whole argument for the library, and it explains why the trade-off is a much smaller surface: no engine abstraction, no ORM, no connection pool, no query builder. You keep SQLAlchemy for the rest of the application and use ConnectorX for the bulk read, which is a normal arrangement. What you give up is the flexibility to treat the two interchangeably.

The README's own benchmark section compares solutions that provide read_sql, loading a 10x TPC-H lineitem table (8.6GB) from Postgres into a DataFrame with 4 cores of parallelism, and reports up to 3x less memory and 21x less time, or 3x less memory and 13x less time against Pandas. Those are the project's numbers on its own benchmark, not an independent measurement, and they are specific to that table, that source and that core count.

## Release cadence, licence and what upgrading costs

The repository is not archived, and the last push was on 2026-09-23. Releases have been unevenly spaced: v0.4.4 on 2025-08-26, v0.4.5 on 2026-01-18, v0.4.6 on 2026-09-18. That is roughly two releases a year, with a gap of about eight months between 0.4.5 and 0.4.6. Plan upgrades around that rhythm rather than expecting frequent patch drops.

The version number is still 0.x, which matters for upgrade cost. The public Python surface is small (read_sql and its arguments), so the blast radius of a breaking change is limited, but there is no long-term support branch documented in the README, and no migration guide for the 0.4.x line. Pin the version in your requirements and read the release notes before moving.

The licence is MIT, which is permissive and imposes no copyleft obligation on your code. That is a statement about the licence text, not legal advice; if you redistribute the wheel or ship it inside a product, have your own counsel look at the notices you need to carry.

## Conclusion

Use ConnectorX when the job is a bulk read of a single query into a DataFrame and the source is one of the listed engines; skip it when you need a general SQLAlchemy-style connection layer, transaction control, or partitioning on a column that can be NULL. Verify first that your source is in the supported list, that your partition column is numeric and non-null, and that the destination library you want is among Pandas, PyArrow, Modin, Dask or Polars.

## FAQ

### What is ConnectorX and what is it used for?

It is a library, written in Rust with Python bindings, that loads the result of a SQL query into a DataFrame. The README describes it as loading data from databases into Python in the fastest and most memory efficient way, with Pandas, PyArrow, Modin, Dask and Polars as destinations.

### How do I install ConnectorX in Python?

The README gives one command, pip install connectorx, which installs a prebuilt wheel. Building a Python wheel from source is covered on a separate documentation page linked from the README.

### How do I make ConnectorX read a query in parallel?

Pass partition_on with a numerical column that cannot contain NULL, plus partition_num for the number of partitions. The README's example uses partition_on="l_orderkey" and partition_num=10, and states that one thread is assigned per partition.

### Does ConnectorX work with SQL Server and Oracle?

Yes. The README lists SQL Server and Oracle among the supported sources, along with Azure SQL Database through the mssql protocol. Connection strings and supported data types for each source are documented on the project's databases page.

## Sources

- [License: MIT](https://github.com/sfu-db/connector-x/blob/main/LICENSE)
- [Project website](https://sfu-db.github.io/connector-x)
- [README](https://github.com/sfu-db/connector-x/blob/main/README.md)
- [Releases](https://github.com/sfu-db/connector-x/releases)
- [sfu-db/connector-x on GitHub](https://github.com/sfu-db/connector-x)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/sfu-db-connector-x
