# Citus: Distributed PostgreSQL as an Extension

> Citus turns a single PostgreSQL node into a sharded cluster without leaving Postgres. Here is what the extension actually does, how to run it in Docker, and where the design stops being the right answer.

**citusdata/citus** — Distributed PostgreSQL as an extension

- Repository: https://github.com/citusdata/citus
- Website: https://www.citusdata.com
- Stars: 12,795 · Forks: 798
- Language: C
- License: AGPL-3.0
- Published: 2026-09-21 · Updated: 2026-09-21 · Language: en
- Canonical page: https://hysenlabs.com/projects/citusdata-citus

## What Citus adds to a PostgreSQL installation

Citus is not a fork and not a separate database server. It is a PostgreSQL extension, which means it installs into an existing Postgres and adds catalog tables, planner hooks and background workers rather than replacing the server. The README describes it as transforming Postgres into a distributed database, and lists five capabilities: distributed tables sharded across a cluster of PostgreSQL nodes, reference tables replicated to all nodes, a distributed query engine that routes and parallelizes SELECT and DML, columnar storage for compression and faster scans, and the ability to query from any node.

The audience is narrow and specific. The README names two reasons developers choose it. The first is a single Postgres node that is outgrowing its hardware: high CPU utilization, I/O wait times, queries returning out of memory errors, autovacuum failing to keep up and increasing table bloat. The second is the opposite motivation, wanting Postgres features such as advanced joins, user-defined functions, upsert, constraints, foreign keys, PostGIS and HyperLogLog at a scale a single machine cannot reach. If neither describes you, the extension adds operational surface for no gain.

The workloads the README calls out by name are multi-tenant applications, time series and IoT data for real-time analytics, and general high transaction throughput. Those three share a property that matters: each has an obvious column to shard on, such as tenant ID or time.

## How sharding and the distributed query engine fit together

The mechanism is sharding at the table level. A distributed table is split into shards placed on worker nodes, and the coordinator holds the metadata that maps shard ranges to workers. Reference tables take the other path: they are replicated in full to every node so that joins and foreign keys from distributed tables work without shipping data across the network. That split is the core design decision. A reference table must be small enough to copy everywhere, and a distributed table must be joined on its distribution column to avoid reshuffling rows between workers.

The query engine sits between the client and the shards. According to the README, it routes and parallelizes SELECT, DML and other operations across the cluster. Routing means a query filtered on the distribution column goes to one shard on one worker. Parallelizing means an aggregate or a scan across all shards is split and the results combined. Batch operations and analytical queries spread across all cores; short transactions stay on a single worker when the filter allows it.

Columnar storage is a separate feature layered on the same tables. The README states it compresses data, speeds up scans and supports fast projections, on both regular and distributed tables. The build system reflects this separation: the top-level Makefile has a columnar target that builds src/backend/columnar independently of the distributed code in src/backend/distributed, and install-columnar exists as its own target. That is a useful signal about coupling: columnar can be built and installed on its own.

Scaling out is not a one-way operation. The README says you can add more worker nodes and rebalance the shards when data size and volume grow. It does not document what happens to in-flight queries during a rebalance, and the README does not describe rollback of a rebalance.

## Installing Citus and running a first query

The README gives two paths. The managed path is Azure Cosmos DB for PostgreSQL, where Azure handles backups, high availability through auto-failover, software updates and monitoring. The self-managed path starts with Docker, and the README calls the single-node container the smallest possible Citus cluster.

The container maps host port 5500 to Postgres port 5432 and sets the superuser password through POSTGRES_PASSWORD. The image name is citusdata/citus:

```bash
docker run -d --name citus -p 5500:5432 -e POSTGRES_PASSWORD=mypassword citusdata/citus
```

You can then open psql inside the container, which avoids installing a client locally:

```bash
docker exec -it citus psql -U postgres
```

Or connect from the host with a local psql against port 5500:

```bash
psql -U postgres -d postgres -h localhost -p 5500
```

For a local PostgreSQL installation, the README points at the packaging repo. On Ubuntu or Debian, the script adds the repository and the package installs the extension for a specific Postgres major version:

```bash
curl https://install.citusdata.com/community/deb.sh > add-citus-repo.sh
sudo bash add-citus-repo.sh
sudo apt-get -y install postgresql-17-citus-13.0
```

On Red Hat the equivalent package is named citus130_17, which encodes the Citus and Postgres versions in the same way:

```bash
curl https://install.citusdata.com/community/rpm.sh > add-citus-repo.sh
sudo bash add-citus-repo.sh
sudo yum install -y citus130_17
```

After installing, the extension has to be registered in the database. The README says to add an entry to postgresql.conf; the truncated text cuts off before showing the exact key, so check the installation page for the precise line rather than guessing it. The version pairing in the package names is the part to pay attention to: citus-13.0 is built against postgresql-17, and installing a mismatched pair is the most common way this goes wrong.

## Where Citus is the wrong tool

Cross-shard joins are the boundary. The README presents reference tables as the answer for joins and foreign keys from distributed tables, which works when one side of the join is small and replicated everywhere. It offers no equivalent for joining two large distributed tables on a column that is not the distribution column. If your schema is a graph of large tables joined on several different keys, sharding will turn those joins into network shuffles or reject them outright, and the performance you were chasing disappears.

Schema design is therefore not optional. The distribution column is a decision made early and embedded in every hot query. Choosing tenant ID works for a multi-tenant application where every query is scoped to one tenant. Choosing poorly means rewriting queries later, and the README does not describe a supported path for changing the distribution column of a live distributed table.

There is also an operational cost the README does not quantify. A Citus cluster is a set of PostgreSQL nodes you now have to run. The managed service exists precisely because that is real work. The README mentions a high availability section in its table of contents but the text here does not cover failover, backup or upgrade procedures for a self-managed cluster.

Finally, single-node Postgres is genuinely good. If your data fits in memory on one machine and your queries are not saturating CPU, adding a coordinator, workers and shard metadata makes the system harder to reason about for no measurable benefit.

## Citus compared with other ways to scale PostgreSQL

The closest alternative in spirit is Patroni, and the difference is what gets scaled. Patroni manages PostgreSQL high availability: streaming replication with automatic failover, so one primary serves writes and standbys serve reads and take over when the primary fails. That scales read capacity and availability, not write throughput or storage. Citus splits the data itself across nodes so writes and storage spread too. If your problem is that the primary dies, Patroni addresses it; if your problem is that the primary cannot keep up, Patroni does not.

Against CockroachDB and YugabyteDB, the difference is architectural rather than feature-level. Those are distributed databases that speak the PostgreSQL wire protocol. Citus is PostgreSQL, running real Postgres on every node, so extensions such as PostGIS and the full set of Postgres features are available rather than reimplemented. The trade is that Citus inherits Postgres's single-node storage engine and its sharding model, including the distribution-column constraint described above.

Against Greenplum, the split is workload. Greenplum is built for analytical warehousing, while the README positions Citus for both transaction throughput in multi-tenant apps and analytical queries, with columnar storage as an option on the same tables rather than the whole system. Against Vitess, the difference is the layer: Vitess shards MySQL at the proxy level, while Citus changes the database itself and keeps SQL semantics inside Postgres. Against TimescaleDB, the overlap is time series and the divergence is scope: TimescaleDB optimizes time-partitioned data inside a single Postgres, while Citus distributes across machines.

## Maintenance, versioning and the AGPL licence

The repository is not archived, and the last push was on 2026-09-18, three days before this writing. Recent releases cluster on 2026-08-06: v14.2.0, v13.4.0 and v12.1.14. Three maintained major lines receiving releases on the same day is the relevant fact here. It means a Citus upgrade is not a single linear track; you pick a major line and stay on it, and the packaged builds in the README reflect that pairing, with postgresql-17-citus-13.0 and citus130_17 both encoding a Citus 13 build for Postgres 17.

That pairing is the upgrade cost. Moving to a new PostgreSQL major version means waiting for a matching Citus build, and the README's install commands show the version is baked into the package name rather than resolved automatically. The repository also carries an EXTENSION_COMPATIBILITY.md file at the top level, which is where the compatibility matrix lives; check it before planning an upgrade rather than assuming the newest Citus supports the newest Postgres.

Building from source is the other path. The top-level Makefile errors out with a message that ./configure needs to be run before compiling Citus, and the extension target depends on a generated header, src/include/citus_version.h. The build also requires PG_CONFIG to be set, since the Makefile derives the extension directory from it. That is a normal PostgreSQL extension build, but it means source installs need the matching Postgres server development files.

The licence is AGPL-3.0. The README states the database is 100 percent open source. AGPL is a copyleft licence with a network clause, and its obligations differ from permissive licences in ways that matter when you distribute software or offer the database as a service. That is a question for your own legal review, not something this article can settle; the LICENSE and NOTICE files at the repository root are the authoritative text.

## Conclusion

Adopt Citus when you already run PostgreSQL and the bottleneck is one node's CPU, memory or I/O, and when your workload has a natural distribution key such as tenant ID. Do not adopt it if you need cross-shard joins on arbitrary columns, or if you want a managed database and are unwilling to run the cluster yourself; the README points to Azure Cosmos DB for PostgreSQL for that case. Before committing, verify three things against your own schema: that every hot query filters on the distribution column, that your reference tables are small enough to replicate to every worker, and that your upgrade path stays inside a supported major version, since the extension tracks Postgres releases and the packaged builds in the README are pinned to specific pairings such as postgresql-17-citus-13.0.

## FAQ

### What is Citus?

Citus is a PostgreSQL extension that turns Postgres into a distributed database. It adds distributed tables sharded across a cluster of PostgreSQL nodes, reference tables replicated to every node, a distributed query engine, and columnar storage.

### How is Citus different from standard PostgreSQL?

Standard PostgreSQL runs on one node, while Citus shards distributed tables across a cluster of PostgreSQL nodes to combine their CPU, memory, storage and I/O capacity. Citus is an extension on top of Postgres, so the Postgres features and tools remain available.

### How do I install Citus?

The README gives three routes: the Azure Cosmos DB for PostgreSQL managed service, a single Docker container run with docker run -d --name citus -p 5500:5432 -e POSTGRES_PASSWORD=mypassword citusdata/citus, or the packaging repo, where the deb script installs postgresql-17-citus-13.0 and the rpm script installs citus130_17.

### What is Citus Postgres?

It is the same thing as Citus: a PostgreSQL extension that distributes tables across a cluster of PostgreSQL nodes. The README describes it as transforming Postgres into a distributed database for high performance at scale.

### What is Citus database?

Citus is a PostgreSQL extension rather than a standalone database server, and the README states the Citus database is 100 percent open source. It is licensed AGPL-3.0 and the primary implementation language is C.

### How does Citus compare with CockroachDB?

CockroachDB is a distributed database that speaks the PostgreSQL wire protocol, while Citus runs actual PostgreSQL on every node as an extension. The README's argument for Citus is that Postgres features such as PostGIS, HyperLogLog, user-defined functions and foreign keys work at scale rather than being reimplemented.

## Sources

- [citusdata/citus on GitHub](https://github.com/citusdata/citus)
- [License: AGPL-3.0](https://github.com/citusdata/citus/blob/main/LICENSE)
- [Project website](https://www.citusdata.com)
- [README](https://github.com/citusdata/citus/blob/main/README.md)
- [Releases](https://github.com/citusdata/citus/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/citusdata-citus
