Open-source project
PeerDB-io/peerdb avatar
PeerDB-io/peerdb

PeerDB: Postgres CDC to ClickHouse, and What the Deprecations Mean for Your Pipeline

Fast, Simple and a cost effective tool to replicate data from Postgres to Data Warehouses, Queues and Storage

3,293 stars220 forksGoAGPL-3.0

At a glance

What is it?
PeerDB is an ETL tool built specifically for Postgres, driven through a Postgres-compatible SQL interface. Its README now marks most destination connectors as deprecated, leaving Postgres to ClickHouse and Postgres to Postgres as the maintained paths.
Who is it for?
Adopt PeerDB if your source is Postgres and your destination is ClickHouse or Postgres, and you want to define syncs in SQL rather than a connector UI. Do not adopt it if you need a maintained Kafka, S3, Snowflake or BigQuery destination: those connectors are deprecated and the README points to a migration guide for pinning a release or forking the code.
Can I use it commercially?
Yes, with strict conditions. AGPL-3.0 is a network copyleft licence: if people use a modified version over a network, for example as a hosted service, you must offer them its source code under the same licence.
Is it still maintained?
Yes. The repository received new commits within the last day.
What is it written in?
Mainly Go, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The Postgres-shaped gap PeerDB is aimed at

Most replication tools are built connector-first. They support a long list of sources and sinks, and Postgres is one entry among many. PeerDB inverts that. The README states the project exists because current data tools prioritize a wide range of connectors and often neglect to optimize for Postgres users, which pushes teams storing large volumes in Postgres into building custom pipelines. The intended audience is narrow and specific: teams running Postgres at the heart of their data stack who move data out of it frequently and at volume. The README also frames the interface as the differentiator. Instead of a connector configuration UI, PeerDB exposes a Postgres-compatible SQL interface for ETL, so the same client tools (psql, pgAdmin), the same BI tools (Grafana, Tableau) and the same migration tooling (Flyway) can be pointed at the sync layer. That is a real design commitment, not a marketing line: it means defining a mirror is closer to writing DDL than to filling in a form. If your team already lives in psql, the learning curve is mostly about PeerDB's own SQL surface, which the README does not reproduce and which lives in the docs.

Three streaming modes, and why only one is still maintained

PeerDB supports log-based CDC, cursor-based replication on a timestamp or integer column, and XMIN-based replication. The README is explicit that CDC is the recommended and actively maintained mirror type, while Query Replication (QRep) and XMIN are deprecated and no longer actively maintained. They remain fully functional and no code is currently being removed, so an existing QRep mirror will keep running. The distinction matters for new work. Log-based CDC reads the Postgres write-ahead log through logical replication, which is what allows schema changes and deletes to propagate and keeps latency in the range of tens of seconds. Cursor-based replication polls a column, so it sees inserts and updates that move that column but has no natural way to represent a delete. Choosing QRep for a new pipeline means choosing a mode the project has stopped developing, and the deprecation notice is the clearest signal in the README about where engineering effort is going.

The connector status matrix is the real adoption decision

The connector deprecation notice is the part of this README that should drive a buying decision. Snowflake, BigQuery, ElasticSearch, Kafka (including Confluent and Redpanda variants), Azure Event Hubs, Google Pub/Sub and S3 are all listed as deprecated destinations, no longer actively maintained. They remain functional in the current and all prior releases and no code is being removed, but the actively maintained paths going forward are Postgres to ClickHouse, Postgres to ClickHouse Cloud, and Postgres to Postgres. There is one asymmetry worth reading twice: the deprecation applies to the destination role only, and BigQuery remains a supported source. So a BigQuery-to-ClickHouse pipeline is on the maintained path while a Postgres-to-BigQuery pipeline is not. If your architecture depends on a deprecated destination, the README points to a deprecated connectors migration guide in docs/deprecated-connectors.md covering how to pin to a release or fork the relevant code. Pinning works until you need a fix. Forking means owning the connector. Neither is free, and the notice does not pretend otherwise.

Installing PeerDB with the quickstart Docker stack

The README's get started section clones the repository and runs a shell script that brings up the whole stack with docker-compose: Postgres as the catalog, Temporal, the PeerDB server, the flow API and workers, and the PeerDB UI. Docker and docker-compose are prerequisites. The local development path is different and additionally requires the buf compiler for protobuf generation.

bash
git clone [email protected]:PeerDB-io/peerdb.git
cd peerdb

# Run docker containers: postgres as catalog, temporal, PeerDB server, PeerDB flow API + workers, PeerDB UI
# Requires docker and docker-compose installed: https://docs.docker.com/engine/install/
bash ./run-peerdb.sh

For local development the README gives a second path, with images built locally rather than pulled.

bash
bash ./generate-protos.sh
bash ./dev-peerdb.sh

Once the stack is up, you connect to PeerDB itself over the Postgres wire protocol and issue SQL. The README specifies psql version 14.0 or later and gives the connection string directly.

bash
psql "port=9900 host=localhost password=peerdb"

What you should see is a psql prompt attached to the PeerDB server rather than to your source database. From there the README says you can query away; the actual mirror-creation SQL is not reproduced in the README and lives in the linked five-minute quickstart guide at docs.peerdb.io. That split is worth knowing before you start: the repository gets you a running stack, the documentation site gets you a working mirror.

The MinIO staging step that breaks non-Docker ClickHouse setups

The README flags one configuration detail that will bite anyone whose ClickHouse does not run inside the same Docker network. PeerDB stages PostgreSQL data in MinIO inside the Docker stack before loading it into ClickHouse. If ClickHouse runs on a VM or in ClickHouse Cloud, it needs a resolvable hostname for MinIO or the load step cannot read the staged files. The fix is an environment variable in docker-compose.yml, and the README gives the exact key and an example value.

yaml
AWS_ENDPOINT_URL_S3: http://172.31.26.57:9001 # Change this to IP/host which is accessible by both PeerDB and ClickHouse

The value must be reachable from both PeerDB and ClickHouse, which is why the default host.docker.internal form in the repository's docker-compose.yml does not transfer to a remote ClickHouse. Rerun Docker Compose after changing it. On AWS, GCP or Azure you also need the security group to allow inbound access to MinIO. This is a genuine architectural constraint rather than a bug: the staging step exists, and it imposes a network requirement that a pure streaming design would not. Budget time for it if your ClickHouse is managed.

Where PeerDB is the wrong tool

If your source is not Postgres, the README's own framing works against you. The project describes itself as built for PostgreSQL and as providing value when Postgres is at the heart of your data stack. The status section notes that the connector ecosystem has expanded to include MySQL, MongoDB, BigQuery and CockroachDB as sources, but the optimizations described (tuning Postgres configs, parallel reading of a replication slot, efficient TOAST column handling) are Postgres-specific, so a non-Postgres source does not get the same treatment. The second case is a deprecated destination. If you need Kafka, S3, Snowflake or BigQuery as a sink, the maintained path does not include you, and the honest options are pinning an older release or maintaining a fork. The third case is operational appetite. The stack includes Temporal for workflow orchestration alongside the catalog database, the server, the flow workers and the UI. That is a real cluster to run. A team that wants a single binary and no workflow engine will find this heavier than expected, and the README does not offer a lighter deployment mode. The README also makes no mention of rollback procedures for a mirror, so treat reversibility as something to establish yourself before pointing this at production.

PeerDB against Airbyte and ClickPipes

Airbyte takes the connector-catalog approach: a large library of sources and destinations, each implemented separately, configured through a UI or API. The trade-off is exactly the one PeerDB's README calls out. Breadth comes at the cost of depth in any single source, and Postgres-specific behaviors such as TOAST columns, schema changes and replication-slot tuning are not the center of that design. If you need to move data between ten heterogeneous systems, the catalog model wins. If you have one Postgres and one ClickHouse, you are paying for connectors you will not use. ClickPipes is the closer comparison, and PeerDB's README addresses it directly: PeerDB is available natively in ClickHouse Cloud, generally available, as the Postgres CDC connector in ClickPipes. That means the two are not purely rivals. Running the self-hosted PeerDB gives you control over the deployment, the catalog and the SQL interface; the ClickPipes route gives you the same engine without the Docker stack and Temporal. Which one fits depends on whether you want to operate the pipeline or hand it to the managed service.

Editorial conclusion

Adopt PeerDB if your source is Postgres and your destination is ClickHouse or Postgres, and you want to define syncs in SQL rather than a connector UI. Do not adopt it if you need a maintained Kafka, S3, Snowflake or BigQuery destination: those connectors are deprecated and the README points to a migration guide for pinning a release or forking the code. Before committing, verify that your ClickHouse instance can reach MinIO, since the quickstart stages files there, and check the connector status matrix against your exact source and destination pair.

Frequently asked questions

What is PeerDB?

PeerDB is an ETL tool built for PostgreSQL that streams data from Postgres to data warehouses, queues and storage engines. It supports log-based CDC, cursor-based and XMIN replication modes, and exposes a Postgres-compatible SQL interface for defining and monitoring syncs.

Is PeerDB open source?

The repository PeerDB-io/peerdb is public, with the code licensed under AGPL-3.0 according to the repository metadata. The README's own license badge points at an Elv2 license file, so the two do not agree and you should read LICENSE.md in the repository before relying on either.

How does PeerDB compare with Airbyte?

Airbyte is organized around a broad catalog of connectors, while PeerDB is built specifically for Postgres and optimizes for Postgres features such as TOAST columns, schema changes and replication-slot handling. PeerDB also exposes a Postgres-compatible SQL interface for ETL rather than a connector configuration UI.

How does ClickPipes compare with PeerDB?

They overlap: the README states PeerDB is available natively in ClickHouse Cloud as the Postgres CDC connector in ClickPipes, generally available. The difference is deployment, since self-hosting PeerDB means running the Docker stack with its catalog, Temporal, flow workers and UI.

What can I use as an alternative to PeerDB?

The README positions PeerDB against connector-catalog tools that support many sources but do not optimize for Postgres. If your destination is a deprecated PeerDB connector such as Kafka or S3, a tool that still maintains that sink is the practical alternative; if your destination is ClickHouse, the managed ClickPipes route uses the same engine.

Official sources

  1. License: AGPL-3.0
  2. PeerDB-io/peerdb on GitHub
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/peerdb-io-peerdb.svg)](https://hysenlabs.com/projects/peerdb-io-peerdb)