# ingestr: a single-command data copier between Postgres, BigQuery, DuckDB and more

> ingestr is a Go CLI from Bruin Data that moves a table from one database to another with one command. It is quick to try, but the documentation leaves schema evolution and incremental-load semantics largely to the reader.

**bruin-data/ingestr** — ingestr is a CLI tool to copy data between any databases with a single command seamlessly.

- Repository: https://github.com/bruin-data/ingestr
- Website: https://getbruin.com/docs/ingestr/
- Stars: 3,987 · Forks: 157
- Language: Go
- License: NOASSERTION
- Published: 2026-09-23 · Updated: 2026-09-23 · Language: en
- Canonical page: https://hysenlabs.com/projects/bruin-data-ingestr

## What ingestr actually does, and who ends up using it

ingestr is a command-line application that copies data from a source into a destination. The README frames it as a way to "copy data from any source to any destination without any code", and the whole user interface is a set of flags on one subcommand. There is no configuration file to write for a basic run, no scheduler, and no transformation step.

The people this suits are engineers who already have two systems and want a table to appear in the second one. A Postgres table into BigQuery, a CSV file into DuckDB, a Kafka topic into a warehouse. The repository topics list bigquery, duckdb, mssql, postgresql and snowflake, which matches the README's support table: most of the well-known relational and analytical stores appear as both source and destination, while streaming systems such as Kafka, Pulsar, Amazon SQS, Azure Event Hubs and Google Cloud Pub/Sub appear as sources only. That asymmetry is the clearest statement of intent in the whole repository. ingestr pulls from event streams and writes to storage, not the other way around.

It is not a data integration platform in the sense of a hosted service with connectors you configure in a browser. It is one binary you run, either by hand, from a cron job, or from a CI step.

## The mechanism: URI in, URI out, one subcommand

The core of ingestr is the ingest subcommand, which takes a source URI, a source table, a destination URI and a destination table. Everything else, including credentials, lives inside those URIs as query parameters. The README's BigQuery example passes credentials_path as a query parameter on the bigquery:// URI, and the Postgres example passes sslmode=disable the same way. That is the entire configuration model for a basic run.

The README also lists three incremental loading modes by name: append, merge and delete+insert. It does not describe the semantics of each in the README itself, so anyone relying on merge or delete+insert should read the linked documentation before scheduling a job. What the README does state is that ingestr handles the transfer without a backend, which means the process reads from the source and writes to the destination in one pass rather than staging through a service you have to operate.

The repository layout supports the same reading. There is a cmd directory, an internal directory, a pkg directory, and a Go module that pulls in driver libraries for BigQuery, ClickHouse, Cassandra, DynamoDB, Athena, S3, Spanner, Pub/Sub, Iceberg, Pulsar, Db2 and more. The Dockerfile builds a single static binary with CGO enabled and copies it into a Debian slim image under a non-root user. So the shipped artifact is one executable, not a set of services.

## Installing ingestr and running your first copy

The README gives two installation paths. The install script is the shortest:

```bash
curl -LsSf https://getbruin.com/install/ingestr | sh
```

Alternatively, the pip package installs the same tool, and the README notes that the pip package can also be used from Python:

```bash
pip install ingestr
```

For the Python SDK path, which pulls in pyarrow, the README uses an extra:

```bash
pip install 'ingestr[sdk]'
```

The first real use is the quickstart, which copies a Postgres table into BigQuery. Note that the source URI carries the credentials and the destination URI carries the path to a service account JSON file:

```bash
ingestr ingest \
    --source-uri 'postgresql://admin:admin@localhost:8837/web?sslmode=disable' \
    --source-table 'public.some_data' \
    --dest-uri 'bigquery://<your-project-name>?credentials_path=/path/to/service/account.json' \
    --dest-table 'ingestr.some_data'
```

According to the README, this command reads public.some_data from the Postgres instance and uploads it to BigQuery under the schema ingestr and the table some_data. If the run succeeds you should find that table in BigQuery with the rows from the source. The README does not document what happens on a partial failure or how to resume, so verify the row count on both sides after the first run.

There is also a Python entry point for data that is already in memory. The README shows a list of dictionaries, a generator, and a DataFrame all going through the same call:

```python
import ingestr

ingestr.ingest(
    [{"id": 1, "name": "Ada"}, {"id": 2, "name": "Grace"}],
    dest_uri="duckdb:///tmp/warehouse.duckdb",
    dest_table="main.people",
)
```

The README states that rows, generators and DataFrames are sent to the ingestr CLI binary as Arrow IPC streams by default, and that the pip package downloads and caches the matching GitHub release binary on first use. That download-on-first-use behaviour is worth knowing about if you run in an environment without outbound network access.

## Where the single-command model runs out

The first limitation is that the README documents no rollback. If a copy fails halfway, nothing in the README says whether the destination table is left partially written, whether the previous contents survive, or how to recover. For append mode that may not matter, since a rerun adds rows. For merge and delete+insert, a partial run against a production table is exactly the kind of event you want documented, and it is not.

Second, the README lists the three incremental modes but does not explain them. Anyone who has used a similar tool knows that merge usually needs a primary key and delete+insert usually needs a comparison column, but the README does not say how ingestr selects either. That is a gap you have to close by reading the external documentation, not by reading the repository.

Third, the support table is asymmetric in ways that will surprise people. Apache Iceberg is a destination only. Apache Pulsar, Amazon SQS, Azure Event Hubs, Google Cloud Pub/Sub, Kafka, InfluxDB, IBM Db2, GCP Spanner, Couchbase and HTTP are sources only. If your plan is to write into Kafka or read out of Iceberg, ingestr is the wrong tool and no flag will change that.

Finally, the pip metadata classifies the project as "Development Status :: 4 - Beta". That is the package author's own label, and it is worth weighing against the release cadence. The repository is not archived, and the last push was on 2026-09-23, the same day as the most recent releases v1.1.55, v1.1.54 and v1.1.53 on 2026-09-22, 2026-09-11 and 2026-09-10. Frequent point releases at the 1.1.x level suggest active work, and also suggest that pinning a version is wise.

## ingestr compared with dlt, and what the difference means in practice

The most common comparison people search for is ingestr against dlt. Both move data between systems without asking you to write a pipeline, but they sit at different layers.

dlt is a Python library. You write a Python script, import the library, and run it. The pipeline is code you own and can extend, and the extraction logic is yours to shape. ingestr is a binary. You pass URIs on the command line and the tool does the rest. There is nothing to import unless you deliberately use the Python SDK, and even then the README describes the SDK as a thin front end that serialises your rows to Arrow IPC and hands them to the CLI binary, not as an extensible framework.

That difference decides the choice more than any feature list. If your team already writes Python pipelines and wants to add custom extraction logic, dlt fits that shape. If your team wants a cron job that copies a table and does not want to maintain a Python environment for it, ingestr is the smaller thing to operate. The trade-off is that ingestr gives you less room to intervene when the copy does something you did not expect.

The other comparison in the search data is Bruin against dbt. Bruin is the vendor behind ingestr, and dbt is a transformation tool. That comparison is about the parent project's transformation story, not about ingestr, which does not transform anything. Reading it as an ingestr comparison would be a mistake.

## Maintenance cost, licensing and what to check before a rollout

The repository is not archived and the last push was on 2026-09-23, with three releases in the twelve days before that. The version numbers are still in the 1.1.x range and the Python package declares Beta status, so expect the flag surface and the connector list to keep moving. Pinning a specific release in your automation is the practical response; the install script fetches whatever is current, which is convenient for a first look and less convenient for a scheduled job.

Licensing needs care here. The repository metadata reports the licence as NOASSERTION, which means the automated detection could not classify the LICENSE file. The Python package metadata, by contrast, declares "License :: OSI Approved :: MIT License". Those two statements do not agree, and the repository also ships a THIRD_PARTY_LICENSES.txt and a licenses.lock.yml, which suggests the dependency tree carries its own obligations. If you plan to redistribute ingestr or embed it in a product, read the LICENSE file and the third-party notices yourself rather than trusting either classifier. This is not legal advice.

The upgrade cost is mostly the usual one for a fast-moving CLI: flags and URI parameters can change between point releases. The Python SDK adds a second moving part, since the pip package downloads and caches a matching release binary on first use. In an air-gapped environment that download has to be arranged separately, and the README does not describe an offline path.

## Conclusion

ingestr fits teams that need a repeatable one-command copy between two systems and are willing to test the incremental modes against their own data before putting them on a schedule. It is a poor fit if you need a full transformation layer, CDC, or a documented rollback story, because the README describes none of those. Before adopting it, run the Postgres to BigQuery quickstart against a throwaway schema, then run the same command twice with each of the append, merge and delete+insert modes and inspect the destination table after the second run.

## FAQ

### How do I install ingestr?

The README gives two options: the install script at https://getbruin.com/install/ingestr, or pip install ingestr. The pip package can also be installed with the sdk extra for Python data ingestion.

### Which incremental loading modes does ingestr support?

The README names three: append, merge and delete+insert. It does not describe how each one behaves, so the semantics need to be checked in the external documentation before you schedule a job.

### Can ingestr write data into Kafka?

No. The README's support table lists Kafka as a source only, with no destination entry. The same applies to Apache Pulsar, Amazon SQS, Azure Event Hubs, Google Cloud Pub/Sub, InfluxDB and IBM Db2.

### Is ingestr maintained?

The repository is not archived and the last push was on 2026-09-23. Releases v1.1.55, v1.1.54 and v1.1.53 were published on 2026-09-22, 2026-09-11 and 2026-09-10 respectively, and the Python package metadata still labels the project as Beta.

## Sources

- [bruin-data/ingestr on GitHub](https://github.com/bruin-data/ingestr)
- [Issues](https://github.com/bruin-data/ingestr/issues)
- [Project website](https://getbruin.com/docs/ingestr/)
- [README](https://github.com/bruin-data/ingestr/blob/main/README.md)
- [Releases](https://github.com/bruin-data/ingestr/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/bruin-data-ingestr
