Open-source project
redpanda-data/connect avatar
redpanda-data/connect

Redpanda Connect: a declarative stream processor in one YAML file

Fancy stream processing made operationally mundane

8,771 stars974 forksGoLicense varies

At a glance

What is it?
Redpanda Connect moves data between sources and sinks with Bloblang mapping, at-least-once delivery by default, and CDC connectors for Postgres, MySQL, MongoDB, Oracle and MSSQL. Here is how it installs, what the transaction model actually promises, and where it stops being the right tool.
Who is it for?
Adopt Redpanda Connect when your pipeline is a topology you can express in one YAML file and your delivery requirement is at-least-once: the in-process transaction model with no disk-persisted state makes horizontal scaling and restarts uneventful. Do not adopt it when you need exactly-once semantics across a sink that cannot deduplicate, or when your transformation is a stateful join over long time windows, because that is not what the processor is built around.
Can I use it commercially?
Not without permission. GitHub finds no licence file in the repository, and without a licence all rights are reserved by default: you may read the code but not reuse it. Check the README, or ask the authors, before using it.
Is it still maintained?
Yes. The repository received new commits within the last day.
What is it written in?
Mainly Go, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The problem Redpanda Connect solves, and who actually has it

Most data movement work is not hard because of the transformation. It is hard because of the plumbing around it: credentials for six systems, retry behaviour when a broker blips, a health endpoint so Kubernetes knows when to restart the pod, and a metrics sink so somebody can see throughput at 3am. Redpanda Connect takes that plumbing as its subject. The README describes it as "a stream processor that moves data between a wide range of sources and sinks, with support for hydration, enrichment, transformation, and filtering along the way."

That framing tells you who it is for. Platform and data engineers who already know which systems need to talk to each other and want the connection declared rather than coded. The repository topics list amqp, cqrs, etl, event-sourcing, kafka, nats, rabbitmq and stream-processing, which is a fair summary of the audience. If your job is writing a bespoke Go service that reads from Kafka, calls an HTTP API, writes to S3 and exposes Prometheus metrics, this project is a direct substitute for that service.

It is a weaker fit for teams whose transformation logic is the interesting part. Bloblang is a mapping language, and mapping is what it does. Anything resembling a stateful join across a long time window, or a model inference step with its own lifecycle, belongs somewhere else.

How the pipeline is assembled: inputs, processors, outputs, and Bloblang

The architecture is a topology, not a program. A configuration file declares an input, an optional chain of processors, and an output. The README calls these "declarative pipelines" and notes that a stream topology fits in a single YAML file. The processor list is where hydration, enrichment, transformation and filtering happen, and Bloblang is the language used for the mapping steps.

The README's own example is the clearest illustration of the data flow. A postgres_cdc input connects to a database with a DSN, names a schema, and lists tables to follow. An iceberg output writes to S3 through a Glue catalog, and the table name is set to a Bloblang interpolation: table: ${! meta("table") }. That single line is the mechanism worth understanding. The CDC input attaches metadata to each change event identifying its source table, and the output uses that metadata to route each record to the right Iceberg table, so one pipeline fans out into many destination tables without a separate output block per table.

The same example enables schema_evolution with a table_location, which is the project's answer to the fact that source schemas change. Rather than failing on an added column, the output evolves the destination schema.

Delivery is handled by an in-process transaction model. The README states that Redpanda Connect processes and acknowledges messages using this model with no disk-persisted state, and that when connecting at-least-once sources and sinks it guarantees at-least-once delivery even through crashes, disk corruption, or other server faults. Read that sentence carefully, because it contains the constraint: the guarantee is conditional on both ends being at-least-once. There is no exactly-once claim, and the absence of disk state is a deliberate trade, not an oversight.

Installing Redpanda Connect and running a first pipeline

Three install paths are documented. On Linux the README downloads an rpk archive and unzips it into a local bin directory. On macOS it uses a Homebrew tap. There is also a container image at docker.redpanda.com/redpandadata/connect. The Linux path is the one that gives you the rpk connect run command used throughout the documentation.

bash
curl -LO https://github.com/redpanda-data/redpanda/releases/latest/download/rpk-linux-amd64.zip
unzip rpk-linux-amd64.zip -d ~/.local/bin/

After unzipping, rpk should be on your PATH. The README points to a getting started guide for other options if that path does not suit your environment.

For macOS, the Homebrew tap is a single command, and it installs the same redpanda toolchain rather than a standalone connect binary.

bash
brew install redpanda-data/tap/redpanda

With Docker you pull the image directly, which is useful when you want to try a config without touching the host toolchain.

bash
docker pull docker.redpanda.com/redpandadata/connect

Running a pipeline means pointing the binary at a config file. The README gives this as the primary invocation.

bash
rpk connect run ./config.yaml

If you would rather not write a file first, the Docker form accepts inline overrides with -s flags, which is the fastest way to confirm the image works before committing to a config. The example below starts an HTTP server on port 4195 and forwards to Kafka, with the broker address and topic supplied on the command line.

bash
docker run --rm -p 4195:4195 docker.redpanda.com/redpandadata/connect run \
  -s "input.type=http_server" \
  -s "output.type=kafka" \
  -s "output.kafka.addresses=kafka-server:9092" \
  -s "output.kafka.topic=redpanda_topic"

Once something is running, two HTTP endpoints tell you whether it is healthy. The README documents /ping as a liveness probe that always returns 200, and /ready as a readiness probe that returns 200 once both input and output are connected and 503 otherwise. Wire /ready into your orchestrator rather than /ping, since a process that is alive but has not connected to its output will accept traffic it cannot deliver.

Where the at-least-once model becomes a real constraint

The delivery guarantee is the most consequential design decision in the project, and it is worth being blunt about the failure mode. At-least-once means duplicates are possible. If the process acknowledges a message and then crashes before the sink has durably recorded it, the source will redeliver. The README presents this as the default "with no caveats," which is honest, but it pushes the deduplication problem to the sink.

For some sinks that is free. Writing to an object store with deterministic keys, or to a database table with a primary key and an upsert, absorbs duplicates without extra work. For a sink that appends, such as a log store or a queue with no dedup window, every crash becomes a small amount of duplicated data. That is a property of the pipeline you are building, not a defect in the processor, but it is the thing to check before adopting.

The second constraint follows from the first. Because there is no disk-persisted state, the process cannot remember what it has already delivered across a restart. That is exactly what makes horizontal scaling and redeployment uneventful, and it is also why exactly-once is not on the menu. If your requirement is genuinely exactly-once, this is the wrong tool, and no amount of configuration will change that.

The README also notes that components linking against external C libraries, such as zmq4, are not included in the default build. You pull them in with the x_benthos_extra build tag, and the README warns that if the required system libraries are missing the build fails with an error like ld: library not found for -lzmq. It also states that this tag may change or be split into more granular tags in future releases, so anything depending on it should expect churn.

How it compares with a hand-written consumer and a full stream-processing framework

The obvious alternative is a small Go service that consumes from the source, transforms, and writes to the sink. That gives you total control over the transformation and over exactly-once semantics if you implement them. The cost is that you also own the connector for every system in the path, the retry logic, the health endpoints, the metrics exporter and the OpenTelemetry wiring. Redpanda Connect's README lists metrics to Statsd, Prometheus and a JSON HTTP endpoint, plus native OpenTelemetry traces, as built-in. Rebuilding that surface by hand is weeks of work that produces no product differentiation.

The other alternative is a full stream-processing framework with a distributed runtime, where state is partitioned across workers and the framework manages checkpoints. That approach handles stateful computation over windows in a way this project does not attempt. The difference in approach is where state lives: a framework keeps it in the runtime and coordinates it; Redpanda Connect keeps none and relies on the source and sink to provide the guarantee. If your job is a windowed aggregation with fault-tolerant state, choose the framework. If your job is moving records from A to B with a mapping in between, the framework's coordination machinery is overhead you will pay for and never use.

A third option worth naming is a managed connector service. Those remove the operational burden entirely but usually constrain which connectors you can use and how the mapping is expressed. Redpanda Connect keeps the config file in your repository and the binary in your container, which matters when the pipeline definition needs to go through the same review process as the rest of your code.

Maintenance, build and upgrade cost

The repository is not archived, and the last push was on 2026-09-21. Releases are frequent: v4.110.0 on 2026-09-17, v4.109.0 on 2026-09-10 and v4.108.0 on 2026-09-03, roughly a weekly cadence. That cadence is good for connector fixes and bad for anyone who wants a quiet dependency. Pin your image tag and read the CHANGELOG before moving.

Building from source requires a currently supported Go version, per the README, and go.mod declares go 1.26.6. The build uses task, not make. The Makefile in the repository opens with a deprecation warning directing you to taskfile.dev, so treat make targets as legacy. The documented build is task build:all, and the development loop is task fmt, task lint and task test, with golangci-lint for linting and gofumpt for formatting.

One upgrade consideration specific to this project: the Makefile shows an explicit bump-benthos target that runs go get -u github.com/redpanda-data/benthos/v4@latest. The project tracks Benthos as an upstream dependency, so some behaviour changes will arrive through that dependency rather than through a connector you chose to use. The CHANGELOG.md at the repository root is where to look before upgrading.

On licensing, the README badges reference an Apache V2 API and a separate Enterprise API, and the repository carries a licenses/ directory and a README-FIPS.md. The license identifier is not stated in the project's own top-level metadata, so treat the boundary between the Apache-licensed API surface and any enterprise components as something to confirm against the repository's licence files before you depend on a specific component. That is a factual check, not legal advice.

Editorial conclusion

Adopt Redpanda Connect when your pipeline is a topology you can express in one YAML file and your delivery requirement is at-least-once: the in-process transaction model with no disk-persisted state makes horizontal scaling and restarts uneventful. Do not adopt it when you need exactly-once semantics across a sink that cannot deduplicate, or when your transformation is a stateful join over long time windows, because that is not what the processor is built around. Before committing, verify two things yourself: that the connector you need is in the catalog at docs.redpanda.com/connect/home, and that your sink tolerates duplicate records, since the README states at-least-once as the default with no caveats.

Frequently asked questions

What is Redpanda Connect used for?

It is a stream processor that moves data between sources and sinks, with hydration, enrichment, transformation and filtering in between. The README highlights declarative pipelines, first-class change-data-capture connectors, and at-least-once delivery by default.

How do I install Redpanda Connect on Linux?

The README downloads an rpk archive with curl from the redpanda releases page and unzips it into a bin directory such as ~/.local/bin. macOS uses brew install redpanda-data/tap/redpanda, and a container image is available at docker.redpanda.com/redpandadata/connect.

Does Redpanda Connect guarantee exactly-once delivery?

No. The README states that it uses an in-process transaction model with no disk-persisted state and guarantees at-least-once delivery when connecting at-least-once sources and sinks. Duplicates are therefore possible and must be handled by the sink.

What are the /ping and /ready endpoints in Redpanda Connect?

/ping is a liveness probe that always returns 200. /ready is a readiness probe that returns 200 once both the input and the output are connected, and 503 otherwise.

How do I build Redpanda Connect from source?

The README clones the repository and runs task build:all, which requires a currently supported Go version. The Makefile in the repository is deprecated and prints a warning pointing to taskfile.dev.

Official sources

  1. Issues
  2. Project website
  3. README
  4. redpanda-data/connect on GitHub
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/redpanda-data-connect.svg)](https://hysenlabs.com/projects/redpanda-data-connect)