# Broadway for Elixir: concurrent data ingestion pipelines built on GenStage

> Broadway is an Elixir library that turns a producer, a pool of processors and optional batchers into a supervised GenStage topology. It fits teams already running Elixir who need long-lived ingestion from SQS, Kafka, PubSub or RabbitMQ, and it is a poor fit for one-off batch jobs.

**elixir-broadway/broadway** — Concurrent and multi-stage data ingestion and data processing with Elixir

- Repository: https://github.com/elixir-broadway/broadway
- Website: https://elixir-broadway.org
- Stars: 2,684 · Forks: 181
- Language: Elixir
- License: Apache-2.0
- Published: 2026-09-28 · Updated: 2026-09-28 · Language: en
- Canonical page: https://hysenlabs.com/projects/elixir-broadway-broadway

## What Broadway is for, and who should reach for it

Broadway targets a specific shape of problem: a stream of events arriving continuously from an external service, each event needing some work done to it, and the results needing to be grouped before they are written somewhere. The README describes it as a way to "build concurrent and multi-stage data ingestion and data processing pipelines with Elixir", consuming from sources it calls producers, naming Amazon SQS, Apache Kafka, Google Cloud PubSub and RabbitMQ among others.

The audience is narrow on purpose. You need to be running Elixir, because Broadway is a library and not a service you install and point at a queue. You need a pipeline that stays up, since the README states that Broadway pipelines are long-lived. And you need the operational parts of ingestion (back-pressure, acknowledgements, failure handling, metrics) to be the library's job rather than yours. A team writing a nightly ETL script that reads a file and writes a table is not the audience. A team consuming a queue that never empties is.

## How the GenStage topology is assembled from configuration

The mechanism is a supervision tree built from a keyword list. When you call Broadway.start_link/2 you pass a producer, a set of processors and an optional set of batchers, and Broadway translates that into a GenStage topology where each stage is a separate process (or pool of processes) supervised together. The README puts it plainly: Broadway "takes the burden of defining concurrent GenStage topologies and provides a simple configuration API that automatically defines concurrent producers, concurrent processing, batch handling, and more".

Data flows in one direction. The producer pulls from the external source and emits events downstream. Processors receive individual messages and run handle_message/3, where you transform the payload and decide which batcher, if any, the message belongs to. Batchers receive messages in groups and run handle_batch/4. Acknowledgement happens at the end of that chain, which is why the README lists "automatic acknowledgements at the end of the pipeline" as a built-in feature rather than something you wire up.

The concurrency numbers are the configuration. In the README's SQS example, processors run with concurrency: 50 and the batcher runs with concurrency: 5, batch_size: 10 and batch_timeout: 1000. Those four numbers decide how much parallelism you get and how long a batch waits before it is flushed. Ordering and partitioning, rate-limiting and custom failure handling are also listed as features, which means the library exposes knobs for the cases where a plain concurrent pool is not enough.

## Installing Broadway and running a first SQS pipeline

Broadway is published on Hex. The README's installation step is a single dependency line in mix.exs. The version constraint given is ~> 1.0.

```elixir
def deps do
  [
    {:broadway, "~> 1.0"}
  ]
end
```

Once that is in place, the README moves straight to a worked example rather than describing further setup steps. That example uses Amazon SQS and assumes you have added broadway_sqs as a dependency and configured your SQS credentials. The module below is the README's example: a producer pointed at a queue URL, 50 concurrent processors, and a batcher named s3 that flushes every 10 messages or every 1000 milliseconds.

```elixir
defmodule MyBroadway do
  use Broadway

  alias Broadway.Message

  def start_link(_opts) do
    Broadway.start_link(__MODULE__,
      name: __MODULE__,
      producer: [
        module: {BroadwaySQS.Producer, queue_url: "https://us-east-2.queue.amazonaws.com/100000000001/my_queue"}
      ],
      processors: [
        default: [concurrency: 50]
      ],
      batchers: [
        s3: [concurrency: 5, batch_size: 10, batch_timeout: 1000]
      ]
    )
  end

  def handle_message(_processor_name, message, _context) do
    message
    |> Message.update_data(&process_data/1)
    |> Message.put_batcher(:s3)
  end

  def handle_batch(:s3, messages, _batch_info, _context) do
    # Send batch of messages to S3
  end

  defp process_data(data) do
    # Do some calculations, generate a JSON representation, process images.
  end
end
```

handle_message/3 is where the per-message work happens; Message.update_data/2 replaces the payload and Message.put_batcher/2 routes the message to the s3 batcher. handle_batch/4 is where the group is written out. The last step is supervision: the README says to add the module as a child of your application supervision tree as {MyBroadway, []}. Without that, the pipeline starts only if something else starts it.

## Where Broadway is the wrong tool

The README's own comparison with Flow draws the boundary. Flow is described as "a more general abstraction than Broadway that focuses on data as a whole, providing features like aggregation, joins, windows, etc." Broadway is described as focusing on events and on operational features. The README then states the split directly: Broadway is recommended for continuous, long-running pipelines, while Flow works with short- and long-lived data processing.

So a job that reads a bounded dataset, joins it against another, and finishes is a Flow job. Reaching for Broadway there means you inherit a long-lived supervision tree, a producer abstraction and acknowledgement semantics you do not need, and you still have no aggregation or windowing primitives to build on.

There is a second boundary that the README does not spell out but the example implies: Broadway is not a queue. It consumes from one. If your source has no official producer, the README points to the docs for "detailed how-tos and supported producers" rather than promising a generic adapter, so the work of writing a producer falls on you. And the acknowledgement model, which fires at the end of the pipeline, means a crash after your side effect but before the acknowledgement is a duplicate-processing problem you have to reason about in your own handler.

## Broadway against Flow, and why the choice is about lifetime

Both libraries are built on GenStage, and the README says so in the first sentence of its comparison. That shared foundation is why the difference is easy to miss. It is not a difference in runtime or in concurrency model. It is a difference in what the library assumes about the data.

Flow treats the data as a whole: you can aggregate, join and window. Broadway treats the data as events: each one arrives, gets processed, gets acknowledged. Because Broadway assumes events, it can offer automatic acknowledgements, custom failure handling, rate-limiting and metrics as built-in features, and the README lists exactly those. Because Flow assumes a dataset, it can offer the aggregation and windowing that Broadway does not list.

The practical test is what happens when the source is empty. A Broadway pipeline waits. A Flow pipeline finishes. If your process is supposed to terminate when the data runs out, Broadway's model is working against you, and the README's recommendation to use Flow for short-lived processing is the honest answer rather than a hedge.

## Maintenance, licence and what upgrading costs

The repository is not archived and the last push was on 2026-09-17, which is recent. The README does not document a release cadence or a versioning policy beyond the ~> 1.0 constraint in the install snippet, and no release notes were available to check, so the upgrade cost cannot be stated from the README.

The licence is Apache-2.0. The README carries the full Apache text and names two copyright holders: Plataformatec from 2019 and Dashbit from 2020. Apache-2.0 is a permissive licence that includes an explicit patent grant and requires preservation of notices; it is not a copyleft licence. That is a description of the text, not legal advice, and the terms that matter for your organisation are in the LICENSE file at the repository root.

One maintenance cost is visible in the repository layout rather than the README: producers live in separate packages, so broadway_sqs, a Kafka producer and a PubSub producer each version independently of Broadway itself. An upgrade therefore has two moving parts, the core library and the producer package, and the compatibility between them is something you check before you bump either.

## Conclusion

Adopt Broadway if you already run Elixir or Erlang and need a long-lived pipeline that consumes from SQS, Kafka, PubSub or RabbitMQ with back-pressure, batching and acknowledgements handled by the topology rather than by your own code. Do not adopt it for short-lived or one-off data processing, and do not adopt it if you are not prepared to run the BEAM in production; the README points that case at Flow instead. Before writing application code, verify three things: that an official producer exists for your source, that your OTP version satisfies the constraint in mix.exs, and that your acknowledgement semantics match what the producer expects, since Broadway acknowledges at the end of the pipeline and a message processed twice is a design question you have to answer yourself.

## FAQ

### What exactly is Broadway?

It is an Elixir library for building concurrent, multi-stage data ingestion and processing pipelines, consuming from sources it calls producers such as Amazon SQS, Apache Kafka, Google Cloud PubSub and RabbitMQ. Pipelines are long-lived and run on the Erlang VM.

### How do I install Broadway in an Elixir project?

Add {:broadway, "~> 1.0"} to the deps list in mix.exs. A producer for your source, such as broadway_sqs, is a separate dependency.

### Does Broadway work with Kafka or RabbitMQ?

The README names Amazon SQS, Apache Kafka, Google Cloud PubSub and RabbitMQ as producers Broadway can consume from, and points to the docs for the list of official producers and their how-tos.

### When should I use Broadway instead of Flow?

The README recommends Broadway for continuous, long-running pipelines and says Flow works with short- and long-lived data processing. Flow also provides aggregation, joins and windows, which Broadway does not list.

## Sources

- [elixir-broadway/broadway on GitHub](https://github.com/elixir-broadway/broadway)
- [Issues](https://github.com/elixir-broadway/broadway/issues)
- [License: Apache-2.0](https://github.com/elixir-broadway/broadway/blob/main/LICENSE)
- [Project website](https://elixir-broadway.org)
- [README](https://github.com/elixir-broadway/broadway/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/elixir-broadway-broadway
