Library / SDK
apache/beam avatar
apache/beam

Apache Beam: a portable pipeline model that outlives its runner

Apache Beam is a unified programming model for Batch and Streaming data processing.

8,672 stars4,662 forksJavaApache-2.0

At a glance

What is it?
Apache Beam separates pipeline definition from execution, so the same Java, Python or Go code can run on DirectRunner, Flink, Spark or Dataflow. The trade-off is a portability layer that adds moving parts between your logic and the cluster.
Who is it for?
Adopt Apache Beam when the same pipeline logic has to survive a change of execution engine, or when you need one programming model across batch and streaming, and you are willing to accept the portability layer between your code and the cluster. Do not adopt it as a scheduler: Airflow orchestrates jobs, Beam defines and executes one, so a team looking for dependency management, retries across systems and backfill calendars is looking at the wrong tool.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 4 days ago.
What is it written in?
Mainly Java, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 25, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The problem Beam solves: writing the pipeline once and choosing the engine later

Most data processing code is welded to its execution engine. A Spark job is written against Spark's APIs, a Flink job against Flink's, and moving between them means a rewrite plus a new set of operational assumptions. Apache Beam attacks that coupling directly. The README describes it as "a unified model for defining both batch and streaming data-parallel processing pipelines, as well as a set of language-specific SDKs for constructing pipelines and Runners for executing them on distributed processing backends".

The audience is split three ways in the README's own framing. End users write pipelines with an existing SDK and run them on an existing runner. SDK writers build a Beam SDK for a specific language community and want to be shielded from runner internals. Runner writers have an execution environment and want to support Beam programs without knowing about every SDK. That second and third group matters for adopters: Beam is not only a library you consume, it is an integration surface that other engines implement against.

The practical consequence is that the choice of engine becomes a deployment decision rather than a design decision. A team that starts on DirectRunner locally and later moves to Flink or Dataflow keeps the pipeline definition. What changes is the runner configuration and, in practice, the parts of the model that runner supports well.

PCollection, PTransform, Pipeline: the shape of a Beam program

The model rests on four concepts the README lists explicitly. A PCollection is a collection of data, bounded or unbounded in size. A PTransform is a computation that turns input PCollections into output PCollections. A Pipeline manages a directed acyclic graph of PTransforms and PCollections ready for execution. A PipelineRunner specifies where and how that graph executes.

That is a functional dataflow graph, and the bounded/unbounded distinction is the load-bearing part. Batch and streaming are not two APIs in Beam; they are the same graph applied to bounded or unbounded input. The model itself descends from Google's MapReduce, FlumeJava and Millwheel work, published originally as the Dataflow Model.

The separation between Pipeline and PipelineRunner is what makes portability possible, and it is also the source of the friction described later. Your transforms describe intent. The runner decides how to schedule, shuffle, checkpoint and checkpoint-recover. When a transform has no efficient implementation on a given runner, the graph still builds; the cost shows up at execution time.

Installing Apache Beam and running a first pipeline locally

The README points new users at language quickstarts rather than giving install commands inline: Java, Python and Go each have a quickstart page on beam.apache.org, and the repository ships a minimal WordCount example plus a Python SDK readme at sdks/python/README.md. The README does not spell out a pip command, so the install path it documents is the quickstart page for your language, and the Python SDK readme in the repository.

bash
pip install apache-beam

That line is the conventional Python install for the SDK, and the README's status section links the PyPI badge for apache-beam. After installation the SDK exposes the beam module, and the DirectRunner is the default local execution path, needing no cluster. The repository's examples directory holds the WordCount variants for each language; the README calls it the "Minimal WordCount example (available in this repository)".

A pipeline is built from the four model concepts. The README's quick start lists them in this order: PCollection, PTransform, Pipeline. Reading a text file, splitting lines into words, counting occurrences and writing the result is the shape the WordCount example takes in each SDK.

For Java, the analogous path is the Java Quickstart and the org.apache.beam artifacts on Maven Central, which the README's status section links. The repository builds with Gradle (build.gradle.kts, settings.gradle.kts, gradlew) and includes start-build-env.sh and local-env-setup.sh for setting up a development environment. Building Beam itself is a different exercise from using it, and the README defers that to CONTRIBUTING.md.

Runners are not interchangeable in practice

The README lists seven PipelineRunners: DirectRunner and PrismRunner locally, plus DataflowRunner, FlinkRunner, SparkRunner, JetRunner and Twister2Runner on clusters. The list reads like a menu of equivalent options. It is not. Each runner is a separate implementation of the same model, and the README does not claim feature parity across them.

This is the real limitation to weigh. A pipeline that uses a transform with no implementation on your target runner will fail or fall back. The Beam project tracks this through its portability work, which is why PrismRunner exists as a local runner that exercises Beam Portability rather than running the classic local path. If your pipeline depends on a specific I/O connector, verify that connector's support for your runner and SDK version before you commit to the architecture.

The second limitation is operational. Beam gives you a pipeline abstraction, not a cluster. You still need Flink or Spark or Dataflow underneath, with their own tuning, versioning and failure modes. The abstraction hides the engine's API, not its behaviour. When a job is slow or a checkpoint fails, the diagnosis happens at the runner layer, and the Beam graph is only part of the picture.

The third case where Beam is the wrong tool is small data. If a job fits comfortably in a single process, the model, the SDK and the runner add ceremony with no payoff. The DirectRunner will run it, but a plain script would too.

Beam compared with Spark, Flink and Airflow

Apache Beam versus Apache Spark is the comparison people search for most, and the difference is architectural. Spark is an execution engine with its own programming APIs; Beam is a model and SDK layer that targets engines including Spark. If you write Spark jobs, you are committing to Spark. If you write Beam pipelines and run them on SparkRunner, you keep the option to move to FlinkRunner or DataflowRunner, at the cost of a translation layer.

Against Apache Flink the same logic applies in reverse: Flink is one of Beam's backends, and FlinkRunner is the Beam runner that submits to a Flink cluster. Choosing Flink directly buys you its native APIs and its own state and checkpointing model without an intermediary. Choosing Beam buys you the ability to leave.

Airflow is a different category entirely, and the confusion is worth clearing up. Airflow schedules and orchestrates work across systems; Beam defines and executes a data processing graph. A Beam pipeline can be one task in an Airflow DAG, and the README's model has nothing to say about dependency management, calendars or cross-system retries. Comparing them as alternatives misses that they occupy different layers.

The same applies to Kafka, which is a log and messaging system rather than a processing model. A Beam pipeline can read from and write to Kafka through I/O connectors; it does not replace the broker.

Licence, releases and what upgrading costs

Apache Beam is licensed under Apache-2.0. The repository carries LICENSE, NOTICE, LICENSE.python and LICENCE.cloudpickle as separate files, which reflects that the codebase bundles components with their own licensing terms. If you redistribute Beam or bundle it into a product, read those files rather than assuming a single licence covers everything. That is an observation about the repository layout, not legal advice.

Release cadence is visible from the tags: v2.74.0, v2.75.0 and v2.76.0-RC4 appear across May, July and August 2026, with the last push to master on 2026-08-11. The project is not archived. Frequent minor releases are convenient for staying current and expensive for staying pinned, because the SDK and the runners version together. A pipeline built against one Beam version and run on a cluster whose runner integration targets another is a configuration problem you will have to solve.

Upgrade cost concentrates in two places: the SDK dependency in your build, and the runner-side integration on your cluster. The CHANGES.md file at the repository root is where release notes accumulate, and it is the first place to look before bumping the version in a production pipeline.

Editorial conclusion

Adopt Apache Beam when the same pipeline logic has to survive a change of execution engine, or when you need one programming model across batch and streaming, and you are willing to accept the portability layer between your code and the cluster. Do not adopt it as a scheduler: Airflow orchestrates jobs, Beam defines and executes one, so a team looking for dependency management, retries across systems and backfill calendars is looking at the wrong tool. Before committing, verify that the transforms you depend on are implemented for your chosen runner and SDK version, because the model is portable while individual runner and SDK combinations are not equally complete. Check the licence file that ships with the artifact you install, since the repository carries LICENSE, LICENSE.python and LICENCE.cloudpickle as separate files.

Frequently asked questions

What is Apache Beam used for?

It is used to define batch and streaming data-parallel processing pipelines that can run on multiple distributed backends. The README describes it as a unified model plus language-specific SDKs for building pipelines and runners for executing them.

What is Apache Beam vs Spark?

Spark is an execution engine with its own programming APIs, while Beam is a model and SDK layer that can target Spark through SparkRunner. Writing Beam pipelines keeps the option of moving to FlinkRunner or DataflowRunner; writing Spark jobs commits you to Spark.

What is the latest version of Apache Beam?

The most recent release tag in the repository is v2.76.0-RC4, dated 2026-08-11, following v2.75.0 in July 2026 and v2.74.0 in May 2026.

What are the key differences between Airflow and Apache Beam?

Airflow schedules and orchestrates work across systems, while Beam defines and executes a data processing graph. A Beam pipeline can sit inside an Airflow DAG as one task; Beam does not provide scheduling or cross-system dependency management.

How to install Apache Beam in Python?

The README's status badges point at the PyPI package apache-beam, and the Python SDK readme lives at sdks/python/README.md in the repository. The README itself defers install steps to the Python Quickstart page on beam.apache.org.

How to run Apache Beam locally?

The DirectRunner runs the pipeline on your local machine, and PrismRunner also runs locally but uses Beam Portability. Both are listed in the README's runner section, and neither needs a cluster.

Official sources

  1. Official documentation
  2. Official README
  3. Project repository
  4. Release notes
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/apache-beam.svg)](https://hysenlabs.com/projects/apache-beam)