Self-hosted service
numaproj/numaflow avatar
numaproj/numaflow

Numaflow: a Kubernetes-native platform for parallel data and streaming jobs

Kubernetes-native platform to run massively parallel data/streaming jobs

2,830 stars179 forksRustApache-2.0

At a glance

What is it?
Numaflow decouples event sources and sinks from processing logic and runs each pipeline vertex as its own autoscaling unit. This review covers the pipeline model, install steps, the exactly-once claim, and where the design stops being the right tool.
Who is it for?
Adopt Numaflow if your event processing already lives on Kubernetes and you want each step to scale on its own without writing consumer boilerplate. Skip it if you need a single-node runner, ordered delivery, or a managed service with a support contract, since the README states that order is not preserved.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 7 days ago.
What is it written in?
Mainly Rust, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 24, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The problem Numaflow targets: wiring sources and sinks to your code

Most streaming jobs on Kubernetes end up as a pod that does three things at once: it talks to a broker, it runs your transformation, and it writes results somewhere. Scaling that pod means scaling the broker client and the transformation together, even when only one of them is the bottleneck. Numaflow splits those responsibilities. Sources and sinks ship with the platform, and a vertex holds only your processing logic. The README describes the goal as decoupling event sources and sinks from processing logic so each component can independently auto-scale based on demand.

The intended audience is a team that already runs Kubernetes and does not want to write consumer loops, offset handling or retry plumbing. The README lists four use cases: event-driven applications such as inventory updates and customer notifications, real-time analytics on social media or observability data, inference on streaming data such as anomaly detection, and workflows that run in a streaming manner. The repository is Apache-2.0, and the code is a mix of Go (the API types and controllers) and Rust (the runtime that carries data between vertices).

How a Numaflow pipeline actually moves data

A pipeline is a directed graph of vertices. Sources read from an external system, map and reduce vertices transform, and sinks write out. Between vertices the data does not travel point to point; it goes through an inter-step buffer service, which the examples refer to as ISB. The repository ships examples/0-isbsvc-jetstream.yaml, so JetStream is one supported backing store for that buffer, and the go.mod file lists both github.com/nats-io/nats-server/v2 and github.com/nats-io/nats.go among the dependencies.

That indirection is what makes the scaling story work. A map vertex can scale down to zero while the source keeps reading, and the buffer absorbs the difference. It is also what makes the semantics story possible: the README claims exactly-once semantics so that no input element is duplicated or lost even as pods are rescheduled or restarted, and separately states that at-least-once is the minimum guarantee. Those two statements sit in different sections, and the exactly-once claim is scoped to unbounded and near real-time sources.

Watermarks are part of the model. The examples directory contains examples/1-simple-pipeline-wm-disabled.yaml, which implies watermark progress is on by default and can be turned off. Windowed operations are covered too: session windows in examples/12-simple-session-pipeline.yaml, accumulator windows in examples/13-accumulator-window.yaml, and streaming reduce in examples/13-streaming-reduce-pipeline.yaml. A separate set of examples covers monovertex, a single-vertex form of the same runtime, which is useful when the graph is trivial and a full pipeline is overhead.

Installing Numaflow and running a first pipeline

The README points to docs/quick-start.md and the examples directory rather than reproducing install commands in the README body, so the exact apply sequence should be read from those files. What the repository does show is that the platform is installed into a cluster as Kubernetes resources, and that pipelines are declared as YAML manifests. The example names follow a numbered progression, so a sensible first read is the simple pipeline manifest.

A pipeline manifest declares the vertices, the edges between them, and the source and sink configuration. The repository's examples/1-simple-pipeline.yaml is the minimal shape. The README's demo and its QUICK_START link are the entry points it gives for a first run, and the examples directory holds the manifests to apply. Read docs/quick-start.md for the exact commands rather than guessing at a sequence.

For a custom step, the examples include UDF-based pipelines such as examples/4-udsink-pipeline.yaml, and the README states that each step can be written in any programming language. The Go client types live under pkg/apis and are documented on GoDoc, which is the reference to use when writing a user-defined function in Go rather than copying a manifest.

Where Numaflow's guarantees stop

The README is explicit that preserving order is not required. If your downstream consumer assumes monotonic ordering per key, Numaflow is the wrong layer, and no amount of tuning the buffer will fix it. Applications that need ordered replay should keep that responsibility elsewhere.

The exactly-once claim is also narrower than the feature list suggests. The README states exactly-once for unbounded and near real-time data sources, and at-least-once as the minimum everywhere else. A bounded or batch-like source falls outside the stronger guarantee. Teams that read the key features list and stop there will overestimate what they get.

There is a second cost that the README does not address: the ISB is a stateful component in the middle of every edge. Operating it is part of operating the pipeline. The roadmap entries for 1.9 mention buffer ownership changes and moving monitor container functionality into numa, and 1.10 lists a Lowest-Watermark First ISB reader. Those are future changes, not current behaviour, and they suggest the buffer path is still being reworked. Anyone planning a long-lived deployment should track those items before depending on the current internals.

Numaflow compared with Kafka Streams and Flink

The closest comparison is a stream processor that runs its own cluster, such as Apache Flink. Flink owns both the runtime and the scheduling, and you deploy it as a job on its own cluster or on Kubernetes. Numaflow does not own the scheduler; it hands that to Kubernetes and expresses the pipeline as custom resources. The difference shows up in operations: with Numaflow you debug with kubectl and read pipeline status from the Kubernetes API, while with Flink you learn a separate job manager and task manager model.

Against Kafka Streams the split is different again. Kafka Streams is a library you embed in your application, so partitioning and state stay inside your process and your deployment. Numaflow is a platform you deploy first, then attach steps to. That is more moving parts up front and less code per step. It is a reasonable trade only if you expect several pipelines, because the fixed cost of running the control plane and the buffer is paid once.

Numaflow's own history is worth noting here: the README says it was created by the Intuit Argo team to address community needs for continuous event processing. That lineage explains the Kubernetes-first design, and it also explains why the project assumes cluster access is available and cheap.

Maintenance, release cadence and licence

The repository is not archived, and the last push was on 2026-09-24. Releases have been frequent: v1.8.4 on 2026-09-10, v1.8.3 on 2026-08-07, and v1.8.2 on 2026-07-24. That is roughly a monthly patch cadence across the 1.8 line, which means upgrade work is recurring rather than occasional. The roadmap names 1.9 and 1.10 for items such as per-message Nack support with redelivery, avoiding numa restarts on UDF crashes, and monovertex streaming. Some of those change internal components, so a 1.8 to 1.9 move is not a pure patch upgrade.

The licence is Apache-2.0, which permits commercial use and modification and includes an explicit patent grant. It does not provide support, warranty or indemnity, and it requires that notices and the licence text be preserved in redistributions. That is a description of the terms, not legal advice; get counsel to review redistribution plans.

The build path is non-trivial if you compile from source. The Dockerfile describes a two-stage build: a base stage that copies a host-built Go binary, and a Rust builder stage on rust:1.98.1-trixie that installs protobuf-compiler, cmake and clang. The Makefile exposes CARGO_PROFILE, defaulting to release, with an image-dev profile for faster local rebuilds, and RELEASE_BASE_IMAGE is set to scratch while DEV_BASE_IMAGE is debian:trixie-slim. Building your own images means keeping a Go toolchain and a Rust toolchain in step with the pinned versions in the repository.

Editorial conclusion

Adopt Numaflow if your event processing already lives on Kubernetes and you want each step to scale on its own without writing consumer boilerplate. Skip it if you need a single-node runner, ordered delivery, or a managed service with a support contract, since the README states that order is not preserved. Before committing, verify that the ISB service you intend to use is one of the ones the examples cover, and check the roadmap entries for 1.9 and 1.10 against the behaviour you depend on.

Frequently asked questions

What programming languages can I write a Numaflow step in?

The README states that each step of the pipeline can be written in any programming language, and that this is intended to let you pick the best language per step. The repository provides Go API types under pkg/apis, documented on GoDoc, for steps written in Go.

Does Numaflow guarantee exactly-once processing?

The README says exactly-once semantics are provided for unbounded and near real-time data sources, and that at-least-once is the minimum guarantee otherwise. It also states that preserving order is not required.

What is the ISB service in Numaflow?

The examples include examples/0-isbsvc-jetstream.yaml, which indicates the inter-step buffer is a separate deployed service that sits between pipeline vertices. The roadmap for 1.10 lists a Lowest-Watermark First ISB reader, so the buffer implementation is still changing.

What is a monovertex in Numaflow?

The examples directory contains several monovertex manifests, including examples/21-simple-mono-vertex.yaml and examples/23-mono-vertex-bypass.yaml, alongside a builtin variant. It is a single-vertex form of the runtime for cases where a full multi-vertex pipeline is unnecessary.

Official sources

  1. License: Apache-2.0
  2. numaproj/numaflow on GitHub
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/numaproj-numaflow.svg)](https://hysenlabs.com/projects/numaproj-numaflow)