Open-source project
ArroyoSystems/arroyo avatar
ArroyoSystems/arroyo

Arroyo: a Rust stream processing engine driven by SQL

Distributed stream processing engine in Rust

5,047 stars381 forksRustApache-2.0

At a glance

What is it?
Arroyo is a distributed stream processing engine written in Rust that runs stateful SQL pipelines over bounded and unbounded sources. It ships as a single binary, but the documentation is thin on operational limits, so the fit depends on your workload.
Who is it for?
Adopt Arroyo if you want SQL-first stream processing with stateful windows and joins, and you are comfortable reading the documentation at doc.arroyo.dev rather than a thick operations manual. Skip it if you need mature, widely deployed operational tooling around the engine, or if your team already has deep Flink or Kafka Streams experience and no reason to move.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 1 day ago.
What is it written in?
Mainly Rust, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 2, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What Arroyo is for, and who should care

Arroyo targets a specific gap: teams that need continuous answers from high-volume event streams but do not want to staff a streaming specialist. The README frames the project as a distributed stream processing engine written in Rust, designed to perform stateful computations on streams of data. Its stated use cases are fraud and security detection, real-time product and business analytics, real-time ingestion into a warehouse or data lake, and real-time ML feature generation. Those four share a shape: events arrive continuously, some state has to be carried across them, and the result needs to land somewhere else quickly.

The intended audience is explicit in the README's comparison section. Under the heading "Designed for non-experts", it says Arroyo cleanly separates the pipeline APIs from its internal implementation, and that you do not need to be a streaming expert to build real-time data pipelines. That is a claim about API surface, not about operational simplicity. A SQL pipeline still has to be deployed, checkpointed and monitored, and the repository does not pretend otherwise. If your team already writes Flink jobs comfortably, the non-expert framing is not aimed at you.

How a SQL pipeline becomes a running dataflow

The workspace layout in Cargo.toml shows the pipeline being split into distinct crates rather than one monolith. crates/arroyo-planner handles planning, crates/arroyo-datastream defines the dataflow representation, crates/arroyo-worker executes work, crates/arroyo-controller coordinates, and crates/arroyo-state plus crates/arroyo-state-protocol deal with checkpointing. A SQL statement is parsed and planned, translated into a dataflow graph, and handed to workers that run it. The README describes the model as the Dataflow model, linking to the O'Reilly article on the world beyond batch, which is the same lineage Flink and Beam draw from.

State is the part worth understanding before adopting. The README lists stateful operations including windows and joins, and separately lists state checkpointing for fault-tolerance and recovery of pipelines. Those two features are coupled: a windowed aggregation or a join holds state that must survive a worker restart, and the checkpoint mechanism is what makes that possible. The crates directory confirms a dedicated state protocol rather than an incidental implementation detail. What the README does not document is the checkpoint interval, the storage backend used for checkpoints, or the recovery time you should expect. Those are questions to answer from doc.arroyo.dev before a production rollout, not from the repository front page.

Installing Arroyo and running a first cluster

Arroyo ships as a single binary, and the README gives three install paths. On macOS, Homebrew:

bash
brew install arroyosystems/tap/arroyo

On macOS or Linux, an install script:

bash
curl -LsSf https://arroyo.dev/install.sh | sh

The third option is downloading a binary for your platform from the releases page. Once the binary is on the machine, the README starts a cluster with a single command:

bash
arroyo cluster

If you would rather not install locally, the README also gives a Docker invocation that publishes the UI port:

bash
docker run -p 5115:5115 \
      ghcr.io/arroyosystems/arroyo:latest

After either of those, the README says to load the Web UI at http://localhost:5115. From there, the getting started guide at doc.arroyo.dev/getting-started and the tutorial at doc.arroyo.dev/tutorial/first-pipeline/ cover creating a pipeline. The README does not reproduce the SQL for a first pipeline, so treat the tutorial as the required next step rather than something you can infer from the install section.

The operational edges the README leaves open

The most concrete limitation is stated by the project itself, in the Cloudflare Pipelines section. Arroyo is available as a managed offering on the Cloudflare Developer Platform, and the README says that at the time of writing, stateless pipelines ingesting into R2 are supported, with stateful pipelines to follow. If stateful processing is the reason you are looking at Arroyo, the managed path does not yet cover it. Self-hosting does, but then you own the cluster.

A second gap is scale claims versus scale evidence. The README lists "Scales up to millions of events per second" as a feature. Nothing in the repository front page describes the hardware that figure assumes, the partitioning model, or how throughput degrades with wide windows or large joins. Stateful operations are exactly where throughput claims get complicated, and the README does not separate them.

The third gap is connector coverage. The README points to doc.arroyo.dev/connectors and names Kafka and Iceberg as examples, which implies the list is longer than two and lives elsewhere. Before designing a pipeline around a specific source or sink, check that page rather than assuming parity with Flink's connector catalogue. A stream engine is only as useful as the systems it can read from and write to.

Finally, there is no documented rollback story on the front page. The README covers installing and starting a cluster, and says nothing about upgrading between releases or downgrading after a bad one. With a dedicated state protocol in the crate list, state compatibility across versions is a real question, and the README does not answer it.

Arroyo against Flink and Kafka Streams

The README names its alternatives directly: Apache Flink, Spark Streaming, and Kafka Streams. The stated differences are serverless operations, high performance SQL as a first-class concern, and separation of pipeline APIs from internals so non-experts can build pipelines. Take those in order, because they are not equally strong.

The SQL-first claim is the substantive one. Flink has a SQL layer, but it sits alongside a DataStream API that many production jobs use, so SQL is one interface among several. In Arroyo, SQL appears to be the interface, with a planner crate and a SQL testing crate in the workspace. Kafka Streams is the clearest contrast: it is a Java library embedded in your application, not a cluster you deploy, and its topology is built through a Java DSL rather than SQL. If your pipeline logic is naturally expressed as application code and you already run Kafka, Kafka Streams removes an entire deployment tier that Arroyo adds.

The serverless claim is the weakest of the three as presented. The README describes pipelines as designed to run in modern cloud environments with scaling, recovery and rescheduling, and the repository contains a k8s/ directory and an arroyo-operator crate, so Kubernetes deployment is a real path. But "designed for" is not the same as "operated at scale by many teams", and the README offers no operational detail to close that gap. Flink's advantage here is not architecture, it is the accumulated body of documentation, managed services and people who have run it in anger.

Licence, releases and what maintenance costs you

Arroyo is dual-licensed under Apache-2.0 and MIT, with LICENSE-APACHE and LICENSE-MIT both present at the repository root and the README badge reading MIT/Apache-2.0. For most adopters that is permissive enough to embed or redistribute without a copyleft obligation. The repository does not include a contributor licence agreement or a separate commercial licence file in its top-level entries, so there is no dual-licensing gate visible on the front page. This is a description of what the repository contains, not legal advice; if you plan to redistribute Arroyo inside a product, read both licence files yourself.

On cadence, the releases listed are v0.15.0 on 2025-12-01, v0.14.1 on 2025-06-23 and v0.14.0 on 2025-03-26. The gap between v0.14.1 and v0.15.0 is roughly five months. The last push to the default branch was on 2026-09-22, so the repository is not dormant, but a version number still below 1.0 after that much development is a signal about API stability. Expect to read release notes before upgrading rather than assuming compatibility.

Upgrade cost is the piece the material cannot settle. The workspace contains a dedicated state protocol crate, which suggests state formats are versioned and therefore may need migration. The README does not document a migration procedure, a compatibility policy, or a supported upgrade path between minor versions. Budget time to test an upgrade on a copy of a real pipeline before touching production, and treat the absence of a documented rollback as a risk you are accepting rather than one the project has mitigated.

Editorial conclusion

Adopt Arroyo if you want SQL-first stream processing with stateful windows and joins, and you are comfortable reading the documentation at doc.arroyo.dev rather than a thick operations manual. Skip it if you need mature, widely deployed operational tooling around the engine, or if your team already has deep Flink or Kafka Streams experience and no reason to move. Before committing, verify which connectors your pipeline actually needs against doc.arroyo.dev/connectors, and check whether the stateful features you rely on are available in the managed Cloudflare Pipelines offering, which the README describes as supporting stateless pipelines into R2 at the time of writing.

Frequently asked questions

What is Arroyo?

Arroyo is a distributed stream processing engine written in Rust that performs stateful computations on streams of data, with SQL as its pipeline interface. The README describes it as letting you ask complex questions of high-volume real-time data with subsecond results.

How do I install Arroyo?

The README gives three paths: Homebrew with brew install arroyosystems/tap/arroyo, a shell script at https://arroyo.dev/install.sh, or a binary downloaded from the releases page. A Docker image at ghcr.io/arroyosystems/arroyo is also documented, publishing port 5115.

Does Arroyo support stateful operations like windows and joins?

Yes. The README lists stateful operations including windows and joins among its features, along with state checkpointing for fault-tolerance and recovery of pipelines. The workspace includes dedicated crates for state and for the state protocol.

Which connectors does Arroyo support?

The README points to doc.arroyo.dev/connectors for the full list and names Kafka and Iceberg as examples. The repository front page does not enumerate the complete set, so check that page before designing a pipeline around a specific source or sink.

Can I run Arroyo without self-hosting?

Yes, as Cloudflare Pipelines on the Cloudflare Developer Platform. The README states that at the time of writing it is in beta and supports stateless pipelines ingesting into R2, with stateful pipelines to follow.

What licence is Arroyo released under?

Arroyo is dual-licensed under Apache-2.0 and MIT, with LICENSE-APACHE and LICENSE-MIT at the repository root. The README badge reads MIT/Apache-2.0.

Official sources

  1. ArroyoSystems/arroyo on GitHub
  2. License: Apache-2.0
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/arroyosystems-arroyo.svg)](https://hysenlabs.com/projects/arroyosystems-arroyo)