Open-source project
spotify/scio avatar
spotify/scio

spotify/scio: a Scala API for Apache Beam and Dataflow

A Scala API for Apache Beam and Google Cloud Dataflow.

2,630 stars535 forksScalaApache-2.0

At a glance

What is it?
Scio wraps Apache Beam in a Scala API that reads like Spark or Scalding, with typed BigQuery support and a REPL. It fits teams already on Scala and Google Cloud, and is a poor fit for anyone who wants to stay in Java or Python.
Who is it for?
Adopt scio if your team already writes Scala and runs pipelines on Google Cloud Dataflow, and you want Beam's batch and streaming model behind an API that resembles Spark. Do not adopt it if your organization is standardized on Java or Python Beam, since you would be maintaining a Scala wrapper around a runtime you barely touch.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 6 days ago.
What is it written in?
Mainly Scala, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 24, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The gap scio fills between Beam and Scala teams

Apache Beam is a portable programming model for batch and streaming pipelines, and Google Cloud Dataflow is a managed runner for it. Beam's Java SDK is the reference implementation. If your codebase is Scala, using that SDK directly means writing Java-flavored builders and anonymous classes inside Scala files, and giving up the collection-style transformations Scala developers expect from Spark or Scalding.

Scio is the bridge. The README describes it as a Scala API for Apache Beam and Google Cloud Dataflow, inspired by Apache Spark and Scalding. The audience is narrow and specific: Scala teams running data pipelines on Google Cloud, or teams that already have Scala libraries (Algebird, Breeze) they want to reuse inside a pipeline. The feature list names unified batch and streaming, integration with Cloud Storage, BigQuery, Pub/Sub, Datastore and Bigtable, a Scio REPL, and type safe BigQuery. That last item is the one most likely to decide the question, because BigQuery schema handling is where Java Beam pipelines tend to accumulate stringly typed code.

How a scio pipeline is put together

The repository is a multi-module sbt build, and the module layout is the clearest description of the architecture. scio-core is the core library. Everything else is an add-on: scio-avro, scio-cassandra*, scio-elasticsearch*, scio-google-cloud-platform, scio-grpc, scio-jdbc, scio-managed, scio-neo4j, scio-parquet, scio-redis, scio-smb, scio-snowflake, scio-tensorflow, plus scio-extra for collections and Breeze utilities, and scio-repl for an extended Scala REPL. Tests live in scio-test, split into scio-test-core, scio-test-google-cloud-platform and scio-test-parquet.

That split matters in practice. A pipeline that reads Parquet and writes BigQuery pulls in scio-parquet and scio-google-cloud-platform on top of scio-core, and each add-on brings its own transitive dependencies. The README also notes that scio-extra is provided with best effort support, which is a meaningful qualifier: it is not held to the same bar as the core artifacts.

The data flow itself is Beam's. A pipeline is constructed, transforms are applied to collections, and the whole thing is submitted to a runner. The README mentions pipeline orchestration with Scala Futures, which is how multi-stage jobs (one pipeline producing input for the next) are expressed without leaving Scala. scio-managed is the add-on for Beam's managed transforms, and the README notes it includes Iceberg.

Installing scio and running the word count example

The README's quick start assumes a JDK of version 11 or higher (it points at adoptium or corretto) and sbt. There is no installer and no global CLI. You create a project from the giter8 template, build it, and run the generated job.

Create the project from the template:

bash
sbt new spotify/scio.g8

The README states the default repository name is scio-job. Change into it and build a staged distribution:

bash
cd scio-job
sbt stage

The template ships a word count example. The README runs it with an output argument:

bash
target/universal/stage/bin/scio-job --output=wc

After the job finishes, the README lists the output directory and reads one part file:

bash
ls -l wc
cat wc/part-00000-of-00004.txt

What you should see is a wc directory containing part files (the README's example names four parts) whose contents are word counts. That is the entire local loop: no cluster, no cloud credentials required for the example itself. Moving to Dataflow means adding the runner and Google Cloud configuration, which the README does not cover in the quick start; it defers to the Getting Started page on the documentation site.

Where scio is the wrong tool

The first limitation is the language boundary. Scio is Scala only. A Java or Python team gains nothing from it, and a mixed team gains a second way of writing the same pipelines. The README's own documentation path makes the cost explicit: it recommends the Beam Programming Guide first for anyone new to Beam, which means scio does not remove the need to understand Beam, it only changes the syntax you write it in.

The second is the artifact surface. The add-on list is long, and each entry is a separate dependency with its own version alignment against Beam. scio-extra is labeled best effort support, so utilities there can lag or change without the guarantees the core carries. The README does not document a compatibility matrix between Scio versions and Beam versions, so that alignment has to be checked per release.

The third is runner choice. The README's headline claim of a fully managed service is footnoted as provided by Google Cloud Dataflow. Scio is a Beam API, so other runners exist in principle, but the documentation and integrations are organized around Google Cloud products. If your target is not Google Cloud, the integration list (BigQuery, Pub/Sub, Datastore, Bigtable, Spanner) is largely dead weight.

Scio against plain Beam Java and against Spark

The honest alternative is Beam's Java SDK. It is the same runtime, the same runners, the same connectors, and the same documentation the Scio README points you to. The difference is entirely in the authoring layer: Java Beam gives you the PTransform and DoFn vocabulary directly, with no Scala compilation step and no wrapper to keep in sync with Beam releases. If your team is not already Scala, Java Beam is the shorter path, and it is the version every Beam tutorial and Stack Overflow answer is written against.

The other alternative is Spark, which the README names as an inspiration rather than a competitor. Spark gives you a Scala API too, but its own execution engine, not Beam's. The practical difference is portability and runner choice: a Spark job runs on Spark, while a Scio job compiles down to a Beam pipeline that can be handed to a runner. The README links a comparison page titled Scio, Scalding and Spark, which is the right place to check the API differences before assuming the migration is mechanical. For teams already on Spark and not on Google Cloud, moving to Scio means adopting Beam as well, which is two changes, not one.

Maintenance, releases and the Apache-2.0 licence

The repository is not archived, and the last push was on 2026-09-23, five days before this writing. Releases are frequent and small: v0.15.10 on 2026-09-18, v0.15.9 on 2026-07-01, v0.15.8 on 2026-06-29. The version numbering stays inside 0.15.x, so upgrades within that line are patch-level, but the project has not declared a 1.0 and the README gives no compatibility promise across minor versions.

Upgrade cost is dominated by the Beam dependency, not by scio itself. Every scio release tracks a Beam release, and the add-on modules each have to move with it. A pipeline using scio-core alone is cheap to upgrade; one using scio-managed, scio-tensorflow, scio-snowflake and scio-parquet is a wider surface. The repository includes a .scala-steward.conf and a Scala Steward badge, so dependency bumps are automated upstream, but that does not remove the need to re-run your own pipeline tests after a bump.

Licensing is Apache-2.0, stated in the README and in the LICENSE file, with the copyright line reading Copyright 2024 Spotify AB. Apache-2.0 permits commercial use and modification and includes a patent grant. The add-ons pull in third-party libraries under their own licences, and the README does not enumerate those, so dependency licence review is on you. Nothing here is legal advice; check the LICENSE and NOTICE files and your own dependency tree.

Editorial conclusion

Adopt scio if your team already writes Scala and runs pipelines on Google Cloud Dataflow, and you want Beam's batch and streaming model behind an API that resembles Spark. Do not adopt it if your organization is standardized on Java or Python Beam, since you would be maintaining a Scala wrapper around a runtime you barely touch. Before committing, verify which artifacts your pipeline actually needs (scio-google-cloud-platform and scio-avro are separate from scio-core), confirm that scio-extra's best effort support is acceptable for anything you depend on, and read the Scio, Scalding and Spark comparison page to check whether the API shift matches your team's expectations.

Frequently asked questions

What is spotify/scio?

It is a Scala API for Apache Beam and Google Cloud Dataflow, described in the README as inspired by Apache Spark and Scalding. It ships as a set of artifacts, with scio-core as the core library and add-ons for Avro, BigQuery, Pub/Sub, Parquet, JDBC and others.

How do I install spotify/scio and run a first job?

Install a JDK of version 11 or higher and sbt, then create a project with sbt new spotify/scio.g8, build it with sbt stage, and run the included word count example with target/universal/stage/bin/scio-job --output=wc. The README shows listing the wc directory and reading a part file to see the result.

Is spotify/scio tied to Google Cloud Dataflow?

The README lists a fully managed service as a feature, footnoted as provided by Google Cloud Dataflow, and the integrations it names are Google Cloud products: Cloud Storage, BigQuery, Pub/Sub, Datastore, Bigtable and Spanner. Scio is a Beam API, so the runtime underneath is Beam's, but the documented integrations are Google Cloud oriented.

Which spotify/scio artifact should I add to my project?

Start from scio-core, then add only the add-ons you need, such as scio-avro, scio-google-cloud-platform or scio-parquet. For tests, scio-test is added as a test dependency, and the README notes scio-extra is provided with best effort support.

What licence does spotify/scio use?

The README states it is licensed under the Apache License, Version 2.0, with the copyright line Copyright 2024 Spotify AB. The README does not list the licences of the third-party libraries the add-on modules pull in.

Official sources

  1. License: Apache-2.0
  2. Project website
  3. README
  4. Releases
  5. spotify/scio on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/spotify-scio.svg)](https://hysenlabs.com/projects/spotify-scio)