Framework
apache/gobblin avatar
apache/gobblin

apache/gobblin: a lake maintenance framework whose build fetches Gradle over an unverified connection

A distributed data integration framework that simplifies common aspects of big data integration such as data ingestion, replication, organization and lifecycle management for both streaming and batch data ecosystems.

2,270 stars750 forksJavaApache-2.0

At a glance

What is it?
A distributed data integration framework whose differentiator is not ingestion but compaction, deduplication, retention and deletion, spread across forty modules and four deployment targets. The build instructions tell you to download the Gradle wrapper jar with certificate checking switched off.
Who is it for?
Adopt Gobblin if the bottleneck in your lake is maintenance rather than movement, because compaction, partitioning, deduplication, retention and fine-grain deletion are the operations it names first, and atomic data publishing plus state management for incremental processing are the parts you would otherwise build yourself.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 6 days ago.
What is it written in?
Mainly Java, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 28, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The build instructions download a jar with verification switched off

Read the build section before anything else, because it contains the sharpest edge in this repository. Gobblin is distributed as a source archive that does not include the Gradle wrapper jar, and the README explains how to get it. One documented route is a wget with the no-check-certificate flag. The other is this:

bash
curl --insecure -L https://github.com/apache/gobblin/raw/${GOBBLIN_VERSION}/gradle/wrapper/gradle-wrapper.jar > gradle/wrapper/gradle-wrapper.jar

Both disable TLS certificate verification for the retrieval of an executable artefact, and the URL contains a variable the reader is told to fill in with the version they downloaded. The manual alternative is also given, as a link to the same jar in the repository, and that is the one to use. Not because the automated route is hostile, but because the jar it fetches is a build tool that will then compile and run your dependency resolution, and disabling verification for that step means you are trusting the network path rather than a checksum. Any team with a supply-chain policy will trip over this line, and it is worth knowing the jar is absent from the archive for a structural reason: foundation projects generally avoid committing binary artefacts into a repository, since every file needs an acceptable licence header and a jar has none. So the exclusion is correct and the remedy in the README is careless. Install a system Gradle, or fetch the wrapper jar from a fixed tag and verify its hash. The rest of the build is more conventional, with a rat target for the licence audit and a build target where the checks can be switched off one at a time.

The rat audit, the four skippable checks, and a stray test file

The build section is short and every line is load-bearing. The licence audit is run with a single target:

bash
./gradlew rat

and the report lands under build/rat. The distribution build has two documented paths. The fast one skips four separate checks by name:

bash
./gradlew build -x findbugsMain -x test -x rat -x checkstyleMain

The slow one is a plain build, and the README notes it requires Maven, which is an odd requirement in a Gradle project and is a leftover of the project's history. That one line is a useful inventory: the project has FindBugs, Checkstyle, the release audit and a test suite, wired as four targets you can exclude individually. A build badge and a codecov badge sit in the header, and there is a .codecov_bash script at the top level, so coverage is collected as well as enforced. Then there is the oddity. A file called FlowTriggerHandlerTest.java sits at the repository root, next to README.md and the build files, rather than inside one of the module directories. It is a single escaped test file, which is the kind of artefact that accumulates over years and is harmless, and it is a useful reminder that a 2016-vintage Java project carries the sediment of its own history.

Forty modules, and the directory list is the roadmap

The top-level listing has around forty gobblin-prefixed modules, and reading them tells you what the project has kept current and what it carries. The current-generation entries are gobblin-iceberg, for the table format, and gobblin-temporal, for a durable workflow engine, and both of those are integrations nobody would have added to a project whose release line stopped in 2017 unless work is genuinely continuing. The deployment entries are gobblin-yarn, gobblin-kubernetes, gobblin-cluster, gobblin-service and gobblin-docker, so the same code runs on YARN, on Kubernetes, as a cluster, as a long-running service and in a container. The source and sink entries are gobblin-aws, gobblin-salesforce and gobblin-metastore, so cloud object storage, a SaaS API and a catalogue are first-class. The operational entries are the interesting ones: gobblin-compaction, gobblin-completeness, gobblin-data-management, gobblin-config-management and gobblin-binary-management, which is where the maintenance operations live. And then there is gobblin-oozie, which is a module for Oozie, the Hadoop workflow scheduler that was superseded long ago. Its presence is the clearest single indicator of the project's age in the whole listing. Underneath all of it sit the build mechanics, with buildSrc, a defaultEnvironment.gradle, a gobblin-flavored-build.gradle and a gradle.properties, which is a variant-based build: the same source is assembled into different distributions for different environments, which is how a project this size stays buildable at all.

The paragraph listing what Gobblin is not is the most useful text in the README

Most READMEs tell you what a project is. This one also tells you what it refuses to be, and that paragraph is worth more than the capability list. Apache Gobblin is not a general-purpose data transformation engine like Spark or Flink, and it can delegate complex processing to Spark or Hive instead. It is not a data storage system like Kafka or HDFS, and it integrates with those as sources or sinks. And it is not a general-purpose workflow execution system like Airflow, Azkaban, Dagster or Luigi. Three exclusions, and each one removes an expectation a newcomer would reasonably arrive with. The delegation point is the important one, because it tells you the intended division of labour: Gobblin moves and maintains data, and if a job needs real computation, that job belongs somewhere else. The orchestration exclusion is reinforced by the control plane described in the highlights, which is described as Gobblin-as-a-service supporting programmatic triggering and orchestration of data plane operations. The phrasing is careful. It orchestrates its own data plane and exposes that to be triggered programmatically, which is not the same as owning your workflow definitions, and the deliberate omission of Airflow and Dagster from a system that schedules jobs is a statement about scope rather than a gap. If you arrive expecting a scheduler, you will be disappointed; if you arrive expecting a managed execution layer that your existing scheduler calls, that is what this is.

Compaction and deletion, not ingestion, are the reason to pick it

The capabilities are listed in four groups and the first one is the least interesting. Ingestion and export from sources and sinks into and out of the data lake is table stakes, and the README frames it as ELT with a small t, meaning transformations happen inline on ingest rather than as a separate stage. The other three groups are where the differentiation is. Data organisation within the lake covers compaction, partitioning and deduplication. Lifecycle management covers data retention. And compliance management covers fine-grain data deletions, with GDPR deletion named explicitly against HDFS and ADLS. Those are the operations that ruin a lake over time. Small files accumulate until a query plan is unusable, duplicates appear because an at-least-once source retried, stale partitions quietly consume budget, and a deletion request becomes a hand-written find-and-delete across every table that ever touched the subject's data. Each of those is unglamorous and each is a multi-week engineering project if you build it yourself. The supporting features named in the highlights reinforce the same point: atomic data publishing, so a reader never sees a partially written partition, and state management for incremental processing, so a retry does not reprocess yesterday. If your pipeline is already moving data and the pain is the fourth week of operations, this framework is aimed at you.

Stream and batch, on four deployment targets, with docs for a laptop

Gobblin supports both stream and batch execution modes, which is the feature that decides whether it fits an architecture with both. Most ingestion tools are one or the other, and a framework that does both means a realtime path and a backfill path share a configuration model, a checkpointing model and a failure model instead of living in two systems. The deployment targets follow from the module list and they are a real choice rather than a single recommendation. YARN is the inherited Hadoop path. Kubernetes is the module for a container deployment. A cluster, a standalone service and Docker cover the small and the local. The documentation makes the entry cost unusually low by including instructions for running Gobblin on Docker from a laptop, alongside a getting started guide, an architecture page, a sample project in a gobblin-example directory, and a list of companies known to use it. The common production patterns listed are the ones to read for fit: Kafka into a data lake on HDFS, S3 or ADLS; bulk loading from the lake into a serving store, with Couchbase named; synchronisation between federated lake locations in both directions; pulling from external vendor APIs such as Salesforce and Dynamics; and enforcing retention and deletion policies on HDFS and ADLS. Read that list as the compatibility statement, because it is far more informative than a features page.

Newest release tag from 2017, commits from 2026, and how to actually depend on it

Here is the version picture, and it needs stating plainly. The three most recent releases are 0.9.0 from 2016-12-19, 0.10.0 from 2017-05-05 and 0.11.0 from 2017-07-20. The last push to the repository was on 2026-09-24, and the repository is not archived. So the tag history visible here stops nine years ago while the code has moved on, and a CHANGELOG.md exists in the tree to record what happened in between. There is a Maven Central badge for the group org.apache.gobblin, which is the practical answer to how you consume this: as a published artefact rather than as a distribution you build. The README build instructions also assume a source archive with a version substituted into a path, which describes a flow that predates the current state. So the evaluation path that costs you the least is the one the repository does not push you toward. Depend on gobblin-api from Maven Central, run the Docker instructions for a laptop, and read the CHANGELOG to see how far past 0.11.0 the code has moved before you decide. The pattern of no releases in nine years with a repository that is still receiving commits is a governance question rather than an abandonment question, and the distinction matters if you are signing up to support it.

Editorial conclusion

Adopt Gobblin if the bottleneck in your lake is maintenance rather than movement, because compaction, partitioning, deduplication, retention and fine-grain deletion are the operations it names first, and atomic data publishing plus state management for incremental processing are the parts you would otherwise build yourself. Do not adopt it as your compute engine or your scheduler, since the README rules out both explicitly and names Spark, Flink, Kafka, Airflow and Dagster as the systems that do those jobs. Two things to settle before you build it. Do not follow the wrapper download line as written, because the documented command retrieves a build tool over a connection with certificate verification disabled, so install Gradle yourself or fetch the jar from a tagged URL and check it. And do not build a distribution from a tag to decide, since the newest release listed is 0.11.0 from 2017-07-20 while the last push was 2026-09-24; consume the gobblin-api artifact from Maven Central instead, and evaluate against the Docker instructions linked from the README.

Frequently asked questions

What is Apache Gobblin and what is it not?

It is a distributed data integration framework for ingestion, replication, organisation and lifecycle management of streaming and batch data. The README explicitly says it is not a transformation engine like Spark or Flink, not a storage system like Kafka or HDFS, and not a workflow system like Airflow, Azkaban, Dagster or Luigi.

How do I build the Apache Gobblin distribution?

Extract the source archive, then run ./gradlew build, which requires Maven for the tests. To skip the checks, run ./gradlew build with findbugsMain, test, rat and checkstyleMain excluded. The distribution appears under build/gobblin-distribution/distributions.

Why does the Apache Gobblin build need a manual Gradle wrapper download?

The source archive does not include gradle-wrapper.jar, so the README gives a wget with certificate checking disabled, a curl with the insecure flag, and a manual download link. The manual route from a fixed tag is the one to use, since the jar is an executable build tool.

What is the latest Apache Gobblin release?

The most recent releases listed are 0.11.0 from 2017-07-20, 0.10.0 from 2017-05-05 and 0.9.0 from 2016-12-19. The last push to the repository was on 2026-09-24, so the default branch is well ahead of the newest tag, and a CHANGELOG.md records the difference.

How do I run the Apache RAT licence audit in Gobblin?

Run ./gradlew rat after extracting the archive. The report is generated under build/rat/rat-report.html.

Official sources

  1. apache/gobblin on GitHub
  2. License: Apache-2.0
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/apache-gobblin.svg)](https://hysenlabs.com/projects/apache-gobblin)