Apache Hudi: upserts, deletes and incremental queries on object storage
Upserts, Deletes And Incremental Processing on Big Data. Apache Hudi Apache Hudi is an open data lakehouse platform, built on a high-performance open table format to ingest, index, store, serve, transform and manage your data across multiple cloud data environments.
At a glance
- What is it?
- Hudi is an open table format and lakehouse platform for Spark and Flink writers. It earns its place when you need record-level updates and change streams, and it costs you table services you have to run.
- Who is it for?
- Adopt Hudi when your pipeline needs record-level updates, deletes, or a change stream read back out of the table, and you already run Spark or Flink. Skip it if you only append immutable files and query them with Trino or Presto; a plain Parquet layout plus a metastore is less machinery.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 4 days ago.
- What is it written in?
- Mainly Java, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 25, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The problem Hudi solves: mutable rows on immutable object storage
Object stores do not let you overwrite a row inside a Parquet file. That single constraint is why data lakes historically forced append-only ingestion: you land every change event, then deduplicate at read time, and the table grows forever. Hudi's premise is that the table format itself should absorb that work. The README describes the project as an open data lakehouse platform built on a high-performance open table format, and the features it lists are all consequences of that stance: fast upsert and delete support using record-level indexes, atomic commits with rollback and restore, and snapshot isolation between writers and readers.
The intended user is not a general analyst. It is a data engineer who owns an ingestion job and has to reconcile late-arriving records, corrections, and GDPR-style deletions against a table that other engines query. Hudi ships built-in ingestion tools for Apache Spark and Apache Flink users, and a Connect sink for Apache Kafka, so the entry point is a job you already run rather than a new service you deploy. If your data is genuinely append-only and never corrected, Hudi's indexing and table services are overhead you are paying for nothing.
Timeline, file groups and the index: how a Hudi write actually lands
The repository layout shows the split clearly. hudi-common and hudi-io hold the format and metadata layer, hudi-client holds the write path, hudi-timeline-service is a separate process, and each engine gets its own module: hudi-spark-datasource, hudi-flink-datasource, hudi-trino, hudi-kafka-connect. The format is not tied to one compute engine, which is why the same table can be written by Flink and read by Trino.
Writes are grouped into commits recorded on a timeline, which the README describes as timeline metadata tracking the history of changes. A commit is atomic, and the README states that rollback and restore are supported, so a failed write does not leave the table in a half-applied state. Within a table, records are routed to file groups, and the index decides which file group owns a given record key. That index is what makes upserts cheap relative to a full shuffle-and-rewrite: the README describes record-level indexing mechanisms built on row-oriented file formats and bloom filters, plus a scalable indexing subsystem that tracks file listings and column-level and partition-level statistics.
Table services run alongside the writers. Compaction turns row-oriented data into columnar files asynchronously, cleaning expires older versions, and clustering rewrites layout. The README says these are configurable scheduling strategies with built-in failure handling, and that they can run inside Spark or Flink writers or independently. That is the real architectural trade-off: you get incremental maintenance instead of batch rewrites, and in exchange you now operate a scheduler whose failure modes affect query latency.
Building Apache Hudi from source and running a first Spark session
The README does not point at a package manager. It documents building from source with Maven, on a Unix-like system with Java 11 or 17, Git, and Maven 3.6.0 or newer. The clone and build command is given directly:
git clone https://github.com/apache/hudi.git && cd hudi
mvn clean package -DskipTests -Dspark3.5 -Dflink2.2The build produces bundle jars under packaging/. The README then shows how to start a Spark shell against the built bundle, passing the bundle jar plus four configuration keys:
spark-3.5.0-bin-hadoop3/bin/spark-shell \
--jars `ls packaging/hudi-spark-bundle/target/hudi-spark3.5-bundle_2.12-*.*.*-SNAPSHOT.jar` \
--conf 'spark.serializer=org.apache.spark.serializer.KryoSerializer' \
--conf 'spark.sql.extensions=org.apache.spark.sql.hudi.HoodieSparkSessionExtension' \
--conf 'spark.sql.catalog.spark_catalog=org.apache.spark.sql.hudi.catalog.HoodieCatalog' \
--conf 'spark.kryo.registrator=org.apache.spark.HooNote that the last line is cut off in the README as published; the Kryo registrator class name is incomplete there, so copy it from the project site rather than from this excerpt. Once the shell starts with those extensions loaded, Hudi tables are addressable through Spark SQL. The README does not walk through creating a table or writing a first record, so the concrete DDL and insert syntax has to come from the documentation site at hudi.apache.org.
Build profiles matter more than they look. The default is Spark 3.5.3 with Scala 2.12, and the README's table maps flags to bundle names: -Dspark3.3, -Dspark3.4, -Dspark3.5 with -Dscala-2.13, and -Dspark4.0 for Spark 4.0 with Scala 2.13. Getting this wrong produces a jar that will not load in your cluster, which is the most common first-run failure. There is also a Dockerfile in the repository root, but it builds from a private base image named apachehudi/hudi-ci-bundle-validation-base and only runs java -version, so it is a CI artifact rather than a supported way for a user to get started.
Five query types on one table, and what each one costs
Hudi's distinguishing feature is that the same physical table answers different questions. The README lists Snapshot Query for the latest committed state, Incremental Query for the latest value of records changed since a point in time, Change-Data-Capture Query for a change stream with before and after images, Time-Travel Query for a view as of a given instant, and Read Optimized Query for columnar-only reads.
The practical distinction is between the first four and the last. Read Optimized Query is the fast path: it reads purely columnar storage such as Parquet and relies on a compaction policy to provide a transaction boundary. That means it can lag the latest writes, because recent changes may still sit in row-oriented files that compaction has not yet merged. Snapshot Query sees the latest committed state but has to merge those files at read time. Choosing between them is a latency decision, not a correctness one, and the README is explicit that the read-optimized view depends on compaction policy rather than being automatically current.
Incremental Query and Change-Data-Capture Query are the ones with no clean equivalent in a plain Parquet layout. The README frames incremental reads as a way to diff table states between two points in time, and CDC queries as returning both before and after images per change record. That is what makes Hudi viable as a source for downstream replication instead of re-scanning full partitions.
Where Hudi is the wrong tool
The cost is operational surface. A Hudi deployment includes writers, a timeline, and table services that must be scheduled: compaction, cleaning, clustering, and index building. The README describes these as automatic and hands-free when integrated into Spark or Flink writers, but it also states they can be operated independently, which is the configuration most production teams end up in. Every one of those services has a scheduling strategy and a failure path. If your team has no one to own that, an append-only layout will be more reliable than a Hudi table whose compaction has silently stopped.
Concurrency is the second boundary. The README lists optimistic concurrency control for Read-Modify-Write style consistent writes, and non-blocking concurrency control for streaming with out-of-order and late data. Those are different guarantees, and a workload that assumes one while configured for the other will see write conflicts or unexpected ordering. The README does not document rollback of a completed commit, only rollback as part of commit handling, so undoing a bad write that already succeeded is not covered there.
Finally, the build story assumes Java and Maven. There is no pip install, no single binary. If your team is Python-first and does not run Spark or Flink, Hudi is a poor fit regardless of how well the table format suits the data.
Hudi compared with Iceberg and Delta Lake
The most common question about Hudi is how it differs from Apache Iceberg and Delta Lake, and the difference is visible in the module list. Iceberg's center of gravity is the table specification and catalog interoperability, with engines implementing reads and writes against that spec. Delta Lake's is a transaction log layered on Parquet, tightly coupled to the Spark and Databricks ecosystem. Hudi's is the write path: record-level indexes, upsert and delete primitives, and a set of table services that maintain the layout in the background.
That shows up in what each project emphasizes. Hudi's README leads with ingestion from change logs and streaming systems, a Kafka Connect sink, and CDC queries that return before and after images. It also lists catalog sync with Apache Hive Metastore, AWS Glue, Google BigQuery and Apache XTable, which is the concession that the surrounding ecosystem expects a catalog it recognizes. If your priority is a minimal format that many engines read, Iceberg's approach is a better match. If your priority is streaming upserts into a table that stays queryable, Hudi's index and table services are doing work the other two leave to the writer.
Maintenance, releases and the Apache-2.0 licence
The repository is not archived, and the last push was on 2026-06-08. Recent releases listed are 0.14.2 on 2026-06-08, 1.2.0 on 2026-05-23, and 0.15.1 on 2026-05-21. Three release lines receiving tags within a few weeks is the notable fact here: 0.14.x, 0.15.x and 1.2.x are all being maintained, so a team on an older line is not forced to jump to 1.x immediately, but it does have to track which line gets fixes.
Upgrade cost is dominated by the engine matrix rather than the table format. The build table ties each Spark version to a specific bundle artifact, and the README notes that Scala 2.13 is supported for Spark 3.5 and above, with Spark 4.0 requiring Scala 2.13. Moving your cluster's Spark version therefore means rebuilding the bundle and re-verifying your writer configuration, not just bumping a dependency. The README does not document an upgrade procedure for existing tables, so compatibility across table versions is something to confirm against the project site before a major move.
Licensing is Apache-2.0, the same licence as the Hadoop and Spark stack Hudi runs on, with the standard patent grant and attribution requirements. The repository carries both LICENSE and NOTICE files, and the ASF header appears on source files. That is a permissive arrangement, but it is not legal advice, and anyone redistributing a bundle should read the NOTICE file rather than assume it is empty.
Editorial conclusion
Adopt Hudi when your pipeline needs record-level updates, deletes, or a change stream read back out of the table, and you already run Spark or Flink. Skip it if you only append immutable files and query them with Trino or Presto; a plain Parquet layout plus a metastore is less machinery. Before committing, decide which table type you need by checking whether the README's Read Optimized Query description matches your latency requirement, and confirm which Spark or Flink profile in the build table matches the engine version you actually run.
Frequently asked questions
What is Apache Hudi used for?
It is an open data lakehouse platform built on an open table format for ingesting, indexing, storing and managing data across cloud environments. Its distinctive capabilities are fast upsert and delete using record-level indexes, incremental and change-data-capture queries, and table services that maintain file layout automatically.
Is Hudi better than Apache Iceberg?
They optimize for different things. Hudi's README leads with ingestion from change logs and streaming systems, record-level indexes for upserts, and CDC queries returning before and after images. Iceberg's approach centers on the table specification and catalog interoperability. Which is better depends on whether your workload is streaming updates or broad engine compatibility.
What does Hudi do?
It provides an open table format plus ingestion tools for Spark and Flink, a Kafka Connect sink, a timeline that records the history of changes, and table services for compaction, cleaning, clustering and index building. On top of one table it answers snapshot, incremental, change-data-capture, time-travel and read optimized queries.
What are the key differences between Apache Hudi, Apache Iceberg, and Delta Lake?
Hudi's emphasis is the write path: record-level indexes, upsert and delete primitives, and background table services such as compaction, cleaning and clustering. Delta Lake layers a transaction log on Parquet and is closely tied to the Spark ecosystem, while Iceberg focuses on a format specification that many engines implement. Hudi also ships catalog sync for Hive Metastore, AWS Glue, BigQuery and Apache XTable.
What is Apache Hudi vs Parquet?
Parquet is a columnar file format; Hudi is a table format that stores data in open formats such as Parquet and adds a timeline of commits, record-level indexes, and table services on top. The README notes that read optimized queries serve purely columnar storage for fast snapshot performance, which is the Parquet-like path.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/apache-hudi)