# Apache Avro: a schema-first serialization format with multi-language code generation

> Apache Avro is a data serialization system whose schemas travel with the data. This review covers what it solves, how the Java implementation generates code, and where it is the wrong choice.

**apache/avro** — Apache Avro is a data serialization system.

- Repository: https://github.com/apache/avro
- Website: https://avro.apache.org/
- Stars: 3,306 · Forks: 1,775
- Language: Java
- License: Apache-2.0
- Published: 2026-09-24 · Updated: 2026-09-24 · Language: en
- Canonical page: https://hysenlabs.com/projects/apache-avro

## The problem Avro solves: schemas that travel with the bytes

A JSON payload carries no contract. The field names are there, but nothing tells a consumer that a field is required, that an integer will never arrive as a string, or that a new field was added last week. Teams paper over this with documentation, validation layers, or defensive parsing. Avro takes a different position: every serialized record is written against a schema, and that schema is part of the data format rather than an external document.

The README states plainly that Apache Avro is a data serialization system, and the repository structure backs that up. There is a lang/ directory holding implementations for Java, Python, C, C++, C#, Ruby, PHP, Perl and JavaScript, plus a share/ directory for shared assets and a doc/ directory for documentation. The serialized form is compact and binary, and the schema is expressed in JSON, which means a schema can be read by a human and validated by a machine.

Who is this for? Primarily teams moving records between services or into storage where both ends can agree on a schema ahead of time. The Java implementation is the reference one and the primary language of the repository. If you are writing a one-off script that reads a config file, Avro is more machinery than you need. If you are building a pipeline where a producer and a consumer evolve independently and you want the mismatch to surface as an error rather than a silently missing field, the schema-first model is the point.

## How the schema, the writer and the reader fit together

The mechanism is a contract between two schemas. A writer uses one schema to encode records. A reader declares the schema it expects. Avro resolves the two: fields present in both are read, fields the reader expects but the writer did not write are filled from the reader's default, and fields the writer wrote that the reader does not know about are skipped. That resolution step is what makes schema evolution possible without a coordinated redeploy.

The repository shows this is not a single-library project. Each language under lang/ has its own test workflow, and the README lists separate CI badges for C, C#, C++, Java, JavaScript, Perl, Ruby, Python and PHP, plus CodeQL analysis for C#, Java, JavaScript and Python. That breadth is the real architectural claim: the same binary format is meant to be readable across those languages, so a Python producer can write records a Java consumer reads.

The schema itself is JSON. That choice matters more than it looks. A JSON schema is diffable in version control, reviewable in a pull request, and inspectable without a special tool. The trade-off is that schema resolution happens at read time, so a compatibility mistake is a runtime failure rather than a compile-time one unless you add a compatibility check to your build.

## Installing Apache Avro with Maven and a first real use

The repository is built with Maven at the top level: there is a pom.xml at the root and a .mvn/ directory. The related searches point at the Maven route, and the documented way to work on Avro itself without installing a toolchain is the docker-compose.yml at the top level, which builds a container from share/docker/Dockerfile and mounts the repository at /avro.

```yaml
version: "3"
services:
  avro:
    build:
      context: .
      dockerfile: ./share/docker/Dockerfile
    volumes:
      - ".:/avro"
```

Running `docker compose up` against that file builds the image and mounts your checkout, so edits to the repository are visible inside the container. This is the path the repository documents for developing Avro, not for consuming it as a library in your own project.

The README also points at devcontainers as a development route, with an Open in Visual Studio Code link and an Open in Github Codespaces link, so a contributor can start from a browser without a local Maven install. Beyond those two routes, the README does not give install steps for consuming Avro from another project, and it does not document the code generation plugin or its configuration keys. Treat anything you read about a Maven plugin or a Gradle plugin as coming from somewhere other than this repository's README.

## The Rust SDK moved out, and other limits worth knowing

The most concrete limitation is stated at the top of the README in an important note: the Rust SDK is moving to https://github.com/apache/avro-rs, and the README asks that new issues and pull requests go there instead. If you arrived looking for Rust support inside apache/avro, you are in the wrong repository. The related searches include a Rust query, which suggests people land here first and have to be redirected.

A second limitation is structural. The multi-language promise depends on each lang/ implementation keeping pace with the format. The README lists separate test workflows per language, which tells you the project treats them as independently verified rather than as one uniform product. A feature that lands in the Java implementation is not automatically available in the Perl or PHP one, and the README does not claim otherwise.

The third is the one that trips people up most often: Avro is a row-oriented serialization format, not a columnar store. If your workload is analytical scanning over a few columns of a wide table, Avro will read and decode every field you stored, and that is the wrong shape of work for it. The related searches pair Avro against Parquet for exactly this reason. Avro is the right tool for record-at-a-time transport and for landing data where the schema needs to evolve; it is the wrong tool when the query pattern is columnar aggregation.

## Avro against Parquet, JSON, Protobuf and Thrift

The comparisons people search for are not all the same kind of comparison, and it is worth separating them.

Against JSON, the difference is the schema. JSON is self-describing in a loose sense: field names are present, but types and requiredness are conventions. Avro writes a compact binary payload and relies on the schema to interpret it, so the same record is smaller and the contract is explicit. The cost is that you cannot open an Avro file in a text editor and read it. You need a tool or a library, which is why "how can I open an Avro file" is a common question.

Against Parquet, the difference is layout, not schema philosophy. Parquet stores data by column, which makes selective analytical reads cheap. Avro stores records row by row, which makes whole-record writes and reads natural and makes column pruning ineffective. Both are common in the same pipeline: Avro for the transport and landing layer, Parquet for the query layer.

Against Protobuf and Thrift, the difference is where the schema lives and how it evolves. Protobuf and Thrift generate code from an interface definition language and rely on field numbers for compatibility. Avro keeps the schema in JSON, writes it alongside the data, and resolves writer and reader schemas at read time. That read-time resolution is what allows a reader to be built against a schema the writer never saw. The trade-off is a resolution step on every read and a class of errors that only appear when data is actually decoded. The related searches also pair Avro with Arrow and gRPC, but those address different layers: Arrow is an in-memory columnar representation, and gRPC is a remote procedure call framework. Neither is a drop-in substitute for a serialization format.

## Maintenance, releases and the Apache-2.0 licence

The repository is not archived, and the last push was on 2026-09-21. The most recent release tag is release-1.12.2, dated 2026-08-23, following release-1.12.0 on 2024-08-05 and release-1.11.3 on 2023-09-25. That release history shows a gap of roughly two years between the 1.11.3 and 1.12.0 tags, so a team that pins a version should expect long intervals between feature releases and plan upgrades around that cadence rather than assuming a steady stream.

Upgrade cost is dominated by the schema, not the library. Because Avro resolves writer and reader schemas at read time, a library upgrade is usually less disruptive than a schema change. The reverse is also true: a careless schema change can break readers in a way no dependency bump would. The repository's own build is Maven-based at the root, so a team already on Maven has a short path; a Gradle user will need to find the Gradle plugin elsewhere, since the README does not document one.

The project is licensed under Apache-2.0, and the README carries a trademark notice: Apache, Apache Avro and the Apache Avro airplane logo are trademarks of The Apache Software Foundation. Apache-2.0 is a permissive licence that permits commercial use and modification, but the trademark notice is separate from the licence and governs use of the name and logo, not the code. Whether your redistribution needs a NOTICE file is a question for your own legal review; the repository does ship LICENSE.txt and NOTICE.txt at the top level, which is the pattern to follow.

## Conclusion

Adopt Apache Avro when you control both the writer and the reader and want schemas checked at build time rather than at runtime, and when a JVM stack with Maven is already in place. Do not adopt it as a columnar store for analytical scans; that is Parquet's job, and Avro's row layout will not serve those queries. Before committing, verify the reader and writer schema compatibility rules for your own records, confirm which of the lang/ directories your team will actually maintain, and check whether the Rust SDK at apache/avro-rs covers your use case, since the README states it has moved out of this repository.

## FAQ

### What is Apache Avro used for?

It is a data serialization system, per the README, used to encode records against a schema so that producers and consumers agree on the shape of the data. The repository ships implementations under lang/ for Java, Python, C, C++, C#, Ruby, PHP, Perl and JavaScript.

### Why use Apache Avro instead of JSON?

Avro writes a compact binary payload interpreted through a JSON schema, so the contract is explicit and the records are smaller than their JSON equivalents. The cost is that an Avro file is not readable in a text editor and needs a tool or library to inspect.

### Is Apache Avro better than Parquet?

They solve different problems. Avro stores records row by row, which suits record-at-a-time transport and schema evolution; Parquet stores data by column, which suits analytical scans over a few columns of a wide table. The related searches pair them because they are often used in the same pipeline rather than as substitutes.

### How can I open an Avro file?

You need a library or tool rather than a text editor, because the payload is binary and interpreted through a schema. The repository provides implementations under lang/, and the Java one is the primary language of the project.

### How do I use Apache Avro?

The README points to the project website at https://avro.apache.org/ for learning more, and the repository documents two development routes: the docker-compose.yml at the top level, and devcontainers with Open in Visual Studio Code or Open in Github Codespaces links.

### Is Apache Avro still used?

The repository is not archived, and the last push was on 2026-09-21. The most recent release tag is release-1.12.2, dated 2026-08-23, which followed release-1.12.0 on 2024-08-05 and release-1.11.3 on 2023-09-25.

## Sources

- [apache/avro on GitHub](https://github.com/apache/avro)
- [License: Apache-2.0](https://github.com/apache/avro/blob/main/LICENSE)
- [Project website](https://avro.apache.org/)
- [README](https://github.com/apache/avro/blob/main/README.md)
- [Releases](https://github.com/apache/avro/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/apache-avro
