# AutoMQ: Diskless Kafka on S3, and What That Means for Your Broker Bill

> AutoMQ keeps the Kafka API but moves stream storage to object storage, so brokers hold no durable data. Here is how the pieces fit, what the repository actually documents, and where the approach stops being the right answer.

**AutoMQ/automq** — Diskless Kafka® on S3. 10x Cost-Effective. No Cross-AZ Traffic Cost. Autoscale in seconds. Single-digit ms latency. Multi-AZ Availability.

- Repository: https://github.com/AutoMQ/automq
- Website: https://www.automq.com
- Stars: 10,874 · Forks: 778
- Language: Java
- License: Apache-2.0
- Published: 2026-09-21 · Updated: 2026-09-21 · Language: en
- Canonical page: https://hysenlabs.com/projects/automq-automq

## The problem AutoMQ targets: broker disks that hold data you would rather not own

In a conventional Kafka deployment, each broker owns a set of local disks, and a partition lives on the brokers that host its replicas. Scaling storage means adding brokers or attaching volumes, and cross-availability-zone replication between replicas shows up on the network bill. The README frames the project around exactly that: "Diskless Kafka on S3, Offering 10x Cost Savings and Scaling in Seconds." The cost claim is the vendor's, and the repository does not contain the benchmark harness behind it, so treat the multiplier as a marketing figure rather than a measured result.

The audience is narrower than "anyone running Kafka." AutoMQ is for teams whose stream data is already a derived copy of something durable elsewhere, and who want broker count to track CPU and network rather than retained bytes. Log pipelines, change-data-capture fan-out and event buses fit. A system where the broker disk is the only copy of the data, with no upstream replay available, is a different conversation, and the README does not claim to solve it.

## How the diskless design splits the brokers from the bytes

The repository layout tells most of the story. There is a top-level s3stream/ directory, and a storage/ directory alongside the familiar core/, metadata/, raft/, connect/ and streams/ modules inherited from the Kafka codebase. The naming points at the architecture: stream data is written to object storage through an S3 stream layer, while the Kafka protocol surface stays where clients expect it.

That split is what makes the scaling claim possible. A broker that does not own durable bytes can be replaced without copying partitions, because a new broker reads from the same object storage the old one wrote to. It also explains the project's framing around multi-AZ availability and the absence of cross-AZ traffic cost: replicas that would otherwise replicate bytes between zones are instead reading a shared object store. The trade-off is that every fetch path now depends on object storage latency and on the credential and endpoint configuration being correct. Single-digit millisecond latency is the claim in the repository description; the README does not publish the conditions under which it holds, so a reader should measure it against their own access pattern rather than assume it.

Metadata is a separate concern from stream data. The presence of metadata/ and raft/ at the top level suggests coordination state is still managed by the cluster rather than pushed to S3, which is consistent with how a Kafka-compatible system has to behave: consumer group state and partition leadership cannot live in a blob store with the same semantics.

## Getting AutoMQ running when the README has no quickstart

The README does not carry an install block. It points to the documentation site and to an AutoMQ Playground account for a hosted trial, and the repository itself ships the build and packaging paths. The Gradle wrapper is present at the root, so building from source is the first route. The repository also includes a container/ directory, a docker/ directory and a chart/ directory, which is where a container image and a Helm chart would come from; the README does not spell out the image name or the chart repository URL, so check those directories rather than guessing.

The repository pins a Java version in .java-version, and the build uses the wrapper:

```bash
./gradlew build
```

Running that produces the distribution under the build output, and the bin/ directory at the root holds the launch scripts. Configuration lives in config/, which is where the S3 stream settings belong. The README does not reproduce those keys, so read the files in config/ before editing them; the ones that matter are the bucket, region and credential entries the s3stream layer needs.

Once a broker is up, the client side is unchanged, which is the point of the project. The examples/ directory at the root holds a README.md, a bin/ directory and a src/ directory, and that is where the repository puts runnable sample code. Read examples/README.md before writing your own producer, since the README at the root does not show a client snippet and the repository does not document a specific bootstrap flag or port in the text available here.

If a topic can be created and records accepted, the broker is speaking the Kafka protocol. What that does not prove is that records landed in object storage rather than a local volume; the README does not describe a command to confirm the storage backend, so verify it by inspecting the bucket you configured in config/.

## Where the diskless model becomes the wrong choice

The first limitation is operational, not architectural. A diskless broker is only as available as its object storage and its credentials. If the S3 endpoint is unreachable or the IAM role expires, the failure is not a slow disk, it is a stalled fetch path across the cluster. The README's multi-AZ availability claim addresses zone failure; it does not address a regional object storage outage, and the repository does not document a fallback mode for one.

The second is the release cadence. The listed releases are 1.7.4 from 2026-08-29, 1.7.5-rc0 from 2026-09-02, and a nightly build from 2024-08-23. That is a young line with a release candidate in flight, and the README does not describe a long-term support policy or a compatibility guarantee across minor versions. A team that needs a version it can sit on for two years should read the release/ and docs/ directories before assuming one exists.

The third is scope. AutoMQ is not a drop-in answer for every Kafka workload. If your consumers rely on features the project has not carried over, or if your deployment is a single node on a laptop, the object storage dependency adds configuration and a network hop for no benefit. The README does not enumerate unsupported Kafka features, which is itself a gap worth noting: the only way to know is to test your own client and connector set.

## AutoMQ vs WarpStream and vs stock Kafka: three different storage bets

The comparison that matters is about where bytes live and who pays for moving them. Stock Apache Kafka keeps partition data on broker disks and replicates between brokers, often across zones. AutoMQ keeps the Kafka protocol and API but moves stream data to object storage through its s3stream layer, leaving brokers without durable local state. WarpStream, frequently named alongside AutoMQ in search queries, takes the same object storage premise but is written in Go rather than Java, which changes the operational footprint and the amount of the Kafka codebase you inherit.

That last difference is the practical one. AutoMQ's repository is a Java project with Kafka's module structure intact (core/, connect/, streams/, raft/, clients/), so existing JVM-side tooling and the Kafka client ecosystem are close at hand. A Go implementation makes a different trade: a smaller runtime to deploy, but a separate codebase from the one your Kafka-adjacent tools were written against. Neither choice is free, and the README does not benchmark AutoMQ against WarpStream, so the repository offers no data to settle it.

## Licence, maintenance and the cost of upgrading

AutoMQ is Apache-2.0, and the repository carries LICENSE, NOTICE and NOTICE-binary files at the root. The NOTICE-binary file is the one to read if you redistribute the built artifacts, since it records third-party components bundled into the binaries. Apache-2.0 permits commercial use and modification; it also means the project carries no warranty, which is standard and worth remembering when the storage layer is your object store.

Maintenance is active by the only measure available here: the last push was on 2026-09-21, and the repository is not archived. That says the code is being touched, not that any given release is stable. Upgrading means tracking the 1.7.x line and reading the release notes in release/ for each step, because the README does not publish an upgrade procedure or a rollback path. If you build from source, note that the pinned Java version in .java-version is part of your upgrade surface, not just your build environment.

There is an operational cost that does not appear in a licence file. Diskless brokers shift spend from provisioned EBS volumes to object storage requests and egress. The README advertises no cross-AZ traffic cost, but it does not model per-request pricing, and a high-throughput small-message workload can move the bill in the other direction. Model that before treating the cost claim as a budget line.

## Conclusion

Adopt AutoMQ if you already run Kafka clients and want broker storage off the critical path, you are comfortable operating on AWS object storage, and you can accept a Java project whose README sends you to the documentation site rather than walking you through a quickstart. Do not adopt it if your workload is a single-node development cluster, if you need a long-term support release rather than the 1.7.x line, or if you cannot run the S3, IAM and network setup the diskless design assumes. Verify four things before committing: that your Kafka client versions and custom serializers work against an AutoMQ broker, that the S3 stream settings in config/ match your bucket and credentials, that the container image or Helm chart you intend to use is the one you tested, and that a broker restart under load leaves your consumer offsets where you expect them. The last one is the whole point of the design, so test it first.

## FAQ

### What are the key differences between Kafka and AutoMQ?

AutoMQ keeps the Kafka protocol and API but stores stream data in object storage through its s3stream layer instead of on broker disks, so brokers hold no durable bytes. The repository description frames the result as lower cost, no cross-AZ traffic cost and scaling in seconds, though the README does not publish the benchmark conditions behind those figures.

### What is AutoMQ?

It is an Apache-2.0 Java project described in its README as diskless Kafka on S3. The repository keeps Kafka's module layout (core/, connect/, streams/, raft/, clients/) and adds an s3stream module for the object storage path.

### Is AutoMQ an alternative to Kafka?

It is Kafka-compatible rather than a replacement protocol: the README presents AutoMQ as a Kafka-compatible system and the repository retains the clients/ and connect/ modules from the Kafka codebase. The difference is the storage layer, which moves from broker disks to object storage.

## Sources

- [AutoMQ/automq on GitHub](https://github.com/AutoMQ/automq)
- [License: Apache-2.0](https://github.com/AutoMQ/automq/blob/main/LICENSE)
- [Project website](https://www.automq.com)
- [README](https://github.com/AutoMQ/automq/blob/main/README.md)
- [Releases](https://github.com/AutoMQ/automq/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/automq-automq
