# BadgerDB: an embeddable Go key-value store built around the WiscKey design

> BadgerDB is a pure Go, embeddable key-value database from the Dgraph project, aimed at SSD-backed workloads that need ACID transactions without Cgo. This review covers its LSM-plus-value-log design, how to install it, and where it is the wrong choice.

**dgraph-io/badger** — Fast key-value DB in Go.

- Repository: https://github.com/dgraph-io/badger
- Website: https://dgraph-io.github.io/badger/
- Stars: 15,780 · Forks: 1,324
- Language: Go
- License: Apache-2.0
- Published: 2026-09-21 · Updated: 2026-09-21 · Language: en
- Canonical page: https://hysenlabs.com/projects/dgraph-io-badger

## What BadgerDB solves, and who it is actually for

BadgerDB is an embeddable key-value database written in pure Go. The word embeddable is doing most of the work in that sentence: there is no server process to start, no port to expose, no separate deployment to keep alive. You link the library into your Go program and the database lives in your process. The README states the project is the underlying database for Dgraph, a distributed graph database, and that it is meant as a performant alternative to non-Go-based key-value stores like RocksDB.

That framing tells you who the library is for. If you are building a Go service and you want local persistence with sorted key access, atomic batches and transactions, Badger is a candidate. If you are building in another language, the pure Go requirement is irrelevant to you and the main reason the project exists evaporates. The Cgo constraint matters because it changes your build pipeline: a Cgo dependency means a C toolchain, cross-compilation friction, and a different story for static binaries. Badger removes all of that.

The design goals listed in the README are narrow and consistent: write a key-value database in pure Go, use recent research to build the fastest KV database for data sets spanning terabytes, and optimize for SSDs. Note the last one. This is not a general-purpose storage engine tuned for spinning disks or for tiny embedded devices. It is tuned for the access pattern and hardware that the WiscKey paper describes.

## The LSM tree plus value log split, and what it costs

Badger's design is based on the WiscKey paper, whose central idea is separating keys from values in SSD-conscious storage. The README's comparison table describes Badger as an LSM tree with value log, against RocksDB as an LSM tree only and BoltDB as a B+ tree.

The mechanism is worth stating plainly. Keys and their metadata live in the log-structured merge tree, which keeps them sorted and compact. Values are written to a separate append-only value log. Because the LSM tree no longer carries the value payload, compaction rewrites far less data, and the README attributes that to the WiscKey paper's finding of significantly reduced write amplification compared to a typical LSM tree.

The cost is on the read path. A point lookup has to find the key in the LSM tree, then follow the pointer into the value log to fetch the value. For large values this is a clear win. For small values it is overhead you pay on every read for a benefit you do not receive. The README does not quantify where the crossover sits, and the benchmark logs it points to live in a separate repository. Treat the value-size threshold as something you measure on your own data rather than something the project hands you.

The other consequence is file layout. A value log is an additional set of files on disk with its own lifecycle, and the README does not document that lifecycle inline. It directs readers to the documentation site at badger.dgraph.io. If you are evaluating Badger for an environment where you must control exactly what files appear on disk, that is a gap you have to close before adopting.

## Installing BadgerDB and writing a first key

The README requires Go 1.23 or above and states that Badger v3 and above needs Go modules. The repository's go.mod declares go 1.24.0, so the module itself is built against a newer toolchain than the README's minimum. From your project directory, fetch the library:

```bash
go get github.com/dgraph-io/badger/v4
```

This retrieves the library and records the dependency in your go.mod. There is no daemon to start afterward.

For the offline command line tool, the README says to retrieve the repository, check out the desired version, and run go install from the repository root:

```bash
cd badger
go install .
```

The README states this installs the badger command line utility into your $GOBIN path. The README describes the tool as able to perform certain operations like offline backup and restore. It does not enumerate the full command set, so check the help output after installing rather than assuming a flag exists.

A minimal program follows the usual embedded-database shape: open a database at a directory path, run a transaction, set a key, read it back, close. The README does not print a full code sample in the section shown, so the API surface you should expect is the one documented on pkg.go.dev under github.com/dgraph-io/badger/v4. Two things are worth knowing before you write that code. First, the comparison table lists TTL support and sorted key access, so both are properties of the store rather than things you build on top. Second, the table's footnote on 3D access says Badger provides direct access to value versions via its Iterator API, and that you can specify how many versions to keep per key via Options. That is a configuration decision you make at open time, not a query-time choice, and getting it wrong means either retaining more history than you need or losing versions you wanted.

## Transactions, isolation, and the testing the project relies on

Badger supports concurrent ACID transactions with serializable snapshot isolation guarantees, and the README states this as a property rather than a goal. It also describes how the project tries to earn that claim: a Jepsen-style bank test runs nightly for eight hours with the --race flag, and Badger has been tested against filesystem-level anomalies to check persistence and consistency.

That is a more specific testing story than most embedded stores publish, and it is the kind of claim you can inspect, because the repository ships the test files. The top-level listing includes db_test.go, db2_test.go, db_corrupt_test.go, iterator_test.go, batch_test.go, merge_test.go, manifest_test.go, and a set of level-specific tests including l0_backpressure_test.go and l0_blockwrites_regression_test.go. The presence of a corruption test file and a backpressure test file tells you the maintainers treat write stalls and on-disk damage as first-class concerns rather than edge cases.

The comparison table also makes a pointed claim about the competition: RocksDB is listed as having transactions that are non-ACID, while Badger and BoltDB are listed as ACID. If your application logic depends on snapshot isolation across concurrent readers and writers, that row is the one that decides the choice. If it does not, you are paying for isolation you never use, and a simpler store may serve you better.

One caveat about the nightly test. The README does not state what hardware it runs on or how long the suite takes in a developer's environment. A nightly eight-hour job is not something you will reproduce on a laptop before a release.

## When BadgerDB is the wrong tool

The clearest failure mode is value size. The entire design premise is that moving values out of the LSM tree reduces write amplification. If your workload stores small values, you get the extra indirection of the value log without the compaction savings that justified it, and a B+ tree store like BoltDB, which the README lists as having high read throughput, may fit better.

The second case is read-heavy workloads with few writes. The README's comparison table lists RocksDB as not having high read throughput and BoltDB as having it, while Badger is listed as yes on both read and write throughput. That table is the project's own summary, and the README points to a separate benchmarking repository for the logs behind it. If your access pattern is overwhelmingly reads, run that comparison against your own data before trusting a row in a feature matrix.

The third case is anything requiring more than a key-value interface. The README does not claim a query language, secondary indexes, or replication. Badger is the storage layer beneath Dgraph, not a replacement for it. If you want graph traversal or a declarative query language, you want Dgraph, and the README says so by describing Badger as its underlying database.

Finally, consider operational expectations. There is no server, so there is no health endpoint, no connection pool to tune, and no separate process to restart. If your team's operational model assumes a database you can connect to over a network, embedding Badger means moving that responsibility into your application binary.

## How BadgerDB compares with RocksDB and BoltDB

The README's comparison table is the most concrete comparison available, and the differences it draws are about design lineage rather than benchmarks. RocksDB is an LSM tree only and is not pure Go, so using it means Cgo. The README's footnote argues that RocksDB is an SSD-optimized version of LevelDB, which was designed for rotating disks, and therefore RocksDB's design is not specifically aimed at SSDs. BoltDB is a B+ tree, pure Go, embeddable, and ACID, but the table lists it as not designed for SSDs and as lacking TTL support and the key-value-version access that Badger offers.

So the practical split is this. Choose BoltDB when your data set fits comfortably in memory-mapped files and your workload is read-dominant; you get a simpler tree structure and no value log. Choose RocksDB when you need its ecosystem, its language bindings, or the maturity of a C++ codebase, and you accept Cgo in your build. Choose Badger when you want pure Go, SSD-oriented design, ACID transactions with SSI, TTL, and per-key version retention in one library.

The comparison table is the project's own framing, and the README is explicit that the benchmarking code and detailed logs live in the badger-bench repository. That is the right place to look if you need numbers, because the README itself does not reproduce them.

## Version pinning, licensing and the upgrade surface

Badger is licensed under Apache-2.0. That is a permissive licence, and the repository carries SPDX headers in its files, including the Makefile, which reads SPDX-License-Identifier: Apache-2.0. The practical implication is that you can embed the library in a closed-source product, but you should read the licence text yourself and confirm how your distribution model handles attribution. This is not legal advice.

Version discipline deserves attention. The go.mod file contains two retractions: retract v4.0.0, referencing issues #1888 and #1889, and retract v4.3.0, referencing issues #2113 and #2121. A retraction means the module author is telling the Go toolchain not to select those versions. If your build is pinned to either, you are on a version the project has withdrawn, and the retraction will affect how the module graph resolves.

The module path carries a major version suffix, github.com/dgraph-io/badger/v4, so a v5 would be a separate import path and a deliberate migration rather than a patch bump. Recent releases in the v4.9.x line have shipped at a steady cadence through mid-2026, and the last push to the repository was on 2026-09-21. The README also notes that Badger is built with go 1.23 and that the project refrains from bumping that version to minimize downstream effects for applications built with older Go versions, while go.mod declares go 1.24.0. Reconcile those two statements against your own toolchain before you pin.

Upgrade cost is mostly in the Options you pass at open time. Because version retention per key is configured through Options, changing that setting is a migration decision, not a runtime toggle. There is a VERSIONING.md file in the repository root that the README does not summarize; read it before planning a major upgrade.

## Conclusion

Adopt BadgerDB if you are writing a Go service that needs an embedded, transactional key-value store and you do not want a Cgo dependency on RocksDB or a server process to operate. Do not adopt it if your values are small and your reads dominate, because the value log exists to move large values out of the LSM tree and it costs you a second lookup; do not adopt it if you need a query language, secondary indexes or replication, none of which the README claims. Before committing, verify the Go version your build chain actually uses against the go.mod requirement, and confirm that your deployment target can tolerate the file layout Badger creates on disk, since the README points at the documentation site for operational detail rather than spelling it out inline.

## FAQ

### What is BadgerDB and what problem does it solve?

BadgerDB is an embeddable, persistent key-value database written in pure Go, and it is the underlying database for Dgraph. It exists as a performant alternative to non-Go-based key-value stores like RocksDB, so Go programs can get an embedded transactional store without a Cgo dependency.

### How do I install BadgerDB in a Go project?

The README requires Go 1.23 or above and states that Badger v3 and above needs Go modules. From your project you run go get github.com/dgraph-io/badger/v4, which retrieves the library.

### Does BadgerDB support ACID transactions?

Yes. The README states that Badger supports concurrent ACID transactions with serializable snapshot isolation guarantees, and that a Jepsen-style bank test runs nightly for eight hours with the --race flag.

### What is the difference between BadgerDB and RocksDB or BoltDB?

The README's comparison table describes Badger as an LSM tree with value log, RocksDB as an LSM tree only, and BoltDB as a B+ tree. Badger and BoltDB are listed as pure Go, while RocksDB is not, and the table lists BoltDB as lacking TTL support.

## Sources

- [dgraph-io/badger on GitHub](https://github.com/dgraph-io/badger)
- [License: Apache-2.0](https://github.com/dgraph-io/badger/blob/main/LICENSE)
- [Project website](https://dgraph-io.github.io/badger/)
- [README](https://github.com/dgraph-io/badger/blob/main/README.md)
- [Releases](https://github.com/dgraph-io/badger/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/dgraph-io-badger
