Open-source project
jankotek/mapdb avatar
jankotek/mapdb

MapDB: an embedded Java collection library that spills to disk

MapDB provides concurrent Maps, Sets and Queues backed by disk storage or off-heap-memory. It is a fast and easy to use embedded Java database engine.

5,051 stars876 forksJavaApache-2.0

At a glance

What is it?
A Java and Kotlin library that hands you Map, Set and Queue interfaces backed by a B-tree store with off-heap memory, disk overflow, transactions and MVCC.
Who is it for?
MapDB is at its best when a service needs a durable, ordered key value store inside the same JVM as the code using it, with no second process to operate and no client library to speak a wire protocol. The trade is that you inherit an engine with its own storage format, its own tuning surface, and a large open issue count.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 41 days ago.
What is it written in?
Mainly Java, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 7, 2026, and from our analysis. They are not legal advice.

Editorial analysis

Drop-in collections that outlive the heap

The library's central claim is that it combines an embedded database engine with the Java collections framework, so the API you already know is the API you keep. The example in the README is four lines:

java
DB db = DBMaker.memoryDB().make();
ConcurrentMap map = db.hashMap("map").make();
map.put("something", "here");

There is no schema, no connection string and no driver. You get a `ConcurrentMap` back, and every implementation of the collection interfaces that MapDB offers is backed by the same store.

The interesting part is what sits underneath that choice. The README lists five roles the library can take: a drop-in replacement for maps, lists and queues; off-heap collections unaffected by the garbage collector; a multilevel cache with expiration and disk overflow; a relational database replacement offering transactions, MVCC and incremental backups; and utilities for local data processing over large volumes. Those are five different products sharing one engine, and which one you are really buying depends on which line you care about.

Off-heap is the claim that gets the most attention and deserves the most scrutiny, because it is the one with the sharpest edge. Memory outside the Java heap is invisible to the garbage collector by construction, which is exactly the appeal for a multi-gigabyte cache. It also means the collector will never help you reclaim it, so an off-heap store has to enforce its own limits and its own eviction, and MapDB does that through the cache configuration rather than through heap pressure signals.

The README is a landing page, not a manual

What arrives in the repository root is short. It states the goal, lists the five roles, shows a Maven dependency snippet with a `VERSION` placeholder, shows the four line hello world, and then points outward: a Quick Start on gitbooks, and a documentation site at mapdb.org/doc. Support, screenshots and design discussion live off the repository rather than in it.

That means the README cannot answer the questions that decide whether MapDB fits. It does not say which B-tree implementation backs which collection, does not compare the store engines, and does not explain how the cache levels are configured. Those decisions are documented on the external site, and anyone evaluating the library would need to read it there.

A few concrete details do appear, and they are worth noting because they explain the shape of the project. The project is Apache 2 licensed, with `LICENSE.txt` in the tree and a matching Maven Central badge. The development section says MapDB is written in Kotlin and recommends IntelliJ IDEA 15 Community to edit it, which dates that paragraph considerably even though the library still ships. Builds use Maven with `mvn install`.

The most telling line concerns testing. MapDB is said to have extensive unit tests, of which only a tiny fraction run by default so the build finishes in under ten minutes, while the full suite has over a million cases and runs for hours or days. Full runs need `-Dmdbtest=1`. A million test cases is an unusual amount of verification for a library most people meet through a four line example, and it is a reasonable signal about how the maintainer treats correctness in the storage engine.

Choosing a maker is the configuration that matters

The `DBMaker` name in the hello world is the API's most consequential detail. It is a factory of configurations rather than a constructor, and the string method in the example, `memoryDB()`, is one option among several. A maker decides where data lives, whether it is compressed, how the cache is layered, whether records are serialized in place or through an object factory, and what the transaction and snapshot semantics are.

This is where the library earns its reputation and also where it can surprise you. Because collections are views onto a store rather than copies, a maker that enables caching will return entries that are equal but not identical to what you put in, and code that relies on reference identity will misbehave in ways that are tedious to diagnose. Serialization settings compound the issue, because a record has to become bytes at some point and the way that happens is configurable.

The relational features sit on the same foundation. Transactions and MVCC, plus incremental backups, mean the store keeps versioned records rather than overwriting in place. That is what makes a consistent read possible while a write is in flight, and it is also what makes the file larger than a naive dump would be. Choosing MapDB for the collection interface and then relying on the transaction semantics is a coherent choice, but it is two decisions, not one.

The practical order of operations is to start in memory, because a memory maker has no file to corrupt and no tuning to regret, confirm that the collection types you need exist, and only then move to a file backed maker with cache sizes set deliberately.

Fifteen years of Kotlin behind a Map interface

Two facts about this repository deserve to sit side by side rather than be reconciled. GitHub classifies the project as Java, and the default branch is named `release-3.1`. The README's development section states that MapDB is written in Kotlin and suggests IntelliJ IDEA 15. Both are consistent with a long lived project: a Java facing API maintained for years by an author who moved the implementation to Kotlin, still published to Maven Central under the `org.mapdb` group.

The repository root is small in a way that matters for reading it. Beyond `.github/`, the tree holds `.travis.yml`, `.gitignore`, `LICENSE.txt`, `README.md`, `logger.properties`, `pom.xml` and `src/`. So the build is Maven and single module, with Travis as the declared continuous integration, and there is no submodule, no example directory and no bundled benchmark suite at the root.

There are no tagged releases. The README sends you to the external site for release notes rather than to a changelog in the tree, which means the version history is not something you can read without leaving the repository. For a library this widely embedded, that is a real gap, and it is worth knowing before you pin a version.

Activity itself is current: the repository reports a last push on 2026-08-27, which is recent enough that the project is not in maintenance limbo. The last push date and the age of the development instructions are simply measuring different things.

Where MapDB beats a ConcurrentHashMap, and where it does not

The honest comparison is not with a database. It is with the thing you already have in the standard library, plus the two libraries that dominate in-memory caching in Java.

Against `ConcurrentHashMap`, MapDB wins on three axes: the data can exceed the heap, it can expire and spill under a configured cache policy, and it survives a restart. If your working set fits comfortably in memory, never needs eviction, and losing it on restart is acceptable, a `ConcurrentHashMap` is faster, has no configuration, and has no serialization step. MapDB's own README frames the choice as a drop-in replacement, which is accurate about the interface and misleading about the cost, because persistence and caching are not free.

Against Caffeine, the comparison is closer and more interesting. Caffeine is the better general purpose cache, with far better hit-ratio algorithms and a much smaller surface. MapDB's advantage is that its entries can be larger than you want to keep in memory and still be addressable, because the cache is a layer over a store rather than the store itself.

Against an embedded SQL database, MapDB wins on ergonomics for pure key value work, since there is no schema to migrate and no SQL to write. It loses on everything a query engine gives you: secondary indexes you can express, aggregation, joins across collections, and the ability to hand the file to another tool. There is also the practical cost of a library that keeps its own file format. Anything not written by MapDB cannot read it.

The cases where the answer is clearly MapDB are a durable ordered index inside an application, a cache larger than the heap with a bounded memory budget, or a store that needs transactions over versioned records without a second service. Those are narrower than the README's five roles suggest, and they are the ones worth building on.

Reading the project without the documentation

For anyone who wants to understand MapDB rather than just install it, the repository offers less than it could and one thing it offers well. What it offers well is the test suite: over a million cases behind a flag, with a fast subset on by default. A storage engine with that much coverage is a library where the author expects silent data corruption to be the failure mode worth spending engineering time on.

What it offers less of is narrative. The external documentation site carries the design discussion, the screenshots and the release notes, and the gitbooks Quick Start is the only guided path. That split means the repository cannot be read as a self contained explanation of the storage engine, and anyone assessing it seriously will end up on mapdb.org with the source open in another window.

Two practical notes for a first encounter. First, there is no `logger.properties` guidance in the README even though the file sits in the root, so log level configuration is something you will need to discover rather than read. Second, because the default branch is `release-3.1` rather than `main` or `master`, automated tooling and mirrors that assume conventional branch names will not find what they expect.

The clearest way to evaluate MapDB is empirical and cheap: create a memory maker, put a hundred thousand entries in, read them back across several thread pools, then repeat with a file backed maker and a cache smaller than the data. That exercise takes an afternoon and answers the tuning questions the README leaves open more directly than any amount of reading the feature list will.

Editorial conclusion

MapDB is at its best when a service needs a durable, ordered key value store inside the same JVM as the code using it, with no second process to operate and no client library to speak a wire protocol. The trade is that you inherit an engine with its own storage format, its own tuning surface, and a large open issue count. For a plain cache, Caffeine or a ConcurrentHashMap is less work. For a secondary index, an embedded SQL database, or anything that needs to be queried from another language, MapDB gives up its main advantage. Start with `DBMaker.memoryDB()` to check the API shape, then move to a file backed maker with explicit cache sizes once the semantics have stopped surprising you.

Frequently asked questions

Is MapDB still maintained?

The repository reports a last push on 2026-08-27, so the project is not dormant. Be aware that some instructions in the README are dated, notably the recommendation of IntelliJ IDEA 15, and that there are no tagged releases in the repository, so version history lives on the project documentation site instead.

What language is MapDB written in?

The README's development section states that MapDB is written in Kotlin and recommends IntelliJ IDEA to edit it, while GitHub classifies the repository as Java and the default branch is named `release-3.1`. The practical reading is a Java facing API with a Kotlin implementation, published to Maven Central under the `org.mapdb` group.

How does MapDB compare with an in-memory cache like Caffeine?

They solve different problems. Caffeine is a better pure cache, with better hit ratio algorithms and far less configuration. MapDB is a cache layered over a persistent store, so its entries can exceed the heap and still be addressable, with expiration, disk overflow, transactions and MVCC behind them. The comparison only becomes close when the working set is larger than the memory you want to spend.

Do I need to configure a schema before using MapDB?

No. You get a `ConcurrentMap` back from `DBMaker.memoryDB().make()` and can put entries into it immediately, which is the point of the library. The configuration that does matter is the maker itself, since it decides where data lives, how it is serialized and how the cache is layered, and the README does not document those options.

Can MapDB replace a relational database?

It offers transactions, MVCC and incremental backups, which covers durability and versioning. It does not offer a query engine: no secondary indexes you can express, no aggregation, no joins across collections, and a file format only MapDB itself can read. For anything needing SQL or cross process access, an embedded database is the better fit.

Official sources

  1. Issues
  2. jankotek/mapdb on GitHub
  3. License: Apache-2.0
  4. Project website
  5. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/jankotek-mapdb.svg)](https://hysenlabs.com/projects/jankotek-mapdb)