# Datalevin: a Datalog engine shipped to five package ecosystems

> datalevin/datalevin is a Clojure Datalog database that embeds like SQLite, serves a Raft cluster on port 8898, carries documents, vectors, full text search and llama.cpp, and leaves its history questions to a printed book.

**datalevin/datalevin** — A simple, fast and versatile Datalog database

- Repository: https://github.com/datalevin/datalevin
- Website: https://datalevin.org
- Stars: 1,480 · Forks: 86
- Language: Clojure
- License: EPL-2.0
- Published: 2026-09-15 · Updated: 2026-09-15 · Language: en
- Canonical page: https://hysenlabs.com/projects/datalevin-datalevin

## One Clojure core published to five package ecosystems

datalevin/datalevin is a Datalog database whose primary language is Clojure, under EPL-2.0, on the master branch, with 1480 stars, 86 forks and 29 open issues. Its site is https://datalevin.org. The distribution story is much wider than that language suggests. The badge row at the top of the README points at Maven Central for org.datalevin/datalevin-java, npm for datalevin-node, PyPI for datalevin and Clojars, and the install guide adds a babashka pod. One Clojure core therefore reaches Java, Node.js and Python callers through thin packages, and the documentation splits to match, with cljdoc for the Clojure API and javadoc.io for the Java one. The repository carries bindings for python and javascript, and the examples that ship are bulk-load, java and simple-deps, which points at bulk import and embedding as the intended first contact rather than a long lived server. The name comes with a pronunciation guide and a pun: /ˈdadə ˈlevən/, with levin meaning lightning.

## The Datomic query shape over somebody's LMDB fork

The query language is the Datomic flavour of Datalog, and a whole query fits in a form like this one, where 2026 is the only input:

```Clojure
(d/q '[:find  ?name ?total
       :in    $ ?year
       :where [?sales :sales/year ?year]
              [?sales :sales/total ?total]
              [?sales :sales/customer ?customer]
              [?customer :customers/name ?name]]
      (d/db conn) 2026)
```

The implicit joins and recursive rules are what the project points at as the reasons developers prefer this flavour once they have used it. Storage is the other half of the story, and it is not upstream LMDB: the README credits a fork at github.com/huahaiy/dlmdb, chosen for read performance, with write ahead logging and asynchronous transactions built on top of it. ACID is stated as a deliberate choice, with the README preferring the widely accepted principles over what it calls unusual semantics and linking a Jepsen analysis of Datomic as the counterexample. The query optimizer is cost based, and the performance claims point at benchmark directories in the repository rather than at numbers printed in the README. The README also opens with a complaint rather than a pitch: I love Datalog, why hasn't everyone used this already?

## Deletion is immediate, and there is no time axis

The clearest design decision in the project is what happens when data go away. Datalevin behaves the way most other databases behave: when data are deleted, they are gone. The README names the alternative and steps away from it, pointing at Datomic's temporal features as the thing that confuses some users, with a third party post about Datomic's history as the supporting reading. What you get in exchange is a data model that matches the habits of anyone coming from a relational engine, and a query language without a `:as-of` style time dimension to learn. What you give up is the ability to ask what a fact looked like last month, and any audit trail assembled from historical datoms, because the store keeps none. Two related choices sit next to it: the project targets durable storage and long lived data rather than a scratch dataset, and it advertises itself as a fast key value store for EDN data, which is the format Clojure programs already read and write.

## Embedded like SQLite, networked on port 8898, or a babashka pod

The same engine is offered in three shapes, and the choice is left entirely to the caller. The first is a library embedded in an application to manage state, used the way SQLite is used. The second is a networked client and server mode with default port 8898, and this is where the operational surface grows: the README describes a Raft consensus based high availability cluster configuration and full fledged role based access control. A Datalog engine that can become a replicated service with RBAC in the same release is a wider commitment than an embedded store, and the repository does not explain what the cluster path asks of an operator beyond that sentence. The third shape is a babashka pod, which puts the query engine inside shell scripting, where it becomes a command line data tool rather than a service. The install guide carries a section for that pod, alongside the package manager routes for Java, Python, Node.js and Clojure.

## Documents under 2 GiB, SIMD vectors, and its own search engine

Three capabilities sit inside the same store rather than beside it. The first is document storage: Datalevin holds large documents under 2 GiB and builds indexes by paths for JSON, EDN and Markdown documents, which the README positions as a document database role comparable to MongoDB or a PostgreSQL JSONB column. The second is vector search, built on an efficient SIMD accelerated indexing and search library at github.com/unum-cloud/usearch. The third is a full text search engine written for the project, and its history is visible in the resource list, where a 2021 post titled T-Wand: Beat Lucene in Less Than 600 Lines of Code sits among the later entries. The performance comparisons are all delegated to benchmark directories: JOB for the classic graph workload, LDBC-SNB for graph queries, openrulebench for deductive reasoning, plus separate write and search benchmarks. The claim on the README is that the optimizer is competitive with PostgreSQL and SQLite and with Neo4j, and the claim is verifiable only by running those directories.

## Four build entry points, a fuzz corner, and three releases in a month

The root of the repository carries more build machinery than a library of this size needs. `deps.edn` serves the tools.deps workflow, `project.clj` covers Leiningen, `build.clj` is the tools.build script and `release.clj` handles releases, with a `.build/` directory alongside them and a `script/` directory for shell helpers. Tests exist in three places, `test/`, `test-src/` and `test-jar/`, and two of the sibling directories are worth noticing: `jepsen/` and `fuzz/`. A Jepsen directory fits the README's insistence on conventional ACID semantics, and a fuzz directory sits next to it without any mention in the README at all. A `.clj-kondo/` directory means the Clojure linter is configured as well. The release history is short and recent: 1.0.1 and 1.0.2 both published on 2026-08-11, 1.1.0 on 2026-09-03, and the last recorded push is 2026-10-01. CHANGELOG.md sits at the root as the one place that history is written down.

## A built in MCP server, llama.cpp in process, and five chapters that stay offline

The project calls itself AI native and backs that with two pieces of machinery: a built in local MCP server for tool access, and in database embedding plus text generation through a built in llama.cpp. That puts a language model runtime inside the database process rather than beside it, which is a maintenance surface as much as a feature. Documentation splits the same way. The online user guide at https://datalevin.org is searchable and accepts user submitted examples so that practical patterns can be shared, and the API references live on cljdoc and javadoc.io. The long form material is not there: the complete book is sold in print and ebook formats through an Amazon listing, and it includes five additional chapters on AI memory that are not part of the online guide. The README also names a developer documentation path and a series of design posts in reverse chronological order, ending at a 2020 London Clojurians meetup talk. Its last reference stops inside a javadoc.io package summary URL, so the pointer trail ends there.

## Conclusion

Datalevin fits teams that already want Datomic's query language without Datomic's temporal model, and that need the same engine embedded, networked or in a shell. Before committing, decide three things. If your application depends on as-of queries or audit history, note that deletion here is immediate and unrecoverable. If you need the AI memory chapters, they sit in the printed book rather than the online guide. And if you plan to run the networked mode, check what the embedded fork of LMDB and the Raft layer mean for your own backup and failover story, because the repository leaves those operational choices to the reader.

## FAQ

### What kind of database is Datalevin?

It is a Datalog database written in Clojure, licensed EPL-2.0, that speaks the Datomic flavour of Datalog and stores data durably on a fork of LMDB. It can be embedded in an application, run as a client and server service, or loaded as a babashka pod.

### Does Datalevin keep a history of deleted data?

No. The README says Datalevin behaves like most other databases, so when data are deleted they are gone. It deliberately avoids Datomic's temporal features, which means there is no time axis to query and no audit trail to reconstruct from history.

### How do I install Datalevin from Java, Node.js or Python?

Each language has its own package: org.datalevin/datalevin-java on Maven Central, datalevin-node on npm and datalevin on PyPI, with the Clojure artifact on Clojars. A babashka pod is published as well for shell scripting, and the README points to its installation documentation for details.

### Which languages and clients does Datalevin support?

For embedded use it currently supports Java, Python, Node.js and Clojure. It can also run in networked client and server mode on default port 8898, with Raft consensus based high availability and role based access control.

### What search and AI features does Datalevin include?

It indexes JSON, EDN and Markdown documents under 2 GiB by path, does vector search through a SIMD accelerated indexing library, and ships its own full text search engine. On the AI side there is a built in local MCP server and in database embedding and text generation through a built in llama.cpp.

## Sources

- [datalevin/datalevin on GitHub](https://github.com/datalevin/datalevin)
- [License: EPL-2.0](https://github.com/datalevin/datalevin/blob/master/LICENSE)
- [Project website](https://datalevin.org)
- [README](https://github.com/datalevin/datalevin/blob/master/README.md)
- [Releases](https://github.com/datalevin/datalevin/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/datalevin-datalevin
