Open-source project
slatedb/slatedb avatar
slatedb/slatedb

SlateDB: an embedded LSM storage engine that keeps its data in object storage

A cloud native embedded storage engine built on object storage.

3,458 stars307 forksRustApache-2.0

At a glance

What is it?
SlateDB is a Rust embedded key-value engine that writes SSTs and a manifest to S3, GCS, ABS, MinIO or any object_store backend. The design trades local disk latency and PUT costs for bottomless capacity and durability, and the README is explicit about that trade.
Who is it for?
Adopt SlateDB when you want an embedded key-value engine whose durability target is an object store you already run, and when you can accept PUT and GET billing plus the flush interval as part of your write path. Do not adopt it if you need sub-millisecond point writes, range deletions today, or a pure local-disk engine with no cloud dependency.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository received new commits within the last day.
What is it written in?
Mainly Rust, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The problem SlateDB targets: an embedded engine whose disk is a bucket

An embedded storage engine normally owns a local disk. RocksDB, LevelDB and their relatives assume that the filesystem underneath is fast, cheap per operation and always reachable. SlateDB keeps the embedded API and the LSM shape but replaces the local filesystem with object storage. The README states that SlateDB writes data to S3, GCS, ABS, MinIO, Tigris and similar services, and that this provides bottomless storage capacity, high durability and easy replication.

The audience is narrow but real. If you are building a service that already stores its durable state in a bucket, SlateDB lets you keep a key-value engine in the process rather than running a separate database cluster. The cost model changes with it. The README names the trade-off directly: object storage has higher latency and higher API cost than local disk. That sentence should decide most adoption questions. A workload that issues millions of small writes per second is not what this engine was shaped for.

How the LSM tree is rearranged around PUT and GET billing

The mechanism is a log-structured merge-tree whose levels live as string-sorted tables (SSTs) in the object store, with a manifest tracking which SSTs are current. The repository layout matches that description: slatedb/ holds the engine, slatedb-common/ shared types, slatedb-cli/ a command line tool, slatedb-bencher/ a benchmarking binary, slatedb-dst/ a deterministic simulation testing crate, and slatedb-txn-obj/ transaction support.

Write cost is handled by batching. The README explains that instead of writing every put() call to object storage, MemTables are flushed periodically as SSTs, and that the flush interval is configurable. A put() updates an in-memory WAL and MemTable and returns a WriteHandle. Durability is opt-in per write or per batch: the README says to call handle.await_durable().await for one write, or db.flush().await to flush all pending writes. Read cost is handled with the standard LSM toolkit the README lists: in-memory block caches, compression, bloom filters and local SST disk caches.

One design consequence deserves attention. Because a write returns before it reaches the bucket, the window between put() returning and await_durable() resolving is where data lives only in process memory. The API makes that window explicit rather than hiding it, which is the right call, but it means callers have to decide per write whether they can tolerate it.

Installing SlateDB and running the README example

SlateDB ships on crates.io, and the README gives the dependency block for a Rust project. It lists slatedb and tokio, both with a wildcard version. Pinning an explicit version is your call; the README does not do it for you.

toml
[dependencies]
slatedb = "*"
tokio = "*"

The README then walks through opening a database against an in-memory object store, putting and getting a key, deleting it, and scanning. The example imports Db and Error from slatedb, and ObjectStore and memory::InMemory from slatedb::object_store, which is the re-exported object_store crate. Db::open takes a path and an Arc<dyn ObjectStore>.

rust
use slatedb::{Db, Error};
use slatedb::object_store::{ObjectStore, memory::InMemory};
use std::sync::Arc;

#[tokio::main]
async fn main() -> Result<(), Error> {
    let object_store: Arc<dyn ObjectStore> = Arc::new(InMemory::new());
    let kv_store = Db::open("/tmp/test_kv_store", object_store).await?;
    Ok(())
}

After the put, the README asserts that get returns Some("test_value".into()), and after the delete that it returns None. It also shows a full-range scan with scan(..), a bounded scan with scan("test_key1"..="test_key2"), a seek to a key, and a close call. If you run the example as written, those assertions are what should hold; if any of them fails, the failure is in the example's own expectations, not in a hidden step.

For real deployments, the same Db::open call takes any backend implementing the ObjectStore trait, so an S3 or MinIO store replaces InMemory without changing the rest of the code. The README does not document a configuration file format, environment variables or a CLI invocation for opening a database, so a first real deployment is a code change, not a config change.

Where SlateDB is the wrong engine

Latency is the first wall. The README says object storage has higher latency than local disk, and the write path adds a flush interval on top. A workload that needs each write durable in single-digit milliseconds has to call await_durable() and then wait for an object store round trip, and the README does not claim otherwise.

Range deletions are the second. The feature list marks range queries, block cache, disk cache, compression, bloom filters, manifest persistence, compaction, transactions, merge operator, clones and change data capture as done, and leaves range deletions unchecked with a link to issue #577. If your access pattern depends on deleting a key range in one operation, the README does not describe that capability as available.

The third is operational coupling. Because the manifest and SSTs live in the bucket, the database is only as available as that bucket and the credentials your process holds. A local-disk engine degrades to a slow disk; this one degrades to whatever the object store does. The README does not document rollback or recovery procedures for a corrupted manifest, so anyone running this in production should treat manifest behaviour as something to verify against the source and the specs/ directory rather than assume from the README.

SlateDB compared with RocksDB and with running a server database

RocksDB is the obvious reference point, and the topic list for the repository names it. Both are LSM-tree embedded engines. The difference is where the sorted files go. RocksDB writes SSTs to a local filesystem and expects that filesystem to be fast and locally attached. SlateDB writes SSTs to object storage and pays a network round trip and a per-request charge for the privilege. RocksDB gives you lower latency and no cloud bill; SlateDB gives you capacity that does not depend on disk size and durability that comes from the bucket's replication.

The second alternative is a server database such as a managed key-value service. That removes the embedded engine entirely: your process talks to a network endpoint, and the database team owns compaction, replication and failover. SlateDB keeps compaction inside your process, which the feature list confirms is implemented, so the operational surface is a library plus a bucket rather than a library plus a bucket plus a cluster. The README does not compare SlateDB to any named server database, so treat that as an architectural difference rather than a claim from the project.

Versioning, release cadence and what the licence allows

The workspace Cargo.toml sets version 0.16.0, and the release list shows v0.16.0 on 2026-08-31, v0.15.0 on 2026-07-29 and v0.14.1 on 2026-07-01. The README states that SlateDB follows Semantic Versioning and releases approximately every two months at the end of each even month. A pre-1.0 version number with that cadence means minor releases can carry breaking changes under SemVer, so a wildcard dependency as shown in the README will move you across those boundaries automatically. Pin the version and read the release notes before upgrading.

The licence is Apache-2.0, declared in both the repository metadata and the workspace package section. Apache-2.0 permits commercial use and modification and includes an explicit patent grant, which matters if you are embedding the engine in a product. It also carries attribution and notice requirements. This is a description of the licence text, not legal advice; if your product ships binaries, have counsel confirm what notices you must include.

Maintenance is visible in the repository activity: the last push was on 2026-09-20, three weeks after the v0.16.0 release, and the repository is not archived. The README also points to a Discord server and a docs site at slatedb.io for anything beyond the README.

Editorial conclusion

Adopt SlateDB when you want an embedded key-value engine whose durability target is an object store you already run, and when you can accept PUT and GET billing plus the flush interval as part of your write path. Do not adopt it if you need sub-millisecond point writes, range deletions today, or a pure local-disk engine with no cloud dependency. Before committing, verify the flush interval and cache defaults against your workload, check that your storage backend implements the object_store trait, and confirm the state of range deletions, which the feature list still marks unchecked.

Frequently asked questions

What object storage backends does SlateDB support?

The README names S3, GCS, ABS, MinIO and Tigris, and states that SlateDB uses the object_store crate, so it works with any backend that implements the ObjectStore trait. SlateDB re-exports that crate for convenience.

Which language bindings does SlateDB ship?

The README lists Go, Java, Node.js/TypeScript and Python as official bindings, and links community bindings for .NET, Ruby and TypeScript. The core engine is written in Rust.

How do I make a write durable in SlateDB?

A write returns a WriteHandle after updating the in-memory WAL and MemTable. Call handle.await_durable().await to wait for one write to become durable, or db.flush().await to flush all pending writes.

Does SlateDB support range deletions?

The feature list leaves range deletions unchecked and links to issue #577, while range queries and most other LSM features are marked done. Treat range deletion as unavailable until the README changes.

What is the best object storage option for SlateDB?

The README does not rank backends. It states that SlateDB works with any storage implementing the object_store trait, and that object storage generally has higher latency and higher API cost than local disk.

Official sources

  1. License: Apache-2.0
  2. Project website
  3. README
  4. Releases
  5. slatedb/slatedb on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/slatedb-slatedb.svg)](https://hysenlabs.com/projects/slatedb-slatedb)