Model or dataset
openlake-project/openlake avatar
openlake-project/openlake

OpenLake: a Rust storage engine that turns GPU hosts into a KV cache pool

OpenLake is a high performance storage engine for efficient LLM inference and GPU Training

2,655 stars422 forksRustApache-2.0

At a glance

What is it?
OpenLake is an Apache-2.0 storage engine written in Rust on io_uring, aimed at LLM inference and GPU training workloads. Its vLLM connector and S3-compatible object store are the two entry points, and the README is more concrete about the first than the second.
Who is it for?
Adopt OpenLake if you already serve long-context models on vLLM and want KV cache to outlive a single request, or if you need a persistent store co-located with GPU hosts. Do not adopt it if you need a general-purpose object store with a documented API surface; the README shows the S3 path only far enough to point an AWS CLI at port 90.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 2 days ago.
What is it written in?
Mainly Rust, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The problem OpenLake targets: KV cache that dies with the request

Long-context inference spends most of its compute on prefill. A 128K context window means the model reads a large prompt before it emits a single token, and if the same prompt arrives again the work is repeated. OpenLake exists to keep that computed state around. The README describes the inference engine writing KV once and reading it back in milliseconds from host RAM and disk, which removes prefill for long and repeated prompts. The repository also frames the same engine around checkpoint storage for RL and ML workloads, vector index building, and context storage for agentic retrieval, so the intended user is not only an inference team. It is anyone whose GPU hosts sit idle waiting on I/O. The topics list on the repository includes rdma, gpu, llm-training and model-serving, which matches that framing. What the README does not do is quantify the benefit outside its own blog links; the 66x time-to-first-token figure it prints is attributed to a cached 128K context window, and the underlying benchmark is not reproduced in the README itself.

How the storage engine is built: compio, io_uring and a workspace split

The Cargo.toml workspace is the clearest architectural statement in the repository. Members are crates/* plus cli, so the server, the client connector and the command line tool are separate compilation units sharing one dependency graph. The async runtime is compio 0.18 with default-features disabled and an explicit feature list of runtime, macros, io-uring, fs, net, time and rustls. That is a completion-based runtime rather than a poll-based one, which is why the HTTP layer needs a bridge: the comment in Cargo.toml states that Cyper bridges hyper's hyper::rt::Read and hyper::rt::Write traits onto compio's I/O surface, and that Cyper-axum is the matching axum::serve replacement. Tokio still gets linked transitively through hyper, hyper-util, h2 and axum. The comment also records a version decision: the umbrella compio crate is used instead of pinning sub-crates by hand, with the note that the 0.18 series ships the same compio-runtime 0.11, compio-fs 0.11, compio-io 0.9, compio-buf 0.8, compio-driver 0.11 and compio-macros 0.1.2 that were previously pinned. The one sub-crate that moves is compio-tls, from a direct 0.9 pin to 0.4 pulled transitively. That is a real maintenance trade-off: fewer pins to manage, less control over the TLS adapter version.

Installing the vLLM connector and serving a model with OpenLake

The fastest path in the README is the KV pool on GPU nodes, and it claims no code changes are needed. Install the connector package and start the daemon:

bash
pip install openlake-vllm
openlaked

The package name is openlake-vllm and the binary is openlaked. With the daemon running on the same host, the README's next step is to launch vLLM with a transfer config that names the connector module. Note the PYTHONHASHSEED export before the serve command:

bash
export PYTHONHASHSEED=0
vllm serve <model_name> --kv-transfer-config '{"kv_connector":"OpenLakeConnector","kv_connector_module_path":"openlake_client.openlake_connector","kv_role":"kv_both","kv_connector_extra_config":{"openlake_nodes":["127.0.0.1:9400"],"openlake_device":"local"}}'

The config keys are kv_connector, kv_connector_module_path, kv_role and kv_connector_extra_config, which itself carries openlake_nodes and openlake_device. In this default form the node list is 127.0.0.1:9400 and the device is local, so the store is on the same host. The README states that offloading across a GPU fleet requires starting openlaked with a --config. For Kubernetes, it points at charts/openlake/README.md for a Helm deployment that places one instance per selected node and generates the ordered vLLM peer configuration.

Building the S3-compatible store from source

The second half of the quickstart is a four-step build. Dependencies come from apt, then rustup, then a release build of the openlaked binary:

sh
sudo apt-get install -y build-essential pkg-config clang cmake libhwloc-dev libudev-dev curl git awscli
curl --proto '=https' --tlsv1.2 -sSf https://sh.rustup.rs | sh -s -- -y && . "$HOME/.cargo/env"
git clone https://github.com/openlake-project/openlake.git && cd openlake
cargo build --release --bin openlaked

The README then creates four data directories and starts the store against a checked-in config file:

sh
mkdir -p data/d0 data/d1 data/d2 data/d3
./target/release/openlaked --config crates/openlake_server/configs/storage-tcp-local.toml

Client access is shown as plain AWS CLI environment variables and an endpoint on port 90. The README's snippet is truncated mid-command, ending at `aws --endpoint-url http://127.0.0.1:90`, so the credentials are openlakeadmin for both AWS_ACCESS_KEY_ID and AWS_SECRET_ACCESS_KEY with region us-east-1, but the exact request the reader should issue is not printed. Treat that as the boundary of what the README documents for the object-store path.

Where OpenLake is the wrong tool, and what to check first

The README is explicit that the default configuration offloads to the same host. Anyone reading "petabyte scale KV cache" and expecting a cluster out of the box will be disappointed by the single-node default; the multi-host path requires a config file per node with a self_id, and the README's example runs openlaked --config kv_rdma_0.toml and kv_rdma_1.toml with ids 0 and 1 respectively. The ordering is load-bearing, because the README states a prefix computed on one GPU host is served to any other from the shared pool. Get the id order wrong and the peer configuration generated for vLLM will not line up. The other gap is operational surface. The README does not document rollback, snapshotting, or what happens to an existing KV pool when you upgrade between releases, and it does not document the object store's S3 API coverage beyond pointing the AWS CLI at it. If you need a general-purpose S3 replacement with a documented compatibility matrix, this repository does not give you one. If you need a store that sits next to the GPUs and serves cached prefixes fast, that is the case it was built for.

How OpenLake differs from MinIO and from a local NVMe cache

MinIO is the natural comparison for the object-store half, and the difference is placement rather than protocol. MinIO is a general S3 server you deploy on storage hardware and point clients at over the network. OpenLake's README describes the store as co-located on GPU hosts, with the KV pool built from the GPU nodes themselves, so the data path stays on the machine that needs it and RDMA is available in the multi-host config via openlake_device set to an interface such as mlx5_ib0. A local NVMe cache differs in the other direction: it is fast and simple, but it is not a shared pool, so a prefix computed on one host is not served to another. The README's claim that a prefix is served across hosts is the specific thing neither a per-host cache nor a remote object store gives you without extra plumbing. The cost is that you now run and coordinate a stateful service on every GPU node, and the Helm chart exists precisely because that coordination is not trivial.

Maintenance, release cadence and the Apache-2.0 licence

The repository is not archived and the last push was on 2026-09-06, which is recent. Release history shows 0.9.0 on 2026-09-05, v0.8.1 on 2026-08-14 and v0.8.0 on 2026-08-12, so the project is cutting releases on a short cadence. Note one inconsistency worth knowing before you file issues: the workspace Cargo.toml still declares version 0.8.0 while the latest release tag is 0.9.0, and the README's update entry for ExANS refers to OpenLake v0.8. The Rust toolchain requirement is also split across files: the README badge says rust 1.91+, while Cargo.toml sets rust-version to 1.88 and rust-toolchain.toml is the file the badge links to. Check rust-toolchain.toml rather than either number if you pin a toolchain. The licence is Apache-2.0, which permits commercial use and modification and requires you to preserve notices and state changes; it also includes an explicit patent grant. That is a permissive baseline, and it says nothing about the separate terms of the openlake-vllm Python package distributed on PyPI, which the README does not discuss.

Editorial conclusion

Adopt OpenLake if you already serve long-context models on vLLM and want KV cache to outlive a single request, or if you need a persistent store co-located with GPU hosts. Do not adopt it if you need a general-purpose object store with a documented API surface; the README shows the S3 path only far enough to point an AWS CLI at port 90. Before committing, verify that the openlake_nodes addresses in your kv_connector_extra_config match the self_id ordering of each openlaked instance, because the README states a prefix computed on one GPU host is served to any other from the shared pool only when that ordering holds.

Frequently asked questions

What is OpenLake?

OpenLake is a storage engine for LLM inference and GPU training, written in Rust on io_uring and licensed Apache-2.0. The README describes it as distributed storage for GPU workloads, used for KV cache offload, checkpointing, vector indexing and context storage.

How does OpenLake differ from a closed storage system?

The repository is public under Apache-2.0 and the server is built from source with cargo build --release --bin openlaked, so you can read and modify the engine. The README does not compare OpenLake against any closed-source storage product, so the licence and the source build are the concrete differences it supports.

How do I install OpenLake for vLLM?

Run pip install openlake-vllm and then start the openlaked daemon. After that, launch vLLM with --kv-transfer-config naming OpenLakeConnector and the module openlake_client.openlake_connector, with openlake_nodes pointing at the store.

Does OpenLake work across multiple GPU hosts?

Yes, but not by default. The README states that offloading to the same host is the default and that a fleet-wide setup requires starting openlaked with a --config, such as kv_rdma.toml with a per-node self_id.

Is OpenLake an S3-compatible object store?

The README shows the store being reached with the AWS CLI against an endpoint on port 90 using the openlakeadmin credentials, and describes it as S3 compatible. The README does not publish a list of which S3 operations are covered.

Official sources

  1. License: Apache-2.0
  2. openlake-project/openlake on GitHub
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/openlake-project-openlake.svg)](https://hysenlabs.com/projects/openlake-project-openlake)