Open-source project
hydra-db/hydradb avatar
hydra-db/hydradb

HydraDB: A Rust Graph Database on S3-Compatible Object Storage

HydraDB - fast graph database on object storage

12,725 stars5,539 forksRustAGPL-3.0

At a glance

What is it?
HydraDB is a distributed graph database written in Rust that stores its graph data on S3-compatible object storage and serves queries over Neo4j-compatible Bolt and an HTTPS API. It is an early-stage project licensed under AGPL-3.0, with its last push on 2026-08-19.
Who is it for?
HydraDB is appropriate for teams that want a graph database with disaggregated storage and compute, plan to run on S3-compatible infrastructure, and can work with an AGPL-3.0-licensed codebase. The AGPL-3.0 license means that if HydraDB is used as a network service, derivative works must be released under the same license.
Can I use it commercially?
Yes, with strict conditions. AGPL-3.0 is a network copyleft licence: if people use a modified version over a network, for example as a hosted service, you must offer them its source code under the same licence.
Is it still maintained?
Yes. The repository received new commits within the last day.
What is it written in?
Mainly Rust, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

DEEP OPEN-SOURCE ANALYSIS

What HydraDB Is and the Problem It Solves

HydraDB is an object-store-native distributed graph database written in Rust. The central design idea is that graph storage and compute are fully disaggregated: S3-compatible object storage holds the durable copy of the graph, and compute nodes hold only disposable state in memory and on local SSD. The README describes this as meaning compute nodes can be replaced or scaled without moving the graph itself.

The problem HydraDB solves is the tight coupling between storage and compute in traditional graph databases. When graph storage is tied to a specific compute node's local disk, scaling out requires either replication or migration. HydraDB avoids this by making object storage the canonical source, so any compute node can read from or write to the same graph without coordinating storage migration.

The target audience is infrastructure engineers and data teams who need a graph database that integrates with cloud object storage, want Neo4j-compatible client tooling, and are operating in an environment where Bolt drivers are already used for other systems.

Architecture: Data Nodes and Indexers

HydraDB separates compute into two independent roles. Data nodes, called graph-node, serve queries and canonical mutations. Indexers, called graph-indexer, build immutable traversal indexes in the background and publish them through atomic object-store pointers.

Each data node has a private local SSD or NVMe cache. The object store is the shared layer beneath the whole tier and the only durable copy of the graph. The README states that readers remain correct when an index is absent or behind because the visible write-ahead log (WAL) tail is applied to the indexed base.

Writer coordination uses object-store CAS (compare-and-swap) leases to select the active writer for each cell, while SlateDB writer epochs fence stale writers. This prevents two nodes from simultaneously issuing conflicting writes to the same cell. Every query runs against one pinned SlateDB snapshot, which provides snapshot-consistent reads.

The query engine uses property indexes, reverse adjacency, sparse traversal, and SuiteSparse GraphBLAS where appropriate. Applications connect using either Neo4j drivers over Bolt 5.x or the typed JSON and streaming NDJSON HTTP API.

Running HydraDB with Docker

The fastest way to get a single development node running is the published Docker image. Pull the latest image:

bash
docker pull ghcr.io/hydra-db/hydradb:latest

Images are published for linux/amd64 and linux/arm64. Releases before 0.1.0 were linux/amd64 only. The README notes that pulling an older tag on an ARM host fails with a "no matching manifest" message, which indicates the tag predates multi-architecture publishing rather than a configuration error.

Before running the container, create local directories for the store and cache, and write an auth token file:

bash
mkdir -p hydradb-data/store hydradb-data/cache
printf '%s\n' 'local-development-token-32-bytes' > hydradb-data/auth-token

The docker run command requires the `--user` flag matching the host user's UID and GID, because the image runs as UID 10001 but the bind-mounted directories are owned by the host user. Without this flag, the container cannot write to the store or cache directories. The README states this explicitly: running as the host user makes the mounted directories writable and keeps created files host-owned.

After the node starts, the README recommends running a round-trip write to verify the node is actually functioning rather than just listening on a port.

Building from Source

Building HydraDB from source requires Rust 1.91 or newer, a C/C++ toolchain, libcypher-parser, and SuiteSparse GraphBLAS. On Ubuntu or WSL, install the system dependencies:

bash
sudo apt-get install -y \
  build-essential clang libclang-dev cmake pkg-config \
  libcypher-parser-dev libgraphblas-dev

On macOS with Homebrew:

bash
brew install just cmake pkg-config llvm suite-sparse

The repository uses just (a command runner, configured in justfile) for common development tasks. The justfile defines targets for formatting, linting, building, and running the smoke tests. The repository also uses mise.toml for toolchain management and cargo-chef in the Dockerfile for efficient Docker layer caching of dependencies.

The Rust workspace root is the main crate. Two additional crates live in crates/placement and crates/telemetry. The telemetry crate provides OpenTelemetry tracing behind an off-by-default otlp feature flag.

OpenCypher Compatibility and Bolt Connectivity

HydraDB serves queries using OpenCypher, the graph query language also used by Neo4j. Applications that already use Neo4j drivers for Bolt 5.x can connect to HydraDB without changing client code, since HydraDB implements Neo4j-compatible Bolt connectivity. The repository includes a cypher-compat.md file that documents the compatibility scope.

The HTTPS query API returns results as typed JSON or streaming NDJSON, which allows applications to query the graph without a Bolt driver. The HTTP interface is the alternative path for environments where a Bolt client library is not available or where streaming JSON results are preferable.

The README notes that authentication, authorization, deadlines, result limits, backpressure, cancellation, cache budgets, metrics, and traces are part of the server runtime. These are operational concerns that a production deployment needs, and the README lists them explicitly as features of the runtime rather than optional additions.

Limitations: Early Stage and AGPL-3.0

HydraDB has no GitHub releases. The last push was on 2026-08-19. The project is at an early development stage where the README is the primary documentation and the setup process involves configuring a non-trivial set of environment variables for the Docker run command.

The AGPL-3.0 license is a significant consideration for commercial adoption. AGPL-3.0 is a copyleft license that extends to network use: if a modified version of HydraDB is run as a service accessible over a network, the AGPL-3.0 requires the modified source to be made available to users of that service. Teams building proprietary SaaS products on top of HydraDB need to assess this requirement before committing to the project.

The Docker images before version 0.1.0 do not support ARM64, which means Apple Silicon developers must use emulation with --platform linux/amd64 for those older tags. For current development, the latest tag includes ARM64 support.

The source build requirements (libcypher-parser, SuiteSparse GraphBLAS) are not standard system packages in most Linux distributions and require explicit installation, adding friction to the developer setup process.

HydraDB vs. Neo4j

Neo4j is the most widely known property graph database and shares several interface characteristics with HydraDB: both use OpenCypher as a query language and Bolt as the client protocol. The architectural difference is storage. Neo4j's community and enterprise editions store graph data on local disk attached to the database server. HydraDB stores graph data on S3-compatible object storage with compute handled by separate stateless nodes.

This means HydraDB's compute nodes can be replaced without any storage migration, which is the design's central advantage over a traditional local-storage graph database. Neo4j's local storage model makes it straightforward to run on a single machine with no external storage dependencies, while HydraDB requires an S3-compatible object store as a prerequisite.

Neo4j is a mature product with stable releases, extensive documentation, and commercial support. HydraDB is an early-stage Rust project with no formal releases. Teams that need production reliability and support contracts today should use Neo4j. Teams building infrastructure on cloud object storage and willing to operate an early-stage project may find HydraDB's architecture more suitable.

Editorial conclusion

HydraDB is appropriate for teams that want a graph database with disaggregated storage and compute, plan to run on S3-compatible infrastructure, and can work with an AGPL-3.0-licensed codebase. The AGPL-3.0 license means that if HydraDB is used as a network service, derivative works must be released under the same license. Teams building a SaaS product should review the license conditions before adopting HydraDB. The project has no GitHub releases as of 2026-08-19, and the setup process requires configuring several environment variables. Before adopting HydraDB, run the verification steps documented in the README to confirm that the node responds to an actual round-tripped write, not just a listening port.

Frequently asked questions

What is HydraDB?

HydraDB is a distributed graph database written in Rust that uses S3-compatible object storage as its durable data store. It serves queries over Neo4j-compatible Bolt and an HTTPS API, using OpenCypher as its query language and SuiteSparse GraphBLAS for graph traversal.

Is HydraDB open source?

The source code is publicly available on GitHub under the AGPL-3.0 license. AGPL-3.0 is a copyleft license that requires modified versions run as a network service to make their source available to users of that service.

What is the main alternative to HydraDB?

Neo4j is the most widely known graph database with the same OpenCypher query language and Bolt protocol. The key difference is that Neo4j stores data on local disk attached to the server, while HydraDB stores data on S3-compatible object storage with stateless compute nodes.

Official sources

  1. hydra-db/hydradb on GitHub
  2. Issues
  3. License: AGPL-3.0
  4. Project website
  5. README
For maintainers

Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/hydra-db-hydradb.svg)](https://hysenlabs.com/projects/hydra-db-hydradb)
Community notes

Community notes