# NebulaGraph: a distributed graph database for large graphs

> NebulaGraph is an Apache-2.0 distributed graph database written in C++, with storage and computing separation and RAFT-based consistency. Here is what it does, how a deployment is put together, and where it stops being the right tool.

**vesoft-inc/nebula** —   A distributed, fast open-source graph database featuring horizontal scalability and high availability

- Repository: https://github.com/vesoft-inc/nebula
- Website: https://nebula-graph.io
- Stars: 12,409 · Forks: 1,329
- Language: C++
- License: Apache-2.0
- Published: 2026-09-21 · Updated: 2026-09-21 · Language: en
- Canonical page: https://hysenlabs.com/projects/vesoft-inc-nebula

## What NebulaGraph solves, and who it is aimed at

A relational database answers questions about relationships by joining tables. Joins are fine until the traversal gets deep: a query that walks three or four hops across a large dataset turns into a stack of joins whose cost grows with the size of the intermediate result. NebulaGraph is built for the opposite shape of workload, where the graph is the primary data model and traversal depth is a normal part of the query rather than an accident.

The README lists the intended domains directly: social media, recommendation systems, knowledge graphs, security, capital flows, and AI. These share a property. The interesting questions are about paths and neighbourhoods, not about aggregating rows. The README also states that the database handles large volumes of data with millisecond latency, scales up quickly, and performs fast graph analytics. Those are the project's own claims; the README does not publish the benchmark methodology behind them.

The audience is therefore infrastructure teams with a graph large enough that a single-machine graph store has become the bottleneck, and with the operational capacity to run a distributed system. It is not a library you embed in an application process.

## Storage and computing separation, and what RAFT buys

The architecture is split. The README lists storage and computing separation as a core feature, and the architecture diagram in the repository shows the graph service separated from the storage service. Queries are parsed and planned by the graph layer; the data itself lives in the storage layer. Because the two scale independently, you can add query capacity without moving data, or add storage capacity without changing how queries are planned.

Consistency comes from RAFT. The README states strong data consistency by RAFT protocol, which means writes are replicated through a consensus group rather than asynchronously copied. The practical consequence is that a write is not acknowledged until the replication requirement is met, so read-your-writes behaviour holds across the cluster. The cost is latency on the write path and a minimum number of replicas to keep the group healthy; the README does not specify replica counts or quorum rules.

The query language is OpenCypher-compatible, so anyone with Cypher experience can read a NebulaGraph query without learning a new syntax. Access control is role-based, which the README lists as a distinct feature rather than leaving it to a separate product. Graph analytics algorithms are shipped as part of the feature set, which matters because analytics on a transactional graph database usually means exporting data somewhere else first.

## Installing NebulaGraph and running a first query

The README does not embed installation commands. It points to two paths: a download page at nebula-graph.io/download, and building from source, with a link to the install documentation for the 3.8.0 documentation set. There is also a cloud option on AWS. Because the README gives no shell commands of its own, the honest starting point is the documented quick start workflow rather than a copied command line.

What the repository does show is the shape of a source build. The top level contains CMakeLists.txt, a cmake/ directory, a third-party/ directory, and a scripts/ directory, with conf/ holding configuration and docker/ holding container definitions. A source build therefore goes through CMake and pulls in vendored dependencies from third-party/.

```bash
git clone https://github.com/vesoft-inc/nebula.git
cd nebula
mkdir build && cd build
cmake ..
make -j$(nproc)
```

That is the conventional CMake sequence implied by the repository layout. The README does not state the required compiler version, the supported platforms, or the expected build time, so treat the block above as a starting point and follow the install documentation for the exact flags. If you only want to evaluate the query language, the cloud option avoids the build entirely.

Once a cluster is reachable, the workflow is: create a graph space, define a tag (a vertex type) and an edge type, insert data, then traverse. The README does not reproduce this sequence, so the syntax belongs to the quick start documentation. What is worth knowing before you start is that NebulaGraph separates the schema definition step from the data loading step, so nothing can be inserted until the tag and edge types exist.

## Where NebulaGraph is the wrong tool

The most common mismatch is scale. A distributed database with a consensus protocol and separate graph and storage services carries fixed operational overhead: multiple processes, configuration files under conf/, and a cluster that has to stay healthy. For a graph of modest size on one machine, that overhead buys nothing. An embedded graph library or a single-node graph database will be simpler to run, simpler to back up, and simpler to debug.

The second mismatch is transactional. The README lists strong consistency by RAFT, which is about replication of writes across the cluster. It does not claim general multi-statement ACID transactions across arbitrary reads and writes, and nothing in the README suggests that model. If your workload depends on that, NebulaGraph is the wrong shape of system, regardless of how well it traverses.

The third is operational appetite. The README describes NebulaGraph as a distributed graph database with multiple components, and the repository confirms it: separate configuration, a docker/ directory, scripts/, and a tests/ tree. Running it in production means running a distributed system, with everything that implies for upgrades and failure handling. The README links to an upgrade guide only for the historical v1.x to v2.x transition; it does not document rollback for current versions, and version compatibility between releases is something you have to check in the documentation before you plan a migration.

## How it compares with Neo4j and with a relational database

The obvious alternative for graph workloads is Neo4j. The difference in approach is architectural rather than syntactic. Neo4j's widely used deployment model centres on a single primary instance, with clustering available in its commercial editions. NebulaGraph is distributed by design: storage and computing are separate services, and consistency is handled by RAFT from the start. If your graph fits comfortably on one machine and you value a mature tooling ecosystem, Neo4j is the lower-friction choice. If the graph has outgrown one machine and you need horizontal scale with replication built into the open source core, NebulaGraph's model is the one that matches the problem.

The other alternative is not a graph database at all. A relational database with a recursive CTE can answer traversal questions, and for shallow queries over a modest dataset it will do so with tooling your team already knows. The trade-off is depth: as traversal depth grows, the recursive query plan and the intermediate results grow with it. NebulaGraph exists precisely for the point where that stops working. Choosing between them is a question of how deep your queries actually go, not of which technology is more modern.

## Maintenance, releases and the Apache-2.0 licence

The repository is not archived, and the last push was on 2026-09-08. That is recent enough that the codebase is still changing. The release history tells a different story about tagged versions: v3.8.0 was released on 2024-05-17, before v3.6.0 on 2023-08-11 and v3.5.0 on 2023-05-23. The README's quick start and install links point at the 3.8.0 documentation, so v3.8.0 is the version the project's own front page steers you toward.

The gap between the last push and the last tagged release is worth reading carefully. It means development activity and stable release cadence are not the same thing. If you need a tagged, documented version, v3.8.0 is the current answer. If you build from master, you are on your own for compatibility.

Licensing is straightforward. The README states NebulaGraph is under Apache 2.0, and that you can freely download, modify, and deploy the source code, including deploying it as a back-end service for a SaaS offering. That is a permissive licence with an explicit patent grant and no copyleft obligation on your own code. It does not cover the documentation, the website, or any commercial offering from the vendor, and it says nothing about support. This is a description of what the README states, not legal advice; have counsel review it if the deployment matters commercially.

## Conclusion

NebulaGraph fits teams that already run a graph workload too large for a single machine and are willing to operate a multi-service cluster with RAFT-based storage. It is the wrong choice for a small graph on one node, for workloads that need multi-statement ACID transactions, or for anyone who wants a single binary to start. Before committing, verify the upgrade path from your current version to v3.8.0 against the official upgrade documentation, and confirm that the query patterns you need are expressible in the OpenCypher-compatible language. The repository's last push was on 2026-09-08, so the codebase is still receiving changes, but the most recent tagged release is v3.8.0 from 2024-05-17 and the README points to the 3.8.0 documentation set rather than anything newer.

## FAQ

### What is NebulaGraph?

It is an open source distributed graph database written in C++ and licensed under Apache 2.0. The README describes it as handling large volumes of data with millisecond latency, with storage and computing separation, horizontal scalability, and strong consistency via the RAFT protocol.

### How do I install NebulaGraph?

The README does not list install commands. It points to the download page at nebula-graph.io/download, to building from source via the install documentation for 3.8.0, and to a cloud option on AWS.

### Which version of NebulaGraph should I use?

The most recent tagged release is v3.8.0 from 2024-05-17, and the README's quick start and install links point at the 3.8.0 documentation set. The repository itself is still receiving pushes, but building from master means you are tracking unreleased code.

### What query language does NebulaGraph use?

The README lists an OpenCypher-compatible query language among the core features, so queries follow Cypher syntax. Access control is role-based and listed as a separate feature.

## Sources

- [License: Apache-2.0](https://github.com/vesoft-inc/nebula/blob/master/LICENSE)
- [Project website](https://nebula-graph.io)
- [README](https://github.com/vesoft-inc/nebula/blob/master/README.md)
- [Releases](https://github.com/vesoft-inc/nebula/releases)
- [vesoft-inc/nebula on GitHub](https://github.com/vesoft-inc/nebula)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/vesoft-inc-nebula
