apache/hugegraph: one graph database, two very different deployment shapes
A graph database that supports more than 100+ billion data, high performance and scalability (Include OLTP Engine & REST-API & Backends)
At a glance
- What is it?
- An Apache graph database where the same server runs on an embedded RocksDB for development or on Raft-replicated PD and HStore for production.
- Who is it for?
- HugeGraph is at its best when your data is genuinely relational in the graph sense and your queries are traversals rather than joins. The Gremlin compliance with TinkerPop 3 means existing traversal code and a large body of literature apply, and the Cypher engine means teams arriving from Neo4j do not have to relearn their query language.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository received new commits within the last day.
- What is it written in?
- Mainly Java, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 7, 2026, and from our analysis. They are not legal advice.
Editorial analysis
A graph engine wrapped in three interchangeable front doors
The architecture diagram in the README is the most useful thing in the repository, because it shows that the query layer and the storage layer are separate concerns. At the top is a client layer offering the Gremlin Console, a REST API, Cypher, and SDKs and tools. Below that sits the HugeGraph server on port 8080, and inside it three engines: a REST API built on Jersey 3, a Gremlin engine at TinkerPop 3.5, and a Cypher engine implementing OpenCypher.
All three converge on one component, the graph engine in hugegraph-core. That is the design decision worth understanding: adding Cypher support did not create a parallel database, it created another way to drive the same engine. Anyone coming from Neo4j can keep Cypher, anyone coming from the TinkerPop ecosystem can keep Gremlin, and the storage behaviour underneath is identical.
That also explains the schema emphasis in the feature list. VertexLabel, EdgeLabel, PropertyKey and IndexLabel are first-class managed metadata, and indexes come in three flavours: exact query, range query, and combinations of complex conditions. In a property graph database that lets you define shape after the fact, every query is a scan. In HugeGraph, you declare your labels and indexes, and the payoff is a query engine that knows where to look.
Standalone or distributed, and the line between them
The same server binary runs in two modes, and the README draws them side by side with a data scale attached to each.
Standalone mode means RocksDB embedded in the server process, a single node, for development and testing, with a stated ceiling of under 1TB. There is no second process to coordinate and no replication to reason about. If you are evaluating whether HugeGraph fits your data model, this is the mode you should use, and it is also the mode most people will never leave.
Distributed mode replaces the embedded store with two Raft-replicated components. HugeGraph-PD runs on 3 to 5 nodes on ports 8620 and 8686 and handles the metadata and coordination layer. HStore runs on 3 or more nodes on port 8520 and holds the graph data. Both use Raft, so each has its own quorum to reason about, and both have their own CI pipeline, which tells you they are treated as separate failure domains.
Distributed mode is described as the production, high availability and cluster option with a scale ceiling under 1000TB. Read the backend evolution guide linked from the README before you pick a version. The store layer changed from RocksDB to HStore for clustered deployments, and that guide exists to explain lifecycle and historical compatibility for people who are still on the older backend.
Two query languages is a genuine differentiator
Most graph databases pick one traversal language and ask you to live with it. HugeGraph supports both Gremlin through TinkerPop 3 and Cypher through OpenCypher, and given that Gremlin's ecosystem spans a decade of published traversals, having a second entry point costs the project little and saves an adopter a rewrite.
The choice also tells you something about the audience. Gremlin is the Apache project's own lineage and is what the OLTP orientation is built around, with billions of vertices and edges stored and queried directly. Cypher is the language most people encounter first, in Neo4j or in documentation about knowledge graphs, and supporting OpenCypher means a Cypher query written elsewhere has a reasonable chance of running here unchanged.
The project also states a Big Data integration story with Flink, Spark and HDFS, which is the natural adjacent concern for a graph store: bulk loading and graph algorithms running in a cluster, feeding results back. Whether that integration is smooth in practice is a question for the external documentation rather than the README.
The ecosystem is four more repositories
The server repository is one part of a larger set, and the README is explicit about the boundaries so you know where to look.
hugegraph-toolchain is the tools suite, and it holds four subprojects. Loader is the data import tool. Dashboard is the web visualization platform, referenced by its package name hugegraph-hubble. Tool is a set of command-line utilities. Client is the Java and Python client SDK. Then hugegraph-computer is described as an integrated graph computing system, hugegraph-ai handles graph AI and large language model and knowledge graph integration, and hugegraph-website, whose repository is actually hugegraph-doc, holds the documentation and website.
The visualization piece deserves a note, because it is the answer to the most common objection to graph databases, which is that you cannot see what is in them. A load-it-and-look-at-it dashboard is exactly what you want during schema design, when you are trying to find out whether your label structure is the right shape. The fact that it lives in a separate repository rather than in the server is normal for this project and does not mean it is optional in practice.
The whole set, server plus toolchain plus computer plus AI plus docs, is what you are evaluating when you choose HugeGraph. Only the server is in this repository, and it is Apache-2.0 licensed with 3,188 stars, 637 forks and 372 open issues.
What the releases say about where the work went
Three releases are published, and the gap between them is long enough to say something. Version 1.3.0 shipped on 2024-03-29, 1.5.0 on 2024-12-10, and 1.7.0 on 2025-11-16, so roughly a year and a half separates the first and last of these.
The 1.5.0 notes carry the single most important operational line in the release history: starting from that version, a Java 11 runtime environment is required. Its work is also structural rather than cosmetic, integrating the pd-grpc, pd-common and pd-client modules and their store equivalents, which is the distributed layer being pulled apart into proper modules. It also fixes the RocksDB backend being switched to memory during Gremlin examples, which is the kind of bug that makes an evaluation demo look fast for the wrong reason.
Version 1.7.0 is smaller and more specific. It adds memory management for the graph query framework, and fixes two bugs you would rather not discover in production: dynamic paths with parameters causing an out of memory condition on PUT, GET and DELETE requests, and a NaN value in the JRaft histogram metrics in HStore. The rest of that release is distribution hygiene, updated docs for an older version, license fixes, and dependency adjustments.
One oddity is worth flagging so it does not confuse you. The 1.7.0 release is named without the incubating prefix, but the download links for every release still point at the incubator path on downloads.apache.org, and the two earlier releases still carry the incubating label in their names.
A repository shaped like an Apache release, not a product
The tree is what you would expect from a project that has been through incubation and still cares about process. There is an .asf.yaml for the Apache metadata, a .licenserc.yaml for automated license header checks, a NOTICE file alongside LICENSE, an .editorconfig, and a style directory. A .serena directory and an AGENTS.md file sit at the top level too, which is a sign the project has started wiring tooling for automated code navigation.
The module layout mirrors the architecture rather than a generic Maven layout. There is hugegraph-server for the query and REST layer, hugegraph-core lives inside it, hugegraph-pd for the metadata service, hugegraph-store for the storage layer, hugegraph-commons for shared code, and hugegraph-struct. Two entries stand out as operational rather than library code: hugegraph-cluster-test, which is how you would verify a real multi-node deployment, and install-dist plus a docker directory for packaging.
Nothing in the tree suggests a young or abandoned project. The last push was on 2026-09-20, the repository is not archived, and there are separate CI workflows for the server and for the PD and store modules. What you will not find here is a Dockerfile for the whole cluster, deployment manifests, or tuning guidance for a specific workload. Those live on hugegraph.apache.org, and the README is candid that its job is to get you to the point where that documentation becomes the thing you need.
Editorial conclusion
HugeGraph is at its best when your data is genuinely relational in the graph sense and your queries are traversals rather than joins. The Gremlin compliance with TinkerPop 3 means existing traversal code and a large body of literature apply, and the Cypher engine means teams arriving from Neo4j do not have to relearn their query language. The architecture is honest about scale in a way few projects are, with a stated boundary of under 1TB standalone and under 1000TB distributed, which tells you when to stop. Two things deserve caution before you commit. The project moved from standalone RocksDB to clustered HStore and left a compatibility guide behind for people still on the old backend, so version choice matters more here than usual. And the schema is explicit, with vertex labels, edge labels, property keys and index labels you define up front, which is a deliberate trade of flexibility for query speed. Start standalone, model your graph in the schema layer, and only reach for PD and HStore when single node genuinely stops working.
Frequently asked questions
Is HugeGraph free to use?
Yes. Apache HugeGraph is released under the Apache License 2.0, and the source is in the Apache GitHub organization. The download links for each release are published on the Apache downloads mirror, so there is no separate commercial edition to evaluate.
What query languages does Apache HugeGraph support?
Two. It is compliant with Apache TinkerPop 3 and runs Gremlin traversal queries through a Gremlin 3.5 engine, and it also runs Cypher through an OpenCypher engine. Both drive the same graph engine in hugegraph-core, so the choice of language does not change storage behaviour.
How does HugeGraph differ from Neo4j or NebulaGraph?
The obvious difference is membership. HugeGraph is an Apache project under the Apache License 2.0, while Neo4j is commercial with a community edition and NebulaGraph is a separate Chinese-origin project. Structurally, HugeGraph offers both Gremlin and Cypher on one engine and lets you run the same server on embedded RocksDB for development or on Raft-replicated PD and HStore for a cluster.
Do I need HStore and HugeGraph-PD to use HugeGraph?
No. Standalone mode runs RocksDB embedded in the server process on a single node, which is the intended mode for development and testing, with a stated ceiling under 1TB. HugeGraph-PD and HStore are only needed for distributed deployments, where they run on Raft-replicated nodes for high availability.
What Java version does the current HugeGraph release need?
Java 11 or newer, since that requirement landed with release 1.5.0 in December 2024 and has not been raised. The most recent release published is 1.7.0, on 2025-11-16, which also fixed an out of memory fault on dynamic paths with parameters and a NaN in HStore histogram metrics.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/apache-hugegraph)