JanusGraph: a distributed graph database that keeps storage and indexing separate
JanusGraph: an open-source, distributed graph database
At a glance
- What is it?
- JanusGraph is an open-source, distributed graph database built on Apache TinkerPop, designed for graphs with billions of vertices and edges. Its design trades a single-binary install for a stack you assemble yourself.
- Who is it for?
- Adopt JanusGraph when your graph outgrows a single machine and you already operate Cassandra, HBase or Bigtable, because the project treats those systems as its storage layer rather than hiding them. Skip it for a departmental graph where an embedded or single-server database would suffice, since you would be paying for a cluster you do not need.
- Can I use it commercially?
- Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
- Is it still maintained?
- Yes. The repository last received commits 4 days ago.
- What is it written in?
- Mainly Java, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 27, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What JanusGraph is for, and who ends up running it
The README describes JanusGraph as a graph database "optimized for storing and querying large graphs with billions of vertices and edges distributed across a multi-machine cluster." That sentence is the whole pitch, and it also defines the audience. This is not a database you drop into a small service to store a few thousand relationships. It is infrastructure for teams whose graph has already outgrown one machine, or who know it will.
The project is a transactional database that the README says can support thousands of concurrent users, complex traversals, and analytic graph queries. Those three workloads pull in different directions, and JanusGraph answers them by delegating: storage goes to a wide-column store, indexing goes to a search engine, and the query surface stays Apache TinkerPop's Gremlin. The repository layout confirms this. There are separate modules for Cassandra (janusgraph-cql), HBase (janusgraph-hbase), Bigtable (janusgraph-bigtable), Scylla (janusgraph-scylla), Elasticsearch (janusgraph-es), Solr (janusgraph-solr) and Lucene (janusgraph-lucene), alongside janusgraph-core and janusgraph-server.
That module list is the honest description of the product. JanusGraph is a coordination and query layer over systems you already have to run. If you do not want to run those systems, the project is the wrong shape for you, no matter how well it scales.
Storage, indexing and Gremlin: how the pieces fit together
The architecture separates three concerns that many graph databases fuse into one process.
Storage is pluggable. JanusGraph persists vertices and edges in a backend such as Cassandra, HBase, Bigtable or Scylla, all of which are distributed stores with their own replication and failure behaviour. JanusGraph does not replace that behaviour; it inherits it. The consequence is that your graph's durability and availability are only as good as the backend you selected and the way you configured it.
Indexing is separate from storage. A mixed index is served by Elasticsearch or Solr, which is what allows full-text and predicate queries that a key-value or wide-column store cannot answer on its own. There is also janusgraph-lucene, which the repository shows as a module, and janusgraph-inmemory for non-persistent use. The split matters operationally: a mixed index is a second system to size, monitor and keep in sync with the graph.
Queries go through Gremlin, the traversal language of Apache TinkerPop. JanusGraph implements the TinkerPop stack rather than inventing a query language, so existing Gremlin clients and tooling are the entry point. The README also points to a separate janusgraph-visualizer repository and lists third-party visualizers including Cytoscape, Gephi, Graphexp, Graphlytic and G.V().
One more module is worth naming: janusgraph-cdc, which appears in the repository layout. Change data capture is a real operational need when a graph feeds downstream systems, and its presence as a module signals that the project treats it as part of the product rather than an afterthought.
Installing JanusGraph and running a first traversal
The README does not contain installation steps. It says the project homepage "contains more information on JanusGraph and provides links to documentation, getting-started guides and release downloads," and it links to the GitHub releases page for downloads and to Docker Hub for the janusgraph/janusgraph image. Those are the two documented distribution channels: a release archive and a container image.
Because the README gives no commands, the honest way to start is from the published artifacts rather than from invented flags. The container image is the shortest path to a running server, and it is the one piece of the stack you can pull without deciding on a backend yet.
docker pull janusgraph/janusgraphThe image name is the one the README's Docker badge points at. What you get inside it, and which backend it is configured against by default, is something the README does not state, so check the documentation before assuming a working graph.
The second documented channel is the release archive. The release list shows v1.1.0 published on 2024-11-09, with v1.0.1 the day before and v1.0.0 on 2023-10-22. Downloading from the releases page gives you the distribution, and the repository's BUILDING.md and TESTING.md are the files that describe building from source rather than installing a release. The README does not give a download command, so take the link from the releases page itself rather than editing a URL by hand.
Once a server is running, the query surface is Gremlin, and any TinkerPop-compatible client can connect. The README does not show a connection snippet, so the correct next step is the getting-started guide on the project homepage, not a guess at a port or configuration key.
Where JanusGraph is the wrong tool
The clearest limitation is operational cost. JanusGraph is a layer, not a complete database. To run it in the configuration its README describes, you need a distributed storage backend and, for mixed-index queries, a search engine. That is at minimum two additional distributed systems to operate, and the project does not remove that burden; it depends on it.
A second limitation follows from the first. Because storage and indexing are external, consistency between the graph and its mixed index is a property you have to reason about. A query that relies on a mixed index and a query that relies on the storage backend can behave differently, and the README does not discuss this at all.
Third, the README makes no claims about single-node simplicity. If your graph fits comfortably on one machine, the entire distributed design is overhead: more moving parts, more failure modes, more configuration. An embedded graph database or a single-server product will be easier to run and easier to back up, and JanusGraph will not reward you for the extra complexity.
Finally, the documentation surface is split. The README points outward to the homepage for getting-started guides, and the repository carries its own docs directory and mkdocs.yml with a pinned requirements.txt for the documentation toolchain. That is a normal arrangement for a project of this age, but it means the README alone is not enough to operate JanusGraph, and anyone evaluating it should budget time for the documentation site.
JanusGraph compared with a single-server graph database
The comparison people search for most is JanusGraph against Neo4j, and the difference is structural rather than a matter of tuning.
Neo4j is a graph database with its own storage engine and its own query language, Cypher. You install one product and you have a database. JanusGraph has no storage engine of its own in the default deployment; it writes to Cassandra, HBase, Bigtable or Scylla, and it answers Gremlin traversals through the Apache TinkerPop stack. The query language difference is not cosmetic. Cypher and Gremlin express traversals differently, and moving between them means rewriting queries rather than changing a connection string.
The trade is explicit. JanusGraph buys horizontal scale by delegating storage to systems built for horizontal scale, and it pays for that with operational surface area. A single-server graph database buys a simpler deployment and pays for it with the ceiling of one machine.
That framing also answers the other comparisons that show up in search data, against ArangoDB, Dgraph, Memgraph, NebulaGraph, Neptune and Apache AGE. They are all graph databases, but they differ on the same axis: which of them own their storage, which own their query language, and which expect you to bring a cluster. JanusGraph is firmly in the bring-your-own-backend camp, and that is the question to answer first.
Maintenance, releases and licensing
The repository is not archived, and the last push was on 2026-09-22. Release cadence is visible in the release list: v1.0.0 on 2023-10-22, v1.0.1 and v1.1.0 both in November 2024. That is roughly a yearly major line with patch releases, which is a slower cadence than projects that ship continuously, and it is worth knowing before you plan an upgrade schedule.
Upgrade cost is dominated by the ecosystem rather than by JanusGraph itself. Because the project tracks Apache TinkerPop and depends on a storage backend and a search engine, a version bump can require coordinated changes across all three. The repository includes .backportrc.json, which indicates a structured backport workflow, and RELEASING.md, which documents how releases are produced. Neither of those reduces the work of upgrading a running cluster.
On licensing, the repository root contains LICENSE.txt, APACHE-2.0.txt, CC-BY-4.0.txt and NOTICE.txt, and the GitHub metadata reports the licence as NOASSERTION, meaning GitHub could not classify it automatically from the repository contents. The presence of both an Apache 2.0 text and a Creative Commons text suggests different parts of the repository are under different terms, which is common when documentation and code are licensed separately. Read LICENSE.txt and NOTICE.txt yourself before redistributing or embedding anything; this is a description of what the repository contains, not legal advice.
Governance is visible too. The mailing lists are hosted on lists.lfaidata.foundation, and CONTRIBUTING.md covers CLAs and contribution practices, which places the project under LF AI & Data rather than a single vendor.
Editorial conclusion
Adopt JanusGraph when your graph outgrows a single machine and you already operate Cassandra, HBase or Bigtable, because the project treats those systems as its storage layer rather than hiding them. Skip it for a departmental graph where an embedded or single-server database would suffice, since you would be paying for a cluster you do not need. Before committing, verify on your own hardware which storage backend and index backend combination you will run, and confirm that the version you install matches the backend adapters present in the repository layout.
Frequently asked questions
What is JanusGraph?
JanusGraph is an open-source, distributed graph database optimized for storing and querying large graphs with billions of vertices and edges across a multi-machine cluster. It is transactional and queries run through Apache TinkerPop's Gremlin.
Is JanusGraph open source?
Yes. The source is on GitHub under the JanusGraph organization, the repository carries LICENSE.txt, APACHE-2.0.txt and CC-BY-4.0.txt, and the project's mailing lists are hosted on lists.lfaidata.foundation.
How does JanusGraph differ from Neo4j?
Neo4j is a graph database with its own storage engine and the Cypher query language, while JanusGraph persists data in a separate backend such as Cassandra, HBase or Bigtable and is queried with Gremlin. The difference is architectural, so moving between them means rewriting queries, not just reconnecting.
How does JanusGraph compare with ArangoDB?
Both are graph databases, but the repository shows JanusGraph delegating storage to Cassandra, HBase, Bigtable or Scylla and indexing to Elasticsearch, Solr or Lucene, rather than owning a single storage engine. The practical question is whether you want to operate those backends yourself.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/janusgraph-janusgraph)