Self-hosted service
alibaba/GraphScope avatar
alibaba/GraphScope

GraphScope: three graph engines behind one Python front door

🔨 🍇 💻 🚀 GraphScope: A One-Stop Large-Scale Graph Computing System from Alibaba | 一站式图计算系统

3,558 stars469 forksC++Apache-2.0

At a glance

What is it?
Alibaba's GraphScope bundles an analytics engine, an interactive engine and a graph learning framework into a single distributed system. The interesting question is whether the Python interface hides the seams well enough to be worth it.
Who is it for?
GraphScope is a serious piece of infrastructure built by people who had a graph problem at Alibaba scale, and the repository is honest about which parts came from where. GRAPE does the analytics, MaxGraph does interactive query, Graph-Learn does GNN training, and the coordinator plus Python client exist to make those three look like one system from a notebook.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 14 days ago.
What is it written in?
Mainly C++, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 7, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What is actually bundled in this repository

GraphScope describes itself as a unified distributed graph computing platform that provides a one-stop environment for graph operations across a cluster, through a Python interface. The important word is bundled. The README is explicit that the platform combines several pieces of Alibaba technology: GRAPE for analytics, MaxGraph in `interactive_engine/` for interactive traversal, Graph-Learn for graph neural network work, and the Vineyard project for efficient in-memory data transfer.

So this is not a single novel engine. It is an integration of three existing systems plus the plumbing that makes them look like one API. That framing matters more than it might seem, because it predicts the shape of the code. The top-level tree is a directory per concern:

code
analytical_engine/
interactive_engine/
learning_engine/
coordinator/
flex/
python/
k8s/
proto/

`coordinator/` and `python/` are the parts that make the bundle feel unified, `k8s/` and `charts/` are deployment, `proto/` is the wire contract between client and server, and `flex/` is the newer, separately versioned evolution. There is also a `gsctl.py` at the root and a `V6D_VERSION` file, which is a leftover from the Vineyard era and a small reminder that this repository has been assembled over a long period.

The project is Apache-2.0 licensed, written mostly in C++, and carries about 3,500 stars and 460 open issues. The last recorded push on the default branch is 2026-09-23.

Installing it and what you are actually installing

The standalone path is one pip command, and the README is unusually direct about the constraints around it:

bash
pip3 install graphscope

Python 3.8 or newer and pip 19.3 or newer. The precompiled binaries are built for the common Linux distributions (Ubuntu 20.04 and up, CentOS 7 and up) and macOS 12 and up on both Intel and Apple silicon. Windows is not on that list, and the recommendation is to run Ubuntu under WSL2 rather than trying to build natively.

That platform list is the single most useful thing in the installation section, because the alternative is discovering the limitation halfway through a build. Note what the pip package is and is not: it is not a source distribution that compiles on install. Behind it sit C++ engines that need a toolchain, and the fact that they ship prebuilt is what makes the one-line install honest rather than optimistic.

There is also a hosted Playground, a managed JupyterLab at try.graphscope.io, which is the fastest way to find out whether the API suits your problem before committing a cluster. A Colab badge appears in the README for the same reason.

The build reveals how much is really here

The Makefile is short and its structure is the clearest statement of the architecture in the repository. It defines one variable per engine directory, then a top-level target that builds all four things in dependency order:

code
all: coordinator analytical interactive gsctl
graphscope: all

The comment above that target explains the ordering: the coordinator relies on the client, which relies on learning. So the build graph runs bottom-up through the Python client, the GNN framework, the coordinator, and finally the two query engines. `install` then assembles the same set into a prefix, defaulting to `/opt/graphscope` when `INSTALL_PREFIX` is not set in the environment.

The build switches are worth a look because they reveal what is optional. `NETWORKX ?= ON` for the analytical engine, `BUILD_TEST ?= OFF` for testing support, and `WITH_GLTORCH ?= ON` for the graphlearn-torch extension, with a comment noting graphlearn itself is built by default. Turning the PyTorch extension off is the kind of choice you only need to make if you are not doing model training, and it is the clearest signal in the build that GNN work is a first-class part of the project rather than an afterthought.

The `clean` target is a small indicator of age and scope: it removes the analytical build directory and generated protobufs, then shells out to Maven for the Java components under both `analytical_engine/java` and `interactive_engine`. A C++ repository that has to run `mvn clean` is a repository with JVM services in it.

Flex is where the project is pointing

The README carries a prominent banner for GraphScope Flex, described as a LEGO-inspired, modular and user-friendly evolution, with its own directory in the tree. That framing matters, because it tells you what the maintainers think of the original design: modular and composable, in the way the current tree is not.

The evidence for how seriously Flex is taken runs through the README's news section. Flex shipped a tech preview with v0.23.0 in July 2023, a paper on arXiv that December, and acceptance at SIGMOD 2024's industry track in February 2025. In May 2024 it set new records on the LDBC Social Network Benchmark interactive audit, and in April 2025 GraphScope posted a result on the SNB interactive workload with Cypher as the query language, claiming 2.0 times the throughput of the previous record holder on SF300.

Benchmark records are the currency this project competes in, and the GraphAr file format donated to the Apache Software Foundation as an incubating project in March 2024 is a different kind of contribution: an attempt to standardize how graph data is serialized at scale, which is a problem every system here has to solve.

The releases tell the quieter story. v0.31.0 in January 2025 focused on Flex Interactive: string column support, a refactored interactive runtime, a new `sharding_mode` option defaulting to exclusive so admin requests are not blocked by long queries, an optimization for Scan plus Limit patterns, and multiple properties on edge triplets. v0.30.0 earlier in January was mostly Portal work on graph visualization, with a new exploration module and four built-in layouts. That is a project shipping focused improvements rather than a new architecture every quarter.

Running it on Kubernetes

Standalone mode is the quick start, but the platform target is Kubernetes, and the repository carries the pieces for it: a `k8s/` directory, Helm charts under `charts/`, and a coordinator component whose whole job is to turn that into a multi-user service. The README's own table of contents splits the walkthrough into creating a session, loading graphs and running computation, and closing the session.

Session-oriented allocation is the right abstraction for this kind of system and also its main operational cost. A graph that does not fit in memory has to live in persistent storage, and the coordinator is what decides where. The releases name that layer directly: v0.29.0 in September 2024 covered the Groot persistent storage alongside Flex, the `gsctl` command line utility and the graph interactive engine.

`gsctl` deserves its own paragraph because it shows how the project expects day two to look. It is a Python-installed command line utility that runs in two modes, utility scripts by default and client or server mode after `gsctl connect`, and it covers building images and packages and managing sessions and resources. A tool whose job is session and resource management is a tool for people running the system as shared infrastructure, not a notebook user.

Vineyard is the quiet piece that makes the cluster mode tolerable. Because it handles in-memory data transfer between stages, a computation that would otherwise serialize a large graph between two operators can keep it in memory, which is the difference between a graph pipeline that is usable and one that spends its time moving bytes.

Who this is for, and who should look elsewhere

The README's worked example is node classification on the `ogbn-mag` citation network, a heterogeneous graph of papers, authors, institutions and fields of study with four types of directed relation between them. Each paper node carries a 128-dimensional word2vec vector as its attribute. The walkthrough then goes through loading the graph, running an interactive query, running graph analytics, and training a GNN.

That sequence is the product pitch in miniature, and it is worth noticing that it covers all three engines in one script. If your task only needs one of them, the bundled platform is heavier than the underlying project would be. Graph analytics without interactive traversal or model training is a job for GRAPE alone, and a graph that fits comfortably in one machine does not need a coordinator.

What GraphScope genuinely buys you is the pipeline. When the same graph has to be loaded once, traversed interactively while a human looks at it, then handed to a trainer, the cost of making three systems share one copy of the data is exactly the problem Vineyard and the coordinator exist to solve. Building that yourself is a project. This is the packaged version of that project, from a team that had to build it.

The honest caveats are the release cadence and the documentation boundary. Releases land a few times a year rather than weekly, the newest one recorded here is from January 2025, and the substantive reference material lives on graphscope.io rather than in the repository. The docs site, the FAQ pages in both English and Chinese, and the Portal developer documentation are where the details are.

Editorial conclusion

GraphScope is a serious piece of infrastructure built by people who had a graph problem at Alibaba scale, and the repository is honest about which parts came from where. GRAPE does the analytics, MaxGraph does interactive query, Graph-Learn does GNN training, and the coordinator plus Python client exist to make those three look like one system from a notebook. The parts to weigh carefully are the release cadence, which has slowed to a few times a year, and the direction of travel, since the README now steers new work toward GraphScope Flex rather than the original tree. For a cluster-scale problem with analytics, interactive traversal and model training in the same pipeline, the bundled stack removes a lot of glue. For a single-machine exploration, a graph library in pure Python is a smaller commitment.

Frequently asked questions

What is graph software used for?

In this project's case, for the three things a large graph turns out to need: computing statistics over it with an analytics engine, querying it interactively while someone is looking at the result, and training a neural network over it. GraphScope exists to make those three stages run against one copy of the data rather than three separate copies.

Is GraphScope a graph database?

Not in the sense of a transactional key-value store you query transactionally. It is a graph computing platform: an analytics engine built on GRAPE, an interactive engine built on MaxGraph, a GNN framework from Graph-Learn, and persistent storage, unified behind a Python interface and a coordinator that manages sessions on a cluster.

What does GraphScope Flex change compared to the original?

Flex is described as a LEGO-inspired, modular evolution with its own directory in the tree, and recent releases have concentrated there. v0.31.0 added string column support, a refactored interactive runtime, a `sharding_mode` option defaulting to exclusive, a Scan plus Limit query optimization and multiple properties on edge triplets.

Official sources

  1. alibaba/GraphScope on GitHub
  2. License: Apache-2.0
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/alibaba-graphscope.svg)](https://hysenlabs.com/projects/alibaba-graphscope)