Framework
awslabs/graphstorm avatar
awslabs/graphstorm

GraphStorm: single-command GNN training on billion-scale graphs

Enterprise graph machine learning framework for billion-scale graphs for ML scientists and data scientists.

454 stars76 forksPythonApache-2.0

At a glance

What is it?
GraphStorm is an Apache-2.0 graph machine learning framework from AWS Labs that trains built-in GNN models with one command and scales to distributed clusters. It is strongest when your graph is large, your schema is heterogeneous, and you can live inside the DGL and PyTorch stack it pins.
Who is it for?
Adopt GraphStorm if you have a heterogeneous graph with billions of nodes or edges and want a built-in RGCN-style baseline without writing model code, and if pinning DGL 2.3.0 and torch 2.3.0 is acceptable. Do not adopt it for a few thousand nodes, for a homogeneous graph where a plain PyG model is enough, or if you need a framework that tracks the newest PyTorch release.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 92 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What GraphStorm is for, and who it is actually aimed at

GraphStorm targets a specific gap: graph neural networks that fit on one GPU are well served by existing libraries, but graphs with billions of nodes and edges are not. The README frames the project as an "enterprise-grade graph machine learning (GML) framework designed for scalability and ease of use", and the two claims pull in opposite directions. Ease of use shows up as a single command that trains a built-in model with no Python written by the user. Scalability shows up as distributed training and inference across a cluster, with Amazon SageMaker AI recommended as the way to avoid managing that cluster yourself.

The intended reader is an ML scientist or data scientist who already understands node classification, link prediction and embeddings, and who has a graph too large for a workstation. The repository layout reflects that audience: training_scripts/ holds YAML configs for built-in models, inference_scripts/ holds the inference counterparts, and examples/ covers concrete domains such as ACM citation data, MAG, network traffic prediction, temporal graph learning and LLM-GNN combinations. Someone who wants a library to embed inside a larger Python service will find the command-line entry points and the partition file format to be the real interface, not the Python API.

How training actually runs: partitioned graphs, DGL, and a launcher

The mechanism visible in the README is a three-step pipeline. First, a raw dataset is converted into a partitioned DGL graph with a JSON config describing it. For OGB arxiv, tools/partition_graph.py downloads and processes the data and writes /tmp/ogbn_arxiv_nc_1p/ogbn-arxiv.json. Second, a training entry point consumes that JSON plus a YAML config file and writes a model directory. Third, an inference pass restores a saved checkpoint and writes predictions or embeddings.

The training commands run through graphstorm.run.gs_node_classification, gs_link_prediction and gs_gen_node_embedding. Each takes --num-trainers and --num-servers, which is the clearest signal that the launcher is built for a distributed setup even when both are set to 1 locally. The README states that users can supply their own model implementations and use the GraphStorm training pipeline to scale them, so the built-in model collection is a starting point rather than the ceiling.

The dependency chain is the constraint that shapes everything else. GraphStorm sits on DGL for graph sampling and on PyTorch for the model and optimizer, and the install instructions pin both: torch 2.3.0 and dgl 2.3.0, with separate wheel indexes for CPU and CUDA 12.1. That pinning is what makes distributed runs reproducible, and it is also what will make an upgrade painful when you want a newer PyTorch.

Installing GraphStorm locally and training a first model

The README gives pip as the local install path and lists the compatibility floor: Python 3.8+, PyTorch 1.13+, DGL 1.0+ and transformers 4.3.0+. The commands below are the CPU variant exactly as the README presents it. The first installs PyTorch from the CPU wheel index, the second installs DGL from the DGL wheel index built against torch 2.3, and the third installs the graphstorm package itself.

bash
pip install torch==2.3.0 --index-url https://download.pytorch.org/whl/cpu
pip install dgl==2.3.0 -f https://data.dgl.ai/wheels/torch-2.3/repo.html
pip install graphstorm

For a GPU machine the README substitutes the CUDA 12.1 wheels: torch==2.3.0 from the cu121 index and dgl==2.3.0+cu121 from the matching DGL index. Note that setup.py also pins transformers==4.48.0, torchdata==0.9.0 and ogb==1.3.6 as install requirements, while requirements.txt lists looser floors (transformers 4.3.0, dgl >= 1.0). The two files disagree, and pip will follow setup.py.

The quick start then clones the repository, because the example configs live in it rather than in the wheel. Partition the OGB arxiv graph into a single part first.

bash
python tools/partition_graph.py \
    --dataset ogbn-arxiv \
    --filepath /tmp/ogbn-arxiv-nc/ \
    --num-parts 1 \
    --output /tmp/ogbn_arxiv_nc_1p

That writes a partitioned graph and a JSON config. Then train an RGCN model for node classification, pointing --cf at the YAML shipped in the repository and --part-config at the JSON just produced.

bash
mkdir /tmp/ogbn-arxiv-nc
python -m graphstorm.run.gs_node_classification \
    --workspace /tmp/ogbn-arxiv-nc \
    --num-trainers 1 \
    --num-servers 1 \
    --part-config /tmp/ogbn_arxiv_nc_1p/ogbn-arxiv.json \
    --cf "$(pwd)/training_scripts/gsgnn_np/arxiv_nc.yaml" \
    --save-model-path /tmp/ogbn-arxiv-nc/models

The reader should see training progress in the terminal and a models directory containing epoch checkpoints. Inference is the same entry point with --inference, --restore-model-path pointing at a checkpoint such as /tmp/ogbn-arxiv-nc/models/epoch-7/, and --save-prediction-path for the output. The link prediction example follows the same shape with tools/partition_graph_lp.py and gs_link_prediction, and swaps the final step for gs_gen_node_embedding to write node embeddings instead of class predictions.

The dependency pins are the real adoption cost

The most consequential thing in this repository is not a model. It is the version matrix. setup.py pins transformers==4.48.0, torchdata==0.9.0 and ogb==1.3.6, and the README's install block pins torch 2.3.0 and dgl 2.3.0. DGL wheels are built per PyTorch version and per CUDA version, so moving to a newer torch means waiting for a matching DGL wheel at the exact URL pattern the README uses. If your environment already runs a different PyTorch, you are not adding GraphStorm to it; you are building a separate environment around GraphStorm.

This is a deliberate trade. Reproducible distributed training across many machines is hard, and pinning is how the project keeps it predictable. The cost lands on teams that treat their ML environment as a shared platform. A second, quieter issue is that requirements.txt and setup.py do not agree on transformers and DGL versions. Anyone who installs from requirements.txt first, then runs pip install graphstorm, will get whatever pip resolves and may not match the tested combination.

A third limit is scope. The README describes built-in models and a training pipeline for custom ones, and the examples cover node classification, link prediction, embeddings and temporal graph learning. Nothing in the README describes a serving or online inference tier, and nothing describes rollback or checkpoint compatibility across releases. If you need to know whether a model trained on 0.4.2 loads under 0.5.1, the release notes are where that answer would be, and the README does not carry it.

GraphStorm against PyTorch Geometric and plain DGL

The closest alternative for most readers is PyTorch Geometric, and the difference is architectural rather than cosmetic. PyG is a library you import: you write the model, the training loop and the data loader in Python, and scaling is something you add later. GraphStorm is a pipeline you invoke: the partitioning step, the launcher, the YAML config and the checkpoint format are all part of the contract, and the built-in models are configured rather than coded. If your graph fits in memory on one machine, PyG gives you more freedom per line of code. If your graph does not fit, GraphStorm has already solved partitioning and multi-trainer coordination, and PyG has not.

Plain DGL is the other comparison, and it is closer than it looks, because GraphStorm is built on DGL. Using DGL directly means you keep full control of sampling and model code but you own the distributed orchestration. GraphStorm's value is the layer above DGL: the partition tooling, the run.gs_* entry points, the config schema and the SageMaker integration path. The README also notes AWS integration out of the box, which matters if your graph data already lives in AWS services. That is a real difference in approach, not a marketing line, and it is also the reason the framework feels opinionated about infrastructure.

Maintenance status, licensing, and what to check before you commit

The repository is not archived, and the last push was on 2026-06-30, which is under three months before today's date. Releases are regular: 0.5.1 on 2025-12-31, 0.5 on 2025-09-29 and 0.4.2 on 2025-06-17. That cadence is roughly quarterly, so plan upgrades around release notes rather than continuous tracking. There is no documented upgrade path between minor versions in the README, and no statement about checkpoint or partition format compatibility, so treat each minor release as something to validate against your own pipeline before rolling it into production.

The licence is Apache-2.0, declared both in the repository metadata and in setup.py. Apache-2.0 permits commercial use and modification and includes a patent grant, and the repository also carries a NOTICE file, which is the standard Apache mechanism for attribution requirements when redistributing. Whether you need to surface that NOTICE in your own product is a question for your legal team, not something this article can settle. The practical point is that the permissive licence removes the biggest blocker to adopting a framework from a cloud vendor's lab, and the dependency pins, not the licence, are what will shape your maintenance budget.

Editorial conclusion

Adopt GraphStorm if you have a heterogeneous graph with billions of nodes or edges and want a built-in RGCN-style baseline without writing model code, and if pinning DGL 2.3.0 and torch 2.3.0 is acceptable. Do not adopt it for a few thousand nodes, for a homogeneous graph where a plain PyG model is enough, or if you need a framework that tracks the newest PyTorch release. Before committing, verify that your graph can be partitioned by tools/partition_graph.py into the JSON config the training entry points expect, and check the PyPI version against the 0.5.1 release note, since setup.py derives the version from python/graphstorm/__init__.py.

Frequently asked questions

What is GraphStorm used for?

GraphStorm is a graph machine learning framework for training and running inference on industry-scale graphs with billions of nodes and edges. It provides built-in GML models that train with a single command, plus a programming interface for custom models run through its distributed pipeline.

How do I install GraphStorm with pip?

The README installs PyTorch 2.3.0 and DGL 2.3.0 from their respective wheel indexes first, choosing the CPU or CUDA 12.1 variant, and then runs pip install graphstorm. The package requires Python 3.8+ and setup.py pins transformers==4.48.0, torchdata==0.9.0 and ogb==1.3.6.

Does GraphStorm support distributed training?

Yes. The training entry points accept --num-trainers and --num-servers, and the README recommends Amazon SageMaker AI for distributed runs to avoid managing cluster infrastructure yourself.

Which models does GraphStorm include?

The README refers to a built-in model collection and demonstrates an RGCN model for node classification and for link prediction on the OGB arxiv graph. Users can also supply their own model implementations and train them through the GraphStorm pipeline.

What are the system requirements for GraphStorm?

GraphStorm is compatible with Python 3.8+ and requires PyTorch 1.13+, DGL 1.0+ and transformers 4.3.0+, though the README's install commands pin torch 2.3.0 and dgl 2.3.0. The repository also ships a docker/ directory and SageMaker integration files.

Official sources

  1. awslabs/graphstorm on GitHub
  2. Issues
  3. License: Apache-2.0
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/awslabs-graphstorm.svg)](https://hysenlabs.com/projects/awslabs-graphstorm)