Library / SDK
tensorflow/gnn avatar
tensorflow/gnn

TensorFlow GNN: building heterogeneous graph neural networks on TensorFlow

TensorFlow GNN is a library to build Graph Neural Networks on the TensorFlow platform.

1,546 stars205 forksPythonApache-2.0

At a glance

What is it?
TensorFlow GNN (TF-GNN) gives TensorFlow users a graph tensor type, a Beam-based sampler and a runner API for training. It is a Google OSS port, and its Keras v2 constraint is the first thing to check before you commit.
Who is it for?
Adopt TensorFlow GNN if you already train on TensorFlow, your graphs have more than one node or edge type, and you need to sample subgraphs from a graph too large to fit in memory. Do not adopt it if you are on Keras v3 with no intention of setting TF_USE_LEGACY_KERAS=1, or if you want a framework-agnostic graph library, because the whole design is bound to TensorFlow tensors and Keras v2 layers.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository received new commits within the last day.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

DEEP OPEN-SOURCE ANALYSIS

What TF-GNN solves, and for whom

Most graph neural network code assumes one node type and one edge type. Real production graphs rarely look like that. A recommendation graph has users, items and the interactions between them. A citation graph has papers, authors and venues. Representing those as a single adjacency matrix means flattening the distinctions that carry most of the signal.

TensorFlow GNN addresses that directly. The README describes it as a library to build Graph Neural Networks on the TensorFlow platform, and the first thing it lists is a tfgnn.GraphTensor type that represents graphs with a heterogeneous schema, meaning multiple types of nodes and edges. The intended reader is an engineer who already works in TensorFlow and has a graph problem that does not fit a fixed-shape tensor pipeline.

The second problem it addresses is scale. The README lists a graph sampler that converts a huge database into a stream of reasonably-sized subgraphs for training and inference. That is the part that matters operationally: you rarely train on the whole graph, you train on sampled neighborhoods, and the library ships that sampling step rather than leaving it to you.

The library is described in the README as an OSS port of a Google-internal library used in a broad variety of contexts, on homogeneous and heterogeneous graphs, and in conjunction with other scalable graph mining tools. That provenance explains the shape of the API. It is not a research prototype; it is a library built to run inside a production training stack, and the abstractions reflect that.

GraphTensor, schema and the sampler: how the pieces fit

The core abstraction is the GraphTensor, a single tensor-like value that holds several node sets and edge sets plus the connectivity between them. Because it is a tensor and not a Python graph object, it can be batched, placed on an accelerator and passed through Keras layers the same way an image batch is. A schema, described in the schema guide, declares which node sets and edge sets exist and what features each carries. That declaration is what lets the library keep the type distinctions instead of collapsing them.

On top of that sit two other layers. One is a collection of ready-to-use models and Keras layers, so you can assemble a GNN from existing building blocks or write your own with the same primitives. The other is a high-level API for training orchestration, documented in the runner guide, which handles the training loop rather than the graph math.

Data flow runs in one direction through four stages. Raw records go through data preparation. The Beam sampler turns a large database into a stream of subgraphs. Those subgraphs are parsed into GraphTensors against the schema. The GraphTensor then flows through Keras layers, either the provided models or your own, and out through the runner. The sampler is the interesting design choice: it is built on Apache Beam, which means distributed sampling is a first-class concern rather than something bolted on after the model works on a small graph. The cost is that Apache Beam is a dependency you carry even if you never run sampling in distributed mode.

Installing TensorFlow GNN and running a first model

The README gives the stable release install as a single pip command. Run it in the environment where TensorFlow is already installed, and the package will pull its own dependencies, which the pyproject.toml lists as ml-collections, networkx, pyarrow, tensorflow, apache-beam, crcmod, tensorboard and tf-keras.

bash
pip install tensorflow-gnn

If you are on TensorFlow 2.16 or newer, that install alone is not enough. The README states that TF-GNN does not work with the new multi-backend Keras v3, and that users of TF2.16+ must also install tf-keras and set an environment variable. The Keras version guide in the repository covers the details.

bash
pip install tf-keras
export TF_USE_LEGACY_KERAS=1

Set that variable before importing TensorFlow, not after. If you skip it, the failure shows up as a Keras import or layer-construction error rather than a clear version message.

The README's own quickstart avoids installation entirely: it points at three Google Colab notebooks, one doing molecular graph classification with the MUTAG dataset, one training end to end on the OGBN-MAG benchmark with heterogeneous sampled subgraphs, and one learning shortest paths with an Encoder/Process/Decoder architecture. Running the OGBN-MAG notebook is the fastest way to see a real heterogeneous graph flow from sampling through training, and it is the example to copy from when you build your own pipeline. The repository also carries a Dockerfile that installs tensorflow-gnn==1.0.0 alongside httplib2, notebook and ogb, which gives you a container to experiment in if you would rather not touch your local environment.

On platform support, the README is explicit that TF-GNN is developed and tested on Linux, and that running on other platforms supported by TensorFlow may be possible. The pyproject classifiers do list macOS and Windows, but the README sentence is the one to trust when you are deciding where to run training.

The Keras v2 constraint is the real adoption cost

The single largest limitation is stated plainly in the README and is easy to skim past: TF-GNN does not work with Keras v3. Every team that moved to TensorFlow 2.16 or later and adopted the multi-backend Keras is, by default, on an incompatible stack. The workaround is to install tf-keras and set TF_USE_LEGACY_KERAS=1, which pins you to the legacy Keras behavior for the whole process. That is not a small side effect. If any other part of your training code depends on Keras v3 semantics, you now have two Keras generations in one process, and you have to reason about which layers see which.

There is a second constraint in the version bounds. The pyproject.toml requires tensorflow>=2.17.0, <3 and apache-beam>=2.57, <3, and the README notes that future releases will raise the required TF version. Those bounds mean TF-GNN tracks TensorFlow releases rather than insulating you from them. An upgrade of TensorFlow in your environment can force an upgrade of TF-GNN, and the README does not document a rollback path for that.

A third point is worth being honest about: the latest stable release listed is v1.0.3 from 2024-05-13, while the last push to the repository was on 2026-09-09. Development activity and release cadence are different things here, and the release history gives you no signal about how quickly a fix lands. If you are evaluating this for a long-lived system, look at the repository activity rather than the release tags, and note that the README does not describe a support window or a deprecation policy.

When TensorFlow GNN is the wrong tool

If your graph is homogeneous and fits comfortably in memory, TF-GNN is more machinery than the problem needs. The GraphTensor and schema layers exist to preserve node and edge type distinctions; with one type of each, you are paying the abstraction cost for nothing. A plain adjacency-matrix implementation in TensorFlow, or a small custom Keras model, will be shorter and easier to debug.

If you are not committed to TensorFlow, this is the wrong library. Every abstraction is a TensorFlow tensor and every model is a Keras v2 layer. There is no export path to another runtime described in the README, and the Keras v2 requirement means you cannot simply lift the layers into a Keras v3 project later. Teams that expect to move between frameworks should pick a framework-agnostic graph library instead.

If your bottleneck is graph analytics rather than learned representation, TF-GNN does not help. The README mentions that the internal library is used in conjunction with other scalable graph mining tools, which is a hint that the mining and the learning are separate concerns. Do not reach for a GNN because you want to compute centrality or community structure; those are graph algorithms, not neural networks.

Finally, the sampler is Beam-based. If your organization has no Beam infrastructure and no intention of running distributed sampling, you are adding a dependency and a programming model to your stack for a component you might replace with a simpler in-process sampler. That trade is defensible for large graphs and wasteful for small ones.

Alternatives and how the approach differs

The most direct comparison is with general-purpose graph learning libraries that are not tied to a single deep learning framework. Those libraries typically define their own graph object and their own message-passing interface, then hand tensors to whichever framework you choose. The difference in approach is where the type system lives. In TF-GNN the graph type is a TensorFlow tensor and the schema is part of the model's input contract, so batching, device placement and Keras integration come for free. In a framework-agnostic library the graph object is the library's own, and the framework boundary is a conversion step you manage.

A second alternative is to stay inside TensorFlow and write the message passing yourself. That is entirely feasible for a homogeneous graph, and it avoids both the Keras v2 constraint and the Beam dependency. What you give up is the sampler, the schema handling and the ready-to-use models. For a team that needs one specific GNN architecture on one graph shape, hand-rolling is often the shorter path.

A third option is a graph database or graph analytics engine with its own embedding support. Those systems keep the graph in a queryable store and compute embeddings close to the data. TF-GNN takes the opposite stance: it exports subgraphs out of your store and trains in TensorFlow. Which is right depends on whether your graph changes faster than your model retrains. If the graph is stable and the model is the hard part, TF-GNN's split makes sense. If the graph is constantly mutating and you need fresh embeddings on demand, pulling subgraphs into a training job is the wrong direction.

Licence, maintenance and what an upgrade costs

TF-GNN is licensed under Apache-2.0, stated both in the LICENSE file and in the pyproject.toml license field. That is a permissive licence, and it is the same licence TensorFlow itself uses, so there is no new category of obligation introduced by adding this dependency. The package metadata lists Google LLC as the author with a [email protected] contact. This is a description of the licence terms, not legal advice; if your organization has rules about dependency licences, run it past whoever handles that.

On maintenance, the last push to the repository was on 2026-09-09, and the repository is not archived. The most recent release listed is v1.0.3 from 2024-05-13, preceded by v1.0.3rc0 and v1.0.2. The gap between the last push and the last release is the number that matters for planning. Code is moving; tagged releases are not keeping pace. If you need a pinned version with a documented changelog for each change, that cadence will frustrate you.

The upgrade cost is dominated by the TensorFlow version bound. Because pyproject.toml requires tensorflow>=2.17.0, <3 and the README warns that future releases will raise the required TF version, upgrading TF-GNN can force a TensorFlow upgrade, and a TensorFlow upgrade can force the tf-keras workaround if you are not already on it. Budget for that as a coupled upgrade rather than a routine dependency bump. There is no rollback procedure in the README, so plan to pin both packages together in your lockfile and test the pair before rolling out.

Editorial conclusion

Adopt TensorFlow GNN if you already train on TensorFlow, your graphs have more than one node or edge type, and you need to sample subgraphs from a graph too large to fit in memory. Do not adopt it if you are on Keras v3 with no intention of setting TF_USE_LEGACY_KERAS=1, or if you want a framework-agnostic graph library, because the whole design is bound to TensorFlow tensors and Keras v2 layers. Before writing model code, verify three things on your own machine: that pip install tensorflow-gnn resolves against your TensorFlow version, that the Keras version guide's tf-keras and TF_USE_LEGACY_KERAS=1 steps apply to your setup, and that your graph schema is expressible as a GraphTensor with named node sets and edge sets. If the schema cannot be written that way, the library is the wrong layer and you will spend the integration budget fighting it.

Frequently asked questions

What is a GNN used for?

The README frames TensorFlow GNN as a way to build graph neural networks on TensorFlow, with a GraphTensor type for graphs that have multiple types of nodes and edges, a Beam-based sampler that turns a huge database into streams of subgraphs, and ready-to-use models and Keras layers. The Colab examples apply it to molecular graph classification, the OGBN-MAG benchmark and shortest-path prediction.

Is GNN better than CNN?

The README does not compare the two. It positions TensorFlow GNN for graph-structured data with heterogeneous schemas, and the examples it links are graph tasks such as molecular classification and node property prediction on OGBN-MAG.

What is TensorFlow used for?

TensorFlow GNN is built on the TensorFlow platform: it requires TensorFlow 2.12 or higher for release 1.0, its pyproject.toml pins tensorflow>=2.17.0, <3, and its models are Keras layers. The README does not describe TensorFlow itself beyond those requirements.

Is GNN a deep learning algorithm?

TensorFlow GNN is described as a library to build Graph Neural Networks on the TensorFlow platform, and it ships ready-to-use models and Keras layers for GNN modeling, with a high-level API for training orchestration. The README does not give a formal definition of GNNs beyond that.

Official sources

  1. Issues
  2. License: Apache-2.0
  3. README
  4. Releases
  5. tensorflow/gnn on GitHub
For maintainers

Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/tensorflow-gnn.svg)](https://hysenlabs.com/projects/tensorflow-gnn)
Community notes

Community notes