Library / SDK
graphistry/pygraphistry avatar
graphistry/pygraphistry

PyGraphistry: GPU-Accelerated Graph Visualization and GFQL Queries from Python

PyGraphistry is a Python library to quickly load, shape, embed, and explore big graphs with the GPU-accelerated Graphistry visual graph analyzer.

2,560 stars229 forksPythonBSD-3-Clause

At a glance

What is it?
PyGraphistry is a BSD-3-Clause Python library that turns pandas, Spark, RAPIDS and Arrow tables into interactive graphs and vectorized GFQL queries. It is a strong fit for notebook-driven graph analysis and a poor fit if you want a self-contained local graph database.
Who is it for?
Adopt PyGraphistry if your graph work already lives in pandas, Spark or Arrow notebooks and you want interactive exploration without standing up a graph database. Do not adopt it if you need an embedded, fully offline graph store, or if you cannot send data to Graphistry Hub or a self-hosted server for the visualization half.
Can I use it commercially?
Yes. BSD-3-Clause is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 5 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 27, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The dataframe-to-graph gap PyGraphistry fills

Most teams that end up with graph-shaped data never start with a graph database. They start with an edge list in a dataframe: a source column, a destination column, and a handful of attributes. Answering a relationship question in that format means writing self-joins, and the moment you want to look at the result you are exporting to Graphviz or Gephi and losing the interaction. PyGraphistry targets exactly that gap. The README describes it as an open source Python library for data scientists and developers to "leverage the power of graph visualization, analytics, AI, including with native GPU acceleration", and the intended audience is stated plainly: data scientists and developers working in notebooks. The library ingests data in many formats and shapes, binds it to a graph object, and then either queries it in place or ships it to a Graphistry visual analyzer for interactive exploration. If your workflow is pandas plus a plotting library plus occasional frustration about layout, this is the layer that replaces the export step.

How the graph object, GFQL and the remote analyzer fit together

The mechanism is a bound graph: you attach node and edge tables to a graphistry instance, and that object becomes the thing you query and plot. Ingestion is dataframe-native rather than file-based, with the README pointing to Pandas, Spark, RAPIDS and Apache Arrow as the supported shapes. On top of that sits GFQL, which the README calls "the first fully vectorized dataframe-native graph query language with an open-source GPU runtime". The important architectural detail is that GFQL is vectorized over columnar data, so a traversal is expressed as operations over Arrow-backed tables rather than as a walk that dereferences pointers one hop at a time. Queries can run against the currently bound graph through g.gfql("MATCH ..."), which accepts Cypher-like syntax, or be dispatched with g.gfql_remote([...]) to a Graphistry server. That split matters: the local path keeps computation next to your data, while the remote path lets a notebook stay thin and a production deployment carry the load. Release 0.58.0 added GFQL seeded fast paths, resident indexes and native Polars OLAP, which suggests the query engine now keeps indexes in memory between calls instead of rebuilding them. The README does not document how large those resident indexes grow, so memory planning is on you. The visualization half is a separate hop: the analyzer renders in a browser, either on Graphistry Hub or on a self-hosted server, and the README positions the workflow as prototyping locally and then powering production dashboards and pipelines remotely.

Installing PyGraphistry and binding a first graph

The package is published on PyPI as graphistry, which is the name you install even though the repository is pygraphistry. The base install pulls numpy, pandas, pyarrow, lark, palettable, requests, packaging and setuptools, so a plain install gives you dataframe ingestion and the GFQL parser without any of the optional graph backends.

bash
pip install graphistry

Optional capabilities are extras rather than default dependencies. The setup.py defines extras including igraph, networkx, gremlin, bolt, nodexl, jupyter, spanner, kusto, polars, umap-learn and cugraph, so a NetworkX round trip needs the extra installed explicitly. The README also points at a separate package for LLM coding assistants, installed with npx, which is a JavaScript toolchain rather than a Python one.

bash
npx skills add graphistry/graphistry-skills

The README claims that this skills package improves AI success rates from roughly 50 percent to roughly 90 percent on PyGraphistry tasks. That figure comes from the project itself and the README does not describe the evaluation behind it, so treat it as a vendor claim rather than a measured result. For a first real use, the README points at the 10-minute tutorial for dataframe-native graph processing, and the documentation states that GFQL accepts a Cypher-like MATCH string through g.gfql. The README does not print a full ingestion snippet, so the exact binding call is something to copy from the tutorial rather than from the README. What you should expect at that point is a graph object rather than a rendered picture. Rendering requires a plot call and, for anything beyond the smallest example, a Graphistry Hub account or a self-hosted server. The README does not document the exact plotting call signature, so check the visualization tutorial in the documentation before assuming the method name.

Where PyGraphistry is the wrong tool

The clearest limitation is that PyGraphistry is not a graph database. GFQL queries run over dataframes you have already loaded into memory, so the practical ceiling is what fits in a pandas or Arrow table on the machine running the notebook, or what Spark or RAPIDS can hold across a cluster. If your question requires multi-hop traversal over a graph larger than memory, or transactional writes with isolation guarantees, a graph database is the correct layer and PyGraphistry is a visualization client for it, which is why the README lists Neo4j, Neptune, TigerGraph, ArangoDB and Memgraph connectors rather than positioning the library as a replacement. The second limitation is the network boundary. The interactive analyzer runs in a browser against Graphistry Hub or a self-hosted server, so the visualization path involves sending graph data to a service. The README does not document an offline rendering mode, and teams in regulated environments will need to confirm that self-hosting is available to them before treating this as an option. Third, the optional extras are genuinely optional, which means a fresh install will not have NetworkX, igraph, Polars or the RAPIDS stack. Code that works in one environment can fail in another purely because an extra was missing, and the failure appears at the import or call site rather than at install time.

How PyGraphistry differs from Gephi, Neo4j and Linkurious

The alternatives people search for alongside PyGraphistry fall into two camps, and the difference is where the graph lives. Gephi is a desktop application: you load a file, lay out the graph, and explore it by hand. It is excellent for one-off investigation and useless inside a pipeline, because there is no Python object to hand back to the next step. PyGraphistry inverts that, keeping the graph in a dataframe you can keep computing on. Neo4j is a graph database with its own storage engine and query language. It persists the graph, enforces schema and handles writes, and it is the right answer when the graph is the system of record. PyGraphistry does not persist anything; it queries dataframes and renders them, and the README treats Neo4j as a source to connect to rather than a competitor. Linkurious is a graph visualization and investigation platform aimed at analysts working over a database, closer in spirit to the Graphistry analyzer than to the Python library. The honest summary is that PyGraphistry occupies a narrow band: dataframe in, vectorized query, browser visualization out. If your problem sits outside that band, one of the other three is likely a better fit, and the README's connector list is effectively an admission of that.

Maintenance, release cadence and the BSD-3-Clause licence

The repository is not archived and the last push was on 2026-07-22, which is recent enough that describing the project as actively developed is defensible on the evidence available. The release history backs that up: 0.58.0 landed on 2026-07-22 with GFQL seeded fast paths, resident indexes and native Polars OLAP, 0.57.0 on 2026-06-29 focused on GFQL performance with parse caching and a single-hop fast path, and 0.56.0 on 2026-05-23. Three releases in roughly two months, all with substantive notes, is a cadence that implies upgrade work rather than a frozen dependency. That is the cost side: pinning to an older minor version means missing the query-engine changes, and the changelog is the only reliable way to know what moved. The licence is BSD-3-Clause, which is permissive and permits commercial use and redistribution provided the copyright notice and disclaimer are retained. The licence covers the Python library. It does not cover Graphistry Hub, which the README describes as a separate hosted service, and it does not automatically cover a self-hosted server deployment. Whether a given deployment needs a commercial agreement is a question for Graphistry, not something the repository answers, and nothing here should be read as legal advice.

Editorial conclusion

Adopt PyGraphistry if your graph work already lives in pandas, Spark or Arrow notebooks and you want interactive exploration without standing up a graph database. Do not adopt it if you need an embedded, fully offline graph store, or if you cannot send data to Graphistry Hub or a self-hosted server for the visualization half. Before committing, verify two things: that your data is allowed to leave your environment, and which optional extras (igraph, networkx, polars, umap-learn, cugraph) your pipeline actually calls, because the base install does not include them.

Frequently asked questions

What is PyGraphistry used for?

It loads, shapes, queries and visualizes graph data from Python, using pandas, Spark, RAPIDS or Arrow tables as the input. The README describes it as a library for data scientists and developers to work with graph visualization, analytics and AI, including with native GPU acceleration.

Why would you use a graph database instead of PyGraphistry?

PyGraphistry queries dataframes in memory and renders them, so it does not persist a graph or provide transactional writes. A graph database is the right layer when the graph is the system of record, which is why the README lists connectors to Neo4j, Neptune, TigerGraph, ArangoDB and Memgraph rather than positioning the library as a replacement for them.

How do I install PyGraphistry?

Install it from PyPI with pip install graphistry. Optional backends such as igraph, networkx, polars and umap-learn are extras defined in setup.py, so they are not pulled in by the base install.

Can PyGraphistry run entirely offline?

The README describes prototyping locally with CPUs and GPUs and then deploying to Graphistry Hub or a self-hosted server for production dashboards, and it does not document an offline rendering mode. The visualization path therefore involves a server, which is a constraint to check before adopting it in a restricted environment.

What is GFQL in PyGraphistry?

GFQL is the graph query language bundled with the library, described in the README as a fully vectorized dataframe-native query language with an open-source GPU runtime. It accepts Cypher-like syntax through g.gfql("MATCH ...") and can also run remotely with g.gfql_remote([...]).

Official sources

  1. Official README
  2. Project repository
  3. Release notes
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/graphistry-pygraphistry.svg)](https://hysenlabs.com/projects/graphistry-pygraphistry)