# shenweichen/GraphEmbedding: Five Classic Node Embedding Algorithms in One Python Package

> A compact reference implementation of DeepWalk, LINE, Node2Vec, SDNE and Struc2Vec, distributed as the ge package. It is a reading and teaching codebase, not a production graph learning service, and the README is thin on anything beyond the happy path.

**shenweichen/GraphEmbedding** — Implementation and experiments  of graph embedding algorithms.

- Repository: https://github.com/shenweichen/GraphEmbedding
- Stars: 3,843 · Forks: 995
- Language: Python
- License: MIT
- Published: 2026-09-23 · Updated: 2026-09-23 · Language: en
- Canonical page: https://hysenlabs.com/projects/shenweichen-graphembedding

## What shenweichen/GraphEmbedding actually packages

This repository is a set of from-scratch implementations of five published node embedding methods: DeepWalk (KDD 2014), LINE (WWW 2015), Node2Vec (KDD 2016), SDNE (KDD 2016) and Struc2Vec (KDD 2017). Each model has a paper link and a companion Chinese article in the method table at the top of the README. The intended audience is someone who wants to read the algorithm and the code side by side, or who needs embeddings for a modest graph and does not want to assemble the pipeline from a larger framework. The design principle stated in the README is deliberately narrow: graph in, embedding out. There is no serving layer, no feature store integration and no distributed execution path. The package is published as ge, version 0.1.0, under the MIT license, and requires Python 3.7 or newer.

## Graph in, embedding out: the shared interface across all five models

All five models consume a networkx graph and expose the same three-call shape. You construct the model with its hyperparameters, call train, then call get_embeddings. That uniformity is the main practical convenience of the repository: swapping DeepWalk for Node2Vec is a two-line change in a script.

The input format is an edge list read by networkx, with the README showing the line format as node1 node2 <edge_weight>. The examples read it with nx.read_edgelist and a data argument of [('weight', int)], so the third column is parsed as an integer edge weight rather than left as a string. Underneath, the random-walk methods (DeepWalk, Node2Vec, Struc2Vec) generate walks and then hand sequences to gensim for skip-gram training, which is why gensim>=4.0.0 is a hard dependency. LINE takes a different route through edge sampling with first-order, second-order or combined objectives, and SDNE is the odd one out: it is an autoencoder trained with TensorFlow, which is why the README installs the tf extra before running examples.

## Installing ge and running the DeepWalk wiki example

The README gives a two-step procedure: clone the repository and install dependencies, then run one example script. The install command uses an editable install with the tf extra, which pulls tensorflow>=1.15.5 in addition to the base requirements listed in setup.py (gensim, networkx, joblib, fastdtw, tqdm, numpy, scikit-learn, pandas, matplotlib).

```bash
pip install -e .[tf]
```

After that, the README runs the DeepWalk example against the bundled Wiki edge list. The data directory and the examples directory are both top-level entries in the repository, so the relative paths in the example scripts resolve from the repository root.

```bash
python examples/deepwalk_wiki.py
```

For a first real use, the README's DeepWalk snippet shows the whole flow. Read the graph, construct the model with a walk length and a number of walks, train, then collect vectors.

```python
G = nx.read_edgelist('../data/wiki/Wiki_edgelist.txt',create_using=nx.DiGraph(),nodetype=None,data=[('weight',int)])

model = DeepWalk(G,walk_length=10,num_walks=80,workers=1)
model.train(window_size=5,iter=3)
embeddings = model.get_embeddings()
```

What you should see is a dictionary-like mapping from node id to a vector, produced after the training loop finishes. The other four models follow the same pattern with different constructor arguments: LINE takes embedding_size and order (first, second or all), Node2Vec takes p and q, SDNE takes hidden_size, and Struc2Vec takes the walk length, walk count, workers and a verbose interval.

## Where the repository stops short

The README documents no evaluation protocol. There is no reported accuracy, no link prediction benchmark, no node classification score, and no guidance on how to tell whether an embedding is any good. For a project whose stated purpose includes experiments, that is a real gap: you get vectors, but the repository does not tell you what a correct output looks like.

Dependency weight is the second constraint. SDNE requires TensorFlow 1.15.5 or newer through the tf extra, and that extra is installed by the README's own instructions even if you only intend to run DeepWalk. Nothing in the README describes a lighter install path for the walk-based models alone, so the documented route drags in the full deep learning stack.

The third limitation is scale and dynamism. The examples operate on bundled wiki and flight edge lists, and nothing in the repository describes incremental updates, graph mutation, or a way to add a new node without retraining. If your graph changes hourly, this is the wrong tool; the whole pipeline assumes a static edge list that you read once and embed once. The README is also silent on rollback, checkpointing and versioning of trained embeddings, so there is no documented way to reproduce an embedding produced last week.

## How this differs from a graph neural network library

The natural alternative for someone who outgrows this repository is a graph neural network library, where message passing layers are trained end to end on a downstream task rather than producing task-agnostic vectors from random walks. The difference in approach is structural. DeepWalk, Node2Vec and Struc2Vec here are unsupervised: they generate walks, treat them as sentences, and train skip-gram. A GNN library instead defines a convolution over the adjacency structure and backpropagates a supervised loss, which means the embeddings are shaped by the label you care about.

That has a practical consequence. If you have labels and a fixed task, a GNN will usually be the better fit, and it will handle node features that this repository has no mechanism for. If you want reusable vectors for a graph with no labels, or you want to see the algorithm rather than a framework abstraction, the walk-based implementations here are the more direct route. The trade is that you own the evaluation, the retraining loop and the dependency management yourself.

## Maintenance, licensing and the cost of upgrading

The repository is not archived, and its most recent push was on 2026-04-26. There are no retrieved releases, which means upgrades are tracked through the master branch rather than tagged versions. The version in setup.py is 0.1.0, and installing with the editable flag means you are tracking whatever is on master at clone time.

The practical upgrade cost follows from that. Because there are no releases, there is no changelog to read before pulling. Your exposure is concentrated in three dependencies: gensim>=4.0.0, which the walk-based models rely on for skip-gram training; networkx, which defines the graph object every model consumes; and tensorflow>=1.15.5, which only SDNE needs. A major gensim or networkx release is the event most likely to break your scripts, and the repository's test extra (pytest, pytest-cov, python-coveralls) is what you would run to find out.

The license is MIT, stated in both setup.py and the LICENSE file at the repository root. MIT is permissive and places few conditions on redistribution or modification. This is not legal advice; if you are embedding the package in a commercial product, read the LICENSE file and your own dependency licenses, particularly TensorFlow's, before you ship.

## Conclusion

Adopt it if you want to read or teach the five classic node embedding algorithms, or to produce embeddings from a small edge list without pulling in a deep learning framework for the whole pipeline. Do not adopt it if you need a supported library with a documented API, incremental updates on a changing graph, or a GPU path. Before committing, verify that the dependency set in setup.py installs cleanly on your Python version, that the tensorflow>=1.15.5 extra resolves for SDNE, and that the edge list you intend to feed nx.read_edgelist actually carries the weight attribute the examples assume.

## FAQ

### What is graph embedding in shenweichen/GraphEmbedding?

It is the process the repository implements: turning a networkx graph into a mapping from each node to a numeric vector. The README describes the design as graph in, embedding out, and provides five algorithms (DeepWalk, LINE, Node2Vec, SDNE, Struc2Vec) that each expose train and get_embeddings.

### How does graph embedding in shenweichen/GraphEmbedding differ from vector embedding?

The output is the same kind of object, a vector per item, but the input is a graph rather than text. The repository reads an edge list with networkx and generates embeddings from graph structure, using random walks and skip-gram for DeepWalk, Node2Vec and Struc2Vec, edge sampling for LINE, and an autoencoder for SDNE.

### What is an embedding example in shenweichen/GraphEmbedding?

The README's DeepWalk snippet is the shortest one: read the Wiki edge list with nx.read_edgelist, construct DeepWalk with walk_length=10 and num_walks=80, call train with window_size=5 and iter=3, then call get_embeddings to get the vectors.

## Sources

- [Issues](https://github.com/shenweichen/GraphEmbedding/issues)
- [License: MIT](https://github.com/shenweichen/GraphEmbedding/blob/master/LICENSE)
- [README](https://github.com/shenweichen/GraphEmbedding/blob/master/README.md)
- [shenweichen/GraphEmbedding on GitHub](https://github.com/shenweichen/GraphEmbedding)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/shenweichen-graphembedding
