DGL-KE: training knowledge graph embeddings at scale with dglke_train
High performance, easy-to-use, and scalable package for learning large-scale knowledge graph embeddings.
At a glance
- What is it?
- DGL-KE is an Apache-2.0 Python package from awslabs that trains TransE, TransR, RESCAL, DistMult, ComplEx and RotatE embeddings on top of DGL. It fits teams with a large graph and a training budget, not someone who wants a quick embedding on a laptop.
- Who is it for?
- Adopt DGL-KE when your graph is large enough that single-machine embedding code stops finishing, and when one of the six documented models (TransE, TransR, RESCAL, DistMult, ComplEx, RotatE) matches your task. Do not adopt it for a small graph, for temporal knowledge graphs, or if you need a maintained release cadence: the newest release listed is 0.1.1 from 2020-08-26.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 85 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What DGL-KE solves, and who it is actually for
A knowledge graph stores entities as nodes and relations as edges. To use that graph in a machine learning task, you usually need each entity and relation as a vector, which is what knowledge graph embedding produces. The README describes DGL-KE as a package for learning large-scale knowledge graph embeddings, built on top of the Deep Graph Library. That framing is the whole pitch: not a general graph learning library, and not a small embedding utility, but a trainer aimed at graphs whose size makes naive implementations impractical.
The README states the design target directly, describing knowledge graphs with millions of nodes and billions of edges. The benchmark it reports covers a graph of over 86M nodes and 338M edges, computed in 100 minutes on an EC2 instance with 8 GPUs and 30 minutes on a cluster of 4 machines at 48 cores each. Those are the project's own numbers, published in its README and paper, and they only mean something at that scale. On a graph of a few thousand triples, the distributed machinery buys you nothing and costs you setup time.
The audience is therefore narrow but real: a team that already has a large KG, has decided that link prediction or entity and relation linkage is the task, and needs training to finish in a bounded time on hardware it already owns. If you are prototyping an embedding idea for a paper, the examples directory and the notebook-examples directory give you a shorter path, but you will still be running the same commands.
How the training pipeline is put together
DGL-KE exposes three tasks, each as a separate command. Training uses dglke_train on a single machine or dglke_dist_train in a distributed environment. Evaluation reads pre-trained embeddings and runs a link prediction task on the test set with dglke_eval. Inference reads pre-trained embeddings and either predicts entity and relation linkages with dglke_predict or computes embedding similarity with dglke_emb_sim.
That split matters more than it looks. The trained artifact is a file of embeddings, and evaluation and inference are separate processes that read that file. You can train once and evaluate many times, or hand the embedding file to a different team that only runs dglke_emb_sim. The architecture diagram in the README shows this as an overall pipeline rather than a single monolithic script.
The model set is fixed and named in the README: TransE, TransR, RESCAL, DistMult, ComplEx and RotatE. The command-line surface is wide. The quick start example alone passes model name, dataset, batch size, negative sample size, hidden dimension, gamma, learning rate, max step, log interval, batch size for evaluation, adversarial negative sampling, a regularization coefficient, a test flag, thread count and process count. Each of those is a knob that changes either quality or runtime, and the README does not explain them; the documentation site is where that belongs.
One line in the README deserves attention before anything else. It says that if you just want to train a KGE model using TransE, DistMult or RotatE, you should go to awslabs/graphstorm instead. DGL-KE is still the place for the scale story, but the project itself is pointing a subset of its own users elsewhere. Read that note as a statement about where the maintainers expect new work to happen.
Installing DGL-KE and running a first TransE job
The README gives two installation steps. The first installs DGL itself, the second installs the dglke package. Both use pip3, and the README's example prefixes them with sudo, which you should only do if your environment requires it.
sudo pip3 install dgl
sudo pip3 install dglkeAfter that, the quick start trains a TransE model on the FB15k dataset. Note the DGLBACKEND environment variable set inline, which selects PyTorch as the backend for this process. The command downloads FB15k, trains, and saves the trained embeddings to a file.
DGLBACKEND=pytorch dglke_train --model_name TransE_l2 --dataset FB15k --batch_size 1000 \
--neg_sample_size 200 --hidden_dim 400 --gamma 19.9 --lr 0.25 --max_step 500 --log_interval 100 \
--batch_size_eval 16 -adv --regularization_coef 1.00E-09 --test --num_thread 1 --num_proc 8What you should see: dataset download messages, then training logs every 100 steps because of --log_interval 100, and finally an evaluation pass because --test is present. The run stops after 500 steps, since --max_step 500 caps it. That cap is what makes this a smoke test rather than a real training run; for a usable model you would raise it, and the README does not say to what.
--num_proc 8 is the parallelism knob on a single machine, and --num_thread 1 is set separately. If your machine has fewer cores than that, drop --num_proc rather than wondering why the process is fighting itself. The examples directory has per-dataset folders (fb15k, wn18, freebase, wikikg2, biokg, wn18_weighted), which is the practical place to look for a configuration closer to your own data than FB15k is.
Where DGL-KE stops being the right tool
The clearest limitation is stated by the project itself. The README note redirects anyone who only wants TransE, DistMult or RotatE training to GraphStorm. If that is your entire use case, you are being told, in the README, that this repository is not the recommended entry point anymore.
The second limitation is release cadence. The most recent release listed is 0.1.1 from 2020-08-26, and before that v0.1.0 from 2020-04-02. The repository's last push was on 2026-07-06, so work is happening on the default branch, but the published version numbering has not moved in years. Anyone who pins a version in a requirements file is pinning something from 2020, and anyone who installs from pip gets whatever that release contains rather than the current branch. That gap between branch activity and releases is the practical upgrade question, and the README does not address it.
The third is scope. DGL-KE trains static knowledge graph embeddings. There is nothing in the README about temporal knowledge graphs, edge timestamps or time-aware models, so if your graph's meaning depends on when a fact was true, this is the wrong package. A search for temporal DGL material will not find an answer here.
Finally, the tuning surface is large and undocumented in the README. Gamma, the regularization coefficient, negative sample size and batch size interact, and a first run that produces poor link prediction results is more likely a configuration problem than a code problem. Expect to spend time on the documentation site and the paper before you trust a number.
How DGL-KE differs from PyTorch-BigGraph and GraphVite
The README positions DGL-KE against two specific systems, and the comparison is worth reading literally rather than as marketing. It includes two figures, one titled DGL-KE vs GraphVite on FB15k and one titled DGL-KE vs Pytorch-BigGraph on Freebase. The accompanying claim is a 2x to 5x speedup over the best competing approaches, on the 86M node, 338M edge benchmark.
The difference in approach is where the numbers come from. PyTorch-BigGraph is a Facebook project built around partitioned, sharded training of embeddings, and the Freebase comparison in the README is the like-for-like case: both systems target graphs too large for one machine, and the claim is about how long the job takes. GraphVite is a different kind of comparison. It is a general graph embedding toolkit that covers more than knowledge graphs, and the FB15k figure compares the two on a single, comparatively small dataset where the distributed design of DGL-KE is not the reason it wins or loses.
DGL-KE's own differentiator is the DGL dependency. Embeddings are learned on top of Deep Graph Library, which means the graph representation, sampling and message passing machinery come from DGL rather than being reimplemented. If your team already runs DGL for other graph work, that is a shared dependency and a shared mental model. If you do not, installing DGL first is an extra step and an extra thing to keep in version alignment with dglke, and the README's two-step pip install is the only guidance on that relationship it gives.
Maintenance signals, upgrade cost and the licence
The repository is not archived. The last push was on 2026-07-06, which is recent enough that the default branch is being touched, but the release history tells a different story: 0.1.1 in August 2020 and v0.1.0 in April 2020. There is no release in the list after 2020. Treat the branch and the release as two separate things when you plan an upgrade, because a fix you see in the repository may not be in the version you install.
The upgrade cost follows from that. If you install with pip3 install dglke, you get the published package. If you need something that only exists on master, you are building from source, and the README does not document how. It also does not document a rollback path, a compatibility matrix between dgl and dglke versions, or a migration note between 0.1.0 and 0.1.1. Those are the questions to raise before you put this in a pipeline that other people depend on.
Licensing is straightforward on its face. The README states the project is licensed under Apache-2.0, and the repository carries both a LICENSE and a NOTICE file. Apache-2.0 is permissive and includes a patent grant, but the NOTICE file exists for a reason and should travel with any redistribution. The README also asks for a citation to the SIGIR '20 paper if you use DGL-KE in a scientific publication, which is a request rather than a licence term. None of this is legal advice; check your own obligations against the LICENSE and NOTICE files in the repository.
Editorial conclusion
Adopt DGL-KE when your graph is large enough that single-machine embedding code stops finishing, and when one of the six documented models (TransE, TransR, RESCAL, DistMult, ComplEx, RotatE) matches your task. Do not adopt it for a small graph, for temporal knowledge graphs, or if you need a maintained release cadence: the newest release listed is 0.1.1 from 2020-08-26. Before committing, verify that dglke_train completes on your own dataset with the model you intend to ship, and check whether the README's pointer to GraphStorm for TransE, DistMult and RotatE training changes your plan.
Frequently asked questions
What are knowledge graph embeddings and how are they used?
Knowledge graphs store entities as nodes and relations as edges, and embeddings turn those entities and relations into vectors so they can be used in machine learning tasks. DGL-KE trains such embeddings and then supports link prediction evaluation with dglke_eval and entity or relation linkage inference with dglke_predict.
How do I install DGL-KE?
The README gives two pip3 commands: install dgl first, then install dglke. The README's example prefixes both with sudo, which is only needed if your environment requires it.
Which models does DGL-KE support?
The README lists TransE, TransR, RESCAL, DistMult, ComplEx and RotatE. It also notes that if you only want to train TransE, DistMult or RotatE, you should go to awslabs/graphstorm instead.
Can DGL-KE run on a CPU machine or does it need GPUs?
The README states that developers can run DGL-KE on a CPU machine, a GPU machine, or clusters. Its published scale benchmark used an EC2 instance with 8 GPUs and a 4-machine cluster at 48 cores per machine.
What is the latest DGL-KE release?
The most recent release listed is 0.1.1 from 2020-08-26, following v0.1.0 from 2020-04-02. The repository's last push was on 2026-07-06, so branch activity and published releases are not in step.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/awslabs-dgl-ke)