# awslabs/graphrag-toolkit: building hierarchical lexical graphs for GraphRAG

> The GraphRAG Toolkit is a collection of Python tools for turning unstructured documents into a queryable lexical graph. It ships two separate packages with different jobs, and the split is the first thing to understand before adopting it.

**awslabs/graphrag-toolkit** — Python toolkit for building graph-enhanced GenAI applications

- Repository: https://github.com/awslabs/graphrag-toolkit
- Website: https://awslabs.github.io/graphrag-toolkit/
- Stars: 442 · Forks: 106
- Language: Python
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/awslabs-graphrag-toolkit

## Two packages, two problems, one repository

The name suggests a single library. The repository does not contain one. It contains lexical-graph, which the README describes as a framework for automating the construction of a hierarchical lexical graph from unstructured data and composing question-answering strategies over it, and byokg-rag, described as an approach to Knowledge Graph Question Answering that combines LLMs with structured knowledge graphs so that users can bring their own graph. The audiences overlap but the inputs do not. Lexical-graph starts from text and builds the graph. BYOKG-RAG starts from a graph you already have and asks questions of it. If you pick the wrong one you will spend a day discovering that the ingestion pipeline you wanted is in the other directory. The topics list on the repository points at amazon-neptune, amazon-opensearch-serverless and postgresql, which is a fair summary of where the toolkit expects to sit.

## The hierarchical lexical graph model

The core design decision is that the graph is lexical and hierarchical rather than a general-purpose entity graph. The README links to a graph model page in the docs site and to an arXiv article titled Hierarchical Lexical Graph for Enhanced Multi-Hop Retrieval, which describes the design. The practical consequence is that retrieval is not a flat nearest-neighbour lookup. A question is answered by composing strategies that query the graph, which is what the README means by question-answering strategies. The repository also carries a benchmarks directory and an integration-tests directory, so there is at least an intent to measure retrieval quality rather than assert it. The README does not publish benchmark numbers, and the linked article is the place to look for the model's justification. Treat the hierarchy as the feature you are buying: if your documents have no natural structure to exploit, the extra layer buys you less.

## Installing graphrag-lexical-graph and running a first query

The README points readers at the docs site with the instruction to start there, and it publishes two PyPI badges, one for graphrag-lexical-graph and one for graphrag-byokg. Those are the distribution names to install, not the repository name. A minimal install of the lexical graph package looks like this.

```bash
pip install graphrag-lexical-graph
```

The README does not spell out the import name or the constructor arguments, so the docs site is the next stop rather than guessing. For a working end-to-end example, the repository ships Jupyter notebooks under examples/lexical-graph, including a self-guided workshop described as using GraphRAG with Amazon Neptune to improve generative AI applications. The examples directory also contains lexical-graph-local-dev and lexical-graph-hybrid-dev, which are the two paths worth reading first because they differ in how much infrastructure you need before anything runs.

The second package installs the same way under its own name.

```bash
pip install graphrag-byokg
```

There is also a lexical-graph-contrib directory at the top level, which is where contributed backends live. The README does not document the contribution contract for that directory, so check the docs site before assuming a backend you want is already there.

## Where the toolkit stops being the right choice

The strongest signal in the repository layout is the number of top-level directories tied to specific infrastructure. Topics name Neptune, OpenSearch Serverless and PostgreSQL. The examples are split into local-dev and hybrid-dev variants, which implies the full path expects managed services. If your team runs a graph database that is not in that set, the toolkit is not a neutral layer you can point at anything; you are either writing a backend under lexical-graph-contrib or you are choosing a different tool. The second limitation is documentation depth. The README is a signpost, not a manual: it delegates to the docs site, to a launch blog post, to two videos, to a Medium companion post, to the arXiv article and to the workshop notebooks. That is a lot of surface area to read before you can judge whether the retrieval quality is good enough for your corpus. The README itself does not publish accuracy figures, latency figures or cost figures, and it does not document rollback or re-ingestion behaviour when the underlying documents change.

## GraphRAG Toolkit compared with Microsoft GraphRAG

The related searches show people arriving with Microsoft GraphRAG in mind, so the difference is worth stating plainly. Microsoft's project is a self-contained pipeline: you point it at a corpus and it produces a graph and community summaries as part of one opinionated workflow. The awslabs toolkit takes the opposite stance. It separates graph construction from graph querying into two packages, and it separates the storage layer from the retrieval logic, which is why Neptune, OpenSearch Serverless and PostgreSQL appear as topics. That separation is the point: you can keep your existing graph database and add lexical-graph retrieval on top, or bring a knowledge graph you built elsewhere into byokg-rag. The cost is that you assemble more of the pipeline yourself and you depend on the docs site to tell you how the pieces fit. If you want one command that produces an answer, Microsoft's shape is closer. If you want to control where the graph lives, this one is.

## Maintenance, licensing and upgrade cost

The repository is not archived, and the last push was on 2026-09-09, which is recent enough that the codebase is moving. The most recent releases are tagged graphrag-lexical-graph/v3.19.1 and graphrag-byokg/v3.19.1, both dated 2026-08-26, with a dev tag for the lexical graph package two weeks earlier. The two packages are versioned in lockstep at the same number, which is convenient when you upgrade both but also means a change in one can force a version bump in the other. Because the packages are published separately to PyPI, you can pin them independently; the release tags suggest you should check both before upgrading either. Licence is Apache-2.0, which permits commercial use and modification and requires that you preserve the licence and NOTICE file. The repository carries a NOTICE file at the top level, so redistribution needs to include it. This is a description of the licence terms, not legal advice; if you are embedding the toolkit in a product, have your own counsel read the Apache-2.0 text and the NOTICE.

## What to read before writing code

The repository gives a clear reading order if you follow it: the docs site first, then the launch blog post, then the workshop notebooks under examples/lexical-graph. The RAG Explorer sample, hosted outside this repository, is described as a UI that compares GraphRAG and Vector RAG responses, which is the cheapest way to see whether the graph layer changes answers on your kind of question before you commit engineering time. The two videos cover the design rationale and a production use in a security intelligence center, and the arXiv article covers the graph model itself. If you only have an afternoon, read the graph model page and run the local-dev example; the hybrid-dev example will not tell you anything the local one does not until you have a backend to point it at.

## Conclusion

Adopt it if you already run Amazon Neptune, OpenSearch Serverless or PostgreSQL and want question answering that traverses a document hierarchy rather than only matching embeddings. Do not adopt it if you need a framework that is agnostic about the graph backend, or if you want a single package that does both graph construction and KGQA, because the repository ships two and they are not interchangeable. Verify first that the backend you intend to use is covered by the lexical-graph documentation, that your Python environment can install graphrag-lexical-graph and graphrag-byokg at the version you need, and that the hierarchical lexical graph model described in the arXiv article matches the shape of your documents.

## FAQ

### What is the GraphRAG Toolkit used for?

It is a collection of Python tools for building graph-enhanced Generative AI applications. The lexical-graph package builds a hierarchical lexical graph from unstructured data and composes question-answering strategies over it, while byokg-rag performs knowledge graph question answering over a graph you already have.

### What are the limitations of the GraphRAG Toolkit?

The README does not publish accuracy, latency or cost figures, and it does not document rollback or re-ingestion behaviour when source documents change. The repository layout also ties the examples and topics to Neptune, OpenSearch Serverless and PostgreSQL, so other graph backends require work under lexical-graph-contrib.

### How does the GraphRAG Toolkit differ from Neo4j GraphRAG?

The available documentation does not describe Neo4j GraphRAG, so no comparison can be made from it. What the repository does show is that this toolkit separates graph construction (lexical-graph) from graph querying (byokg-rag) and names Neptune, OpenSearch Serverless and PostgreSQL among its topics.

## Sources

- [awslabs/graphrag-toolkit on GitHub](https://github.com/awslabs/graphrag-toolkit)
- [License: Apache-2.0](https://github.com/awslabs/graphrag-toolkit/blob/main/LICENSE)
- [Project website](https://awslabs.github.io/graphrag-toolkit/)
- [README](https://github.com/awslabs/graphrag-toolkit/blob/main/README.md)
- [Releases](https://github.com/awslabs/graphrag-toolkit/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/awslabs-graphrag-toolkit
