# neo4j-graphrag-python: building GraphRAG pipelines on a Neo4j graph

> The official Neo4j Python package for GraphRAG ships retrievers, a knowledge graph builder and LLM integrations. It is a good fit when your data already lives in Neo4j, and a poor one when it does not.

**neo4j/neo4j-graphrag-python** — Neo4j GraphRAG for Python

- Repository: https://github.com/neo4j/neo4j-graphrag-python
- Website: https://neo4j.com/docs/neo4j-graphrag-python/current/
- Stars: 1,295 · Forks: 247
- Language: Python
- License: NOASSERTION
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/neo4j-neo4j-graphrag-python

## What neo4j-graphrag-python solves, and for whom

Retrieval augmented generation usually means embedding chunks of text and asking a vector index for the nearest ones. That works until a question needs a relationship: which subsidiary reports to which parent, which person inherited which title. A vector index returns passages that look similar, not passages that are connected. neo4j-graphrag-python is Neo4j's first-party answer to that gap. It is a Python package that builds and queries a graph, so retrieval can follow edges instead of only cosine distance.

The audience is narrow and specific. You are a Python developer who already runs Neo4j, or who is willing to. The package assumes a driver connection, a URI, a username and a password. It does not ship a database, a hosting service or a managed index. If your data sits in Postgres and you have no intention of moving it, this library has nothing to offer you that a vector extension would not.

## Retrievers, the KG builder and how the pieces connect

The package splits into two halves. The retrieval half takes a question and returns context for an LLM. The construction half takes raw text or PDFs and turns them into a graph. They meet in the middle: you can build a graph with the pipeline classes, then query it with a retriever.

On the construction side there are two classes. `Pipeline` is the low-level one, described in the README as offering extensive customization for advanced use cases. `SimpleKGPipeline` is an abstraction over it, aimed at getting a graph built without wiring every stage by hand. Both accept text and PDFs. The pipeline needs an LLM to extract entities and relationships, and an embedder to attach vectors to what it extracts. You declare the shape of the graph up front: node types, relationship types, and patterns that constrain which node type may connect to which. In the README example those are `Person`, `House` and `Planet` joined by `PARENT_OF`, `HEIR_OF` and `RULES`. The LLM is then asked to find only those things in the text.

That schema is the interesting design decision. It trades recall for predictability. An open extraction prompt will surface whatever the model notices, including entity types you never planned for and cannot query later. A constrained schema produces a graph whose shape you already know, which is what makes downstream Cypher tractable. The cost is that anything outside the declared types is silently dropped, and the README does not describe a mechanism for reviewing what was discarded.

The README also points at retrievers for the case where the graph is already populated, and links blog posts covering vector search enriched by graph traversal, hybrid retrieval, and a `Text2CypherRetriever` that turns a natural language question into a Cypher query. Those posts are the documentation for the retrieval half; the README itself stays at the level of a feature list.

## Installing neo4j-graphrag-python and running a first pipeline

The base install pulls in the Neo4j driver, pydantic, pypdf, numpy, scipy and a handful of smaller dependencies. Python 3.10 through 3.14 are listed as supported, and `requires-python` in pyproject.toml is `>=3.10.0,<3.15`.

```bash
pip install neo4j-graphrag
```

LLM providers are optional extras, and the README states that at least one is required for RAG and for the KG builder pipeline. The extras listed are `ollama`, `openai`, `google`, `google-genai`, `cohere`, `anthropic`, `mistralai` and `bedrock`. There are also extras for `sentence-transformers`, vector stores (`weaviate`, `pinecone`, `qdrant`), and three that matter later: `experimental`, `nlp` and `fuzzy-matching`.

```bash
pip install "neo4j-graphrag[openai]"
```

Before running anything you need a reachable Neo4j instance and, for graph construction, the APOC core library installed on it. The README states this as a note attached to the knowledge graph construction section. The examples also expect `NEO4J_URI`, `NEO4J_USERNAME`, `NEO4J_PASSWORD` and `OPENAI_API_KEY` in the environment.

The README's own example is the shortest path to something real. It connects a driver, declares the schema, instantiates an embedder and an LLM, then runs the pipeline over one paragraph of text.

```python
import asyncio

from neo4j import GraphDatabase
from neo4j_graphrag.embeddings import OpenAIEmbeddings
from neo4j_graphrag.experimental.pipeline.kg_builder import SimpleKGPipeline
from neo4j_graphrag.llm import OpenAILLM

driver = GraphDatabase.driver("neo4j://localhost:7687", auth=("neo4j", "password"))
kg_builder = SimpleKGPipeline(
    llm=OpenAILLM(model_name="gpt-5"),
    driver=driver,
    embedder=OpenAIEmbeddings(model="text-embedding-3-large"),
    schema={"node_types": ["Person", "House", "Planet"]},
    on_error="IGNORE",
    from_file=False,
)
asyncio.run(kg_builder.run_async(text="Paul is the heir of House Atreides."))
driver.close()
```

What you should see is a graph in your database: nodes for the entity types you declared, relationships between them, and embeddings attached. Note the import path in the README example. `SimpleKGPipeline` is imported from `neo4j_graphrag.experimental.pipeline.kg_builder`, which tells you where that class currently lives.

Also note `on_error="IGNORE"`. The README uses it in the sample without explaining what errors would otherwise surface. For a first run it keeps the pipeline moving; for a real corpus it means malformed extractions disappear quietly.

## Where the package pushes back

The experimental namespace is the first thing to understand before you plan anything. The README is direct about it: features there are under development, may be incomplete, may change without notice, and may be removed. They are not recommended for production, and support is best-effort. The knowledge graph builder example imports from that namespace. So the most visible feature in the README, the one the sample code demonstrates, sits in the part of the package Neo4j declines to call production-ready.

That is not a contradiction so much as a maturity boundary. If your plan is to run a KG construction pipeline over a large corpus in production, you are building on code the README says may break between releases. Version 1.19.0 shipped on 2026-08-26, and the two prior releases were 1.18.0 on 2026-06-24 and 1.17.0 on 2026-05-27. A roughly monthly cadence with a 1.x version number is a reasonable signal, but it also means breaking changes in the experimental namespace have had several opportunities to land.

The second constraint is the APOC dependency. Knowledge graph construction requires APOC core on the Neo4j instance. That is not a pip install; it is a server-side library you or your platform team must add. On a managed Neo4j deployment you need to confirm APOC is available before writing any pipeline code.

The third is the `nlp` extra. The README states it is not supported on Python 3.14 because of an upstream spaCy import-time issue, and points to spaCy issue 13895. The workaround is Python 3.13 or earlier for spaCy-based features. If you are standardizing on 3.14 across a team, that extra is closed to you until spaCy resolves it.

Finally, an LLM is mandatory for RAG and for graph building. There is no local, model-free path described in the README. Every extraction and every answer costs tokens, and the quality of your graph is bounded by the extraction model's judgement about what counts as a `Person` or a `RULES` relationship.

## Choosing between this and a vector store like Qdrant or Pinecone

The obvious alternative is a vector database plus an embedding model, with no graph at all. Qdrant and Pinecone both appear in this package's own optional dependencies, which is a useful signal: Neo4j does not treat them purely as rivals. The `weaviate`, `pinecone` and `qdrant` extras exist so that retrievers can pull vectors from those stores instead of from Neo4j.

The difference in approach is what the retrieval step is allowed to reason over. A vector store answers "which chunks are semantically near this question". A graph retriever can answer "which chunks are near this question, and what is connected to the entities in them". The README's linked posts describe exactly that split: vector search enriched by graph traversal, then hybrid retrieval, then hybrid retrieval with traversal on top. Each step adds structure to the candidate set.

That extra structure is not free. A vector store is a single service with one index to maintain. This package asks you to run Neo4j, keep APOC available, define and maintain an extraction schema, and pay an LLM to build the graph before you can query it. If your questions are answered by a paragraph that happens to be similar to the query, the graph adds cost and no accuracy. If your questions are about how things relate, the graph is the only one of the two that can answer them without the LLM guessing.

## Maintenance, licensing and the upgrade bill

The repository is not archived, and the last push was on 2026-09-09. Releases have landed at a steady pace through 2026. The project is published by Neo4j, Inc, with a contact address of team-gen-ai@neo4j.com, and the README describes it as first-party with long-term support and maintenance directly from Neo4j.

Licensing needs a careful read rather than a summary. The repository root contains `LICENSE.APACHE2.txt`, `LICENSE.PYTHON.txt`, `LICENSE.txt` and `NOTICE.txt`, and pyproject.toml declares `license = {text = "Apache License, Version 2.0"}`. The GitHub metadata reports the licence as NOASSERTION, which means the platform's classifier could not reduce the repository to a single SPDX identifier. The presence of a separate Python licence file and a NOTICE file is the reason. If you redistribute the package or bundle it into a product, read those files rather than relying on the pyproject field alone. This is a description of what the repository contains, not legal advice.

Upgrade cost is concentrated in two places. The base dependencies pin `neo4j>=5.28.4,<7.0.0`, so a major driver release is the kind of change that will require attention. The experimental namespace is the other: the README explicitly says breaking changes and deprecations should be expected there, and the KG builder lives there. Budget for reading the changelog before every minor bump if a pipeline depends on it. The `nlp` extra carries a Python version ceiling that will not lift until spaCy fixes its import-time issue upstream.

## Conclusion

Adopt it if your corpus is already in Neo4j and you want retrievers and a KG builder that speak Cypher and the official driver. Do not adopt it if you have no Neo4j instance, or if a plain vector store plus an embedding model already answers your questions. Before committing, verify that APOC core is installed if you plan to use the knowledge graph construction features, confirm your Python version is 3.10 through 3.14, and check whether the components you need sit in the experimental namespace.

## FAQ

### Can I use GraphRAG with Neo4j?

Yes, that is what this package is for. It is Neo4j's official Python library for building graph retrieval augmented generation applications on a Neo4j database, and it ships both retrievers and knowledge graph construction pipelines.

### Can I use Neo4j with Python?

Yes. This package depends on the official `neo4j` driver and connects through `GraphDatabase.driver` with a URI, username and password. The README examples use `neo4j://localhost:7687` as the default URI.

### How can I use GraphRAG in Python?

Install the package with `pip install neo4j-graphrag`, add an LLM provider extra such as `neo4j-graphrag[openai]`, and then either build a graph with `SimpleKGPipeline` or query an existing one with a retriever. Graph construction also requires the APOC core library on the Neo4j instance.

## Sources

- [Issues](https://github.com/neo4j/neo4j-graphrag-python/issues)
- [neo4j/neo4j-graphrag-python on GitHub](https://github.com/neo4j/neo4j-graphrag-python)
- [Project website](https://neo4j.com/docs/neo4j-graphrag-python/current/)
- [README](https://github.com/neo4j/neo4j-graphrag-python/blob/main/README.md)
- [Releases](https://github.com/neo4j/neo4j-graphrag-python/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/neo4j-neo4j-graphrag-python
