Model or dataset
neo4j-labs/llm-graph-builder avatar
neo4j-labs/llm-graph-builder

Neo4j LLM Graph Builder: turning PDFs and videos into a property graph

Neo4j graph construction from unstructured data using LLMs

5,269 stars879 forksJupyter NotebookApache-2.0

At a glance

What is it?
Neo4j Labs' LLM Graph Builder is a FastAPI and React application that reads unstructured files and writes entities and relationships into Neo4j. It fits teams that want a queryable Cypher graph, not a one-line library call.
Who is it for?
Adopt it if you already run Neo4j 5.23 or later with APOC and want a browser UI plus a FastAPI backend that writes LLM-extracted entities and relationships into a database you control. Skip it if you want a single library call inside an existing Python pipeline, if you are on Neo4j 5.20 or older, or if you expect to run the whole stack from docker-compose against Neo4j Desktop, which the README says is not supported.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 14 days ago.
What is it written in?
Mainly Jupyter Notebook, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The gap between a pile of documents and a graph you can query

A retrieval pipeline over PDFs usually ends as a vector index: chunks in, nearest neighbours out. That answers similarity questions well and structural questions badly. Ask who reports to whom across forty contracts and a flat embedding store has nothing to traverse. Neo4j LLM Graph Builder takes the other route. It ingests unstructured sources, sends them through an LLM with the LangChain framework, and writes the extracted nodes, relationships and their properties into Neo4j as a property graph. The README describes the input side as PDFs, DOCs, TXTs, YouTube videos and web pages, with Google Cloud Storage, S3 buckets and local files as the places those documents come from.

The audience is narrower than the feature list suggests. This is for teams that have already decided the graph belongs in Neo4j and want a working interface rather than a schema design exercise. The repository is dominated by a backend and a frontend, with docker-compose.yml at the top level and a cronjob directory beside them. That layout tells you what it is: an application, not a library. If you wanted to call one function and get a graph back, you are looking at the wrong entry point, and the README offers no such function.

FastAPI, React, and a schema that steers extraction

The backend is a FastAPI app served by uvicorn from the module score. The frontend is a React app built with yarn. Between them sits the extraction step: documents are chunked, an LLM is asked to pull out nodes and relationships, and the result is written into Neo4j through Cypher. The README states that the backend uses the Cypher variable-scope subquery syntax, written as CALL (variable) { ... }, which is why Neo4j 5.23 is the floor and why 5.20 will not work.

Schema support is the part worth understanding before you upload anything. The README lists two modes: use a custom schema, or use existing schemas configured in the settings. That is the difference between a graph shaped the way you intended and one shaped by whatever the model felt like extracting that afternoon. A schema is also the cheapest guard against entity drift across a large corpus, where the same organisation appears under four spellings.

Chat is the read side of the same system. The README lists chat modes named vector, graph_vector, graph, fulltext, graph_vector_fulltext, entity_vector and global_vector, all enabled by default and configurable through VITE_CHAT_MODES. Those names describe retrieval strategies, not models: some search embeddings, some traverse the graph, some combine both. The standalone chat interface lives at the /chat-only route.

Installing the backend and building a first graph

You need Python 3.12 or higher for a separate backend deployment, and a Neo4j database at 5.23 or later with APOC installed. Aura databases, including the free tier, are supported. Neo4j Desktop is not covered by docker-compose, so the README tells Desktop users to deploy backend and frontend separately.

Start by copying the example environment file and filling in the connection details. Pre-configuring credentials here bypasses the login dialog in the UI.

bash
NEO4J_URI=<your-neo4j-uri>
NEO4J_USERNAME=<your-username>
NEO4J_PASSWORD=<your-password>
NEO4J_DATABASE=<your-database-name>

With the environment file in place, create the virtual environment, install the dependencies and start the API. The README gives this sequence exactly.

bash
cd backend
python3.12 -m venv venv
source venv/bin/activate  # On Windows: venv\Scripts\activate
pip install -r requirements.txt
uvicorn score:app --reload

The frontend runs on its own. Copy its example environment file, adjust the variables you need, then install and start the dev server.

bash
cd frontend
cp example.env .env
yarn
yarn run dev

If you prefer the container route, docker-compose.yml builds the backend from ./backend/Dockerfile, mounts the backend directory into the container and reads ./backend/.env. It also passes a long list of LLM_MODEL_CONFIG_ variables, most defaulting to empty, plus KNN_MIN_SCORE defaulting to 0.94 and IS_EMBEDDING defaulting to true. Only OpenAI and Diffbot are enabled out of the box; the README says Gemini needs additional GCP configuration. Which models appear in the UI is controlled by VITE_LLM_MODELS_PROD, and which sources appear by VITE_REACT_APP_SOURCES, whose default set is local, YouTube, Wikipedia, AWS S3 and web. Add gcs to that list and supply VITE_GOOGLE_CLIENT_ID to turn on Google Cloud Storage.

Once the app is up, the flow is: pick a source, upload or point at documents, choose an LLM, optionally supply a schema, and run the extraction. Progress and token use are visible per user and per database connection when TRACK_USER_USAGE is set to true, and the README mentions an API endpoint for checking remaining limits. Graph visualization is delegated to Neo4j Bloom rather than drawn in the app itself.

Where the setup bites: version floors and a split embedding configuration

The Neo4j 5.23 requirement is the first thing that will stop a working deployment. It is not a soft recommendation; the README ties it to a specific Cypher construct the backend emits. Anyone sitting on 5.20 has to upgrade the database before the application will function, and that upgrade touches every other client of that database.

The embedding configuration is the second trap, and it is a design choice rather than a bug. The README gives two mutually exclusive modes. With TRACK_USER_USAGE=true you supply token tracker credentials such as TOKEN_TRACKER_DB_URI and TOKEN_TRACKER_DB_USERNAME, and the embedding model is selected in the frontend under Graph Settings, then saved to the user profile. With TRACK_USER_USAGE=false you set EMBEDDING_MODEL and EMBEDDING_PROVIDER in the backend .env, and the README states plainly that the embedding model cannot be changed from the frontend in that mode. Teams that disable tracking for privacy reasons lose the ability to switch models per user, and teams that enable it take on a second database to operate.

Cost is the third. Every document passes through an LLM, and the token tracking feature exists because that adds up. The README documents the tracking, not a cap on spending. Nothing in the README describes rate limiting, retry behaviour on provider failures, or what happens to a partially extracted document when the model call fails midway. Those are the questions to ask before pointing it at a large corpus, and the README is silent on all of them.

When a plain LangChain pipeline is the better call

The closest alternative is LangChain's own LLMGraphTransformer, which this project builds on. The difference is scope. LLMGraphTransformer is a component you import into a Python process you already control: you own the chunking, the concurrency, the error handling and the writes. LLM Graph Builder is an application around that idea, with a UI, a user model, token accounting, source connectors for S3 and GCS, and a chat layer with several retrieval modes. If your pipeline already exists and you only need entity extraction, the transformer is less machinery to operate. If you want a browser, a schema editor and a chat tab without writing them, this project is the shorter path.

The other realistic comparison is to skip graph extraction entirely and keep a vector store. That is not a worse choice, it is a different one. Vector search answers similarity questions with far less setup and no schema decisions. It cannot answer a multi-hop question like which supplier is two steps removed from a recalled component. Choose based on whether your questions have hops in them.

Editorial conclusion

Adopt it if you already run Neo4j 5.23 or later with APOC and want a browser UI plus a FastAPI backend that writes LLM-extracted entities and relationships into a database you control. Skip it if you want a single library call inside an existing Python pipeline, if you are on Neo4j 5.20 or older, or if you expect to run the whole stack from docker-compose against Neo4j Desktop, which the README says is not supported. Before committing, check three things: that APOC is installed, that the LLM and embedding keys you plan to use are listed in the backend .env, and that TRACK_USER_USAGE is set the way your deployment expects, because that flag decides whether the embedding model is chosen in the frontend or pinned by EMBEDDING_MODEL and EMBEDDING_PROVIDER.

Frequently asked questions

Which AI tool is best for making graphs?

That depends on the graph you want. Neo4j LLM Graph Builder extracts nodes, relationships and their properties from unstructured files with an LLM and stores them in Neo4j, so it produces a property graph you can query with Cypher rather than a chart or a diagram.

How can I use an LLM to create a knowledge graph?

In this project you upload documents from a local machine, GCS, S3 or web sources, choose an LLM, optionally supply a custom schema or use existing schemas from the settings, and run the extraction. The application writes the extracted nodes and relationships into your Neo4j database.

Is Neo4j better than SQL for this?

The project only targets Neo4j, and the README requires Neo4j 5.23 or later with APOC installed. A relational database is not a supported backend here, so the question is settled by the deployment rather than by a feature comparison.

Official sources

  1. License: Apache-2.0
  2. neo4j-labs/llm-graph-builder on GitHub
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/neo4j-labs-llm-graph-builder.svg)](https://hysenlabs.com/projects/neo4j-labs-llm-graph-builder)