Open-source project
rahulnyk/knowledge_graph avatar
rahulnyk/knowledge_graph

rahulnyk/knowledge_graph: turning a text corpus into a concept graph with a local LLM

Convert any text to a graph of knowledge. This can be used for Graph Augmented Generation or Knowledge Graph based QnA

4,087 stars634 forksJupyter NotebookMIT

At a glance

What is it?
A Jupyter notebook that splits a PDF into chunks, asks a locally hosted Mistral 7B model for concepts, and writes the result into a NetworkX graph. It is a working method for graph-based retrieval, not a packaged library.
Who is it for?
Adopt it if you want to see how concept extraction and co-occurrence weighting actually produce a graph, and you are comfortable editing a notebook rather than importing a package. Do not adopt it as a library: there is no published API, no release, and the pyproject.toml still describes a different project.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 46 days ago.
What is it written in?
Mainly Jupyter Notebook, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 25, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The problem: vector retrieval loses the connections between concepts

Retrieval Augmented Generation with a vector database returns passages that are semantically close to a query. What it does not return is the structure between those passages. If a document argues that one concept causes another, a similarity search can surface both concepts without ever exposing the link. The README frames this directly: it describes Graph Retrieval Augmented Generation as "a new and improved version of Retrieval Augmented Generation (RAG) where we use a vectory db as a retriever to chat with your documents", with the graph acting as the retriever instead.

The intended audience is narrow. This is for someone who has a body of text, usually a PDF, wants to inspect the concept structure inside it, and is willing to run a notebook to get there. It is not for a team that needs a graph service behind an HTTP endpoint. The README says the author created "a simple knowledge graph from a PDF document" and that the process is a simplification of the general pipeline. The repository layout matches that claim: extract_graph.ipynb sits at the top level next to helpers/, data_input/ and data_output/, with no server entry point anywhere.

Concepts instead of entities, and why the edges are text chunks

The extraction step deliberately avoids a named entity recognition model. The README draws the distinction: "'Bangalore' is an entity, and 'Pleasant weather in Bangalore' is a concept", and states that in the author's experience concepts make more meaningful graphs than entities. That single choice drives everything downstream. A concept is a phrase, so it is longer, noisier, and much harder to deduplicate than a proper noun. The README acknowledges this in the contributions section, where the first suggested improvement is to "use embeddings to deduplicate semantically similar concepts".

The pipeline in extract_graph.ipynb has four steps. Text is split into chunks and each chunk gets a chunk_id. An LLM extracts concepts and their semantic relationships from each chunk, and those relations get a weight of W1. Concepts appearing in the same chunk are treated as related by contextual proximity, with a weight of W2. Finally, similar pairs are grouped, their weights summed, and their relation names concatenated, so one edge remains between any two distinct concepts, carrying a total weight and a list of relations as its name.

The consequence is worth stating plainly: an edge in this graph is not a typed relationship. It is a bag of relation strings plus a number, and the number mixes two different signals, one from the model's reading and one from mere co-occurrence. That makes the graph good for centrality and community work, which the README lists, and weaker for queries that need a precise predicate. The notebook also computes node degree and communities for sizing and colouring, which is the part that makes the Pyvis output readable rather than a hairball.

Installing with Docker and running the notebook

The README lists Docker as the only prerequisite for the containerised path and calls it recommended. Clone the repository and move into it first.

bash
git clone https://github.com/rahulnyk/knowledge_graph.git
cd knowledge_graph

Build the image from the Dockerfile at the repository root. It uses python:3.11-slim, installs Poetry, runs poetry install --no-root, and copies the application code into /app.

bash
docker build -t knowledge-graph .

Run it, publishing port 8888. The Dockerfile's CMD starts JupyterLab with --ip=0.0.0.0 --port=8888 --no-browser --allow-root, so the container serves a lab interface rather than a notebook file.

bash
docker run -p 8888:8888 knowledge-graph

After that, open the JupyterLab URL printed in the container logs and open extract_graph.ipynb. The README is explicit that this is the notebook you have to tweak: "To generate a graph this the notebook you have to tweak." There is no command-line entry point to run instead.

The model side is separate from the container. The README instructs the reader to install Ollama, then run the following in a terminal, which pulls the model and starts the Ollama server.

bash
ollama run zephyr

Note the mismatch: the setup step pulls zephyr, while the tech stack section says the project uses Mistral 7B OpenOrca for extraction. Both statements are in the README, and the notebook is the only place that resolves which model tag is actually called. Check the model name in the notebook before you assume the pull above is sufficient.

Local inference keeps the cost at zero, and that is the whole economic argument

The README states the project "adopted a no-GPT approach here to keep things economical" and that generating the graph is "basically free (No calls to GPT)" because Mistral 7B runs locally through Ollama. That is the strongest reason to pick this over a hosted extraction pipeline: you can run it on a personal machine without a per-token bill, which matters when you are iterating on chunk sizes and prompts and will re-run the extraction dozens of times.

The trade-off is that extraction quality is now your problem. The README credits Mistral 7B OpenOrca with following system prompt instructions well, but a 7B model will produce inconsistent concept phrasing across chunks, and the deduplication step that would fix it is listed as an unimplemented suggestion, not a feature. If your corpus is large, the constraint that bites is not money but wall-clock time on local hardware, and the README offers no throughput figures. Everything in the graph is only as good as the concept strings the model returns.

Where this is the wrong tool

There is no packaging. pyproject.toml declares the name knowledge-graph with version 0.1.0, but its description reads "Knowledge Graph package for code analysis", which does not describe this project, and no recent releases were retrieved. So there is no version to pin, no changelog, and no importable module documented anywhere in the README. The Dockerfile installs dependencies with --no-root, meaning the project itself is not installed as a package inside the image. You are meant to work inside the notebook.

There is also no persistence layer. The README says the schema is held in Pandas dataframes and notes that a graph database "can be used at a later stage". That means the graph lives in memory and in whatever files you export. If you need concurrent queries, transactions, or a Cypher-style query language, this is not that, and the README does not pretend otherwise.

Finally, the extraction is one-shot. Re-running the notebook rebuilds the graph from scratch. Nothing in the README describes incremental updates, so adding a document to the corpus means re-running the whole pipeline and paying the full local inference cost again.

Compared with putting the documents in a vector store

The obvious alternative is the standard RAG stack: embed the chunks, store them in a vector database, and retrieve by cosine similarity. The difference is what gets retrieved. A vector store returns chunks ranked by similarity to the query, and the relationships between those chunks are never materialised. This project returns nodes and weighted edges, so you can compute degree and communities and ask which concepts bridge otherwise separate parts of the corpus. The README presents the graph as a replacement retriever for chat, not as a companion index.

The cost of that difference is the extraction pass. A vector store needs an embedding model and nothing else. This pipeline needs a chat model to read every chunk and emit concepts and relations, which is slower and far more sensitive to prompt wording. Searches around this project include "knowledge graph vs vector database", and the honest answer is that they retrieve different things: one retrieves text, the other retrieves structure that was inferred from text. Neither is a drop-in for the other, and the README does not claim the graph replaces the vector store in a hybrid setup.

Licence and maintenance

The repository is MIT licensed, with the LICENSE file at the top level. MIT is permissive: it allows reuse, modification and redistribution provided the copyright notice and permission notice are kept. Note that the licence covers this repository's code, not the Mistral 7B OpenOrca weights it calls, which are distributed separately on Hugging Face under their own terms. Anyone shipping a product built on the extraction step should read the model card, not just the LICENSE file here.

On maintenance, the last push was on 2026-08-15. The repository is not archived. The README ends with a contributions section that lists unimplemented improvements, including the embedding-based deduplication of similar concepts, and the truncated list suggests more items follow. Treat the notebook as the deliverable and the surrounding files as scaffolding. Upgrading means re-reading extract_graph.ipynb, because that is where the method lives; the pyproject.toml pins langchain at ^0.0.335, and that is the kind of dependency where a version bump is more likely to break the notebook than to fix it.

Editorial conclusion

Adopt it if you want to see how concept extraction and co-occurrence weighting actually produce a graph, and you are comfortable editing a notebook rather than importing a package. Do not adopt it as a library: there is no published API, no release, and the pyproject.toml still describes a different project. Before running anything, confirm that Ollama is installed and that the model tag you pull matches the one the notebook calls, because the README names zephyr in the setup step and Mistral 7B OpenOrca in the tech stack.

Frequently asked questions

What is meant by a knowledge graph in rahulnyk/knowledge_graph?

The README quotes IBM's definition: a network of real-world entities such as objects, events, situations or concepts, illustrating the relationships between them, usually stored in a graph database and visualised as a graph structure. In this project the nodes are concepts extracted by an LLM and the edges are the text chunks in which two concepts appear.

How do I use a knowledge graph with an LLM in this project?

The LLM does the extraction rather than the querying. The notebook sends each text chunk to Mistral 7B OpenOrca through Ollama and asks for concepts and their semantic relationships, then weights those relations as W1 and adds a W2 weight for concepts that merely co-occur in the same chunk.

How do I use a knowledge graph in RAG with rahulnyk/knowledge_graph?

The README describes Graph Retrieval Augmented Generation, where the graph replaces the vector database as the retriever so you can chat with your text. It presents this as an improved version of RAG rather than a hybrid, and the graph is built from the same chunks you would otherwise embed.

How do I set up a knowledge graph with rahulnyk/knowledge_graph?

Install Docker, clone the repository, build the image with docker build -t knowledge-graph ., and run it with docker run -p 8888:8888 knowledge-graph to get JupyterLab on port 8888. Separately install Ollama and pull a model, then edit extract_graph.ipynb, which the README says is the notebook you have to tweak.

Official sources

  1. Issues
  2. License: MIT
  3. rahulnyk/knowledge_graph on GitHub
  4. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/rahulnyk-knowledge-graph.svg)](https://hysenlabs.com/projects/rahulnyk-knowledge-graph)