Model or dataset
TheAiSingularity/graphrag-local-ollama avatar
TheAiSingularity/graphrag-local-ollama

GraphRAG Local Ollama: Microsoft's GraphRAG pipeline driven by Ollama models

Local models support for Microsoft's graphrag using ollama (llama3, mistral, gemma2 phi3)- LLM & Embedding extraction

1,108 stars162 forksPythonMIT

At a glance

What is it?
A fork of Microsoft's GraphRAG that swaps hosted OpenAI calls for local Ollama models, adds a Gradio web UI and five query modes. The trade-off is that indexing quality now depends on whichever local model you can run.
Who is it for?
Adopt it if you need Microsoft's two-stage graph index over private documents and cannot send text to a hosted API, or if you want a browser UI around indexing and graph inspection without writing the plumbing yourself. Do not adopt it if you need a supported, versioned release: pyproject.toml still reads 0.1.1, there are no releases, and the README's own contributing section asks for help with llama integration.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 145 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What GraphRAG Local Ollama changes about Microsoft's GraphRAG

Microsoft's GraphRAG answers questions over a private corpus by building a graph-based text index in two stages: an LLM derives an entity knowledge graph from the source documents, then pregenerates community summaries for groups of closely related entities. At query time each community summary produces a partial response, and those partial responses are summarized again into the final answer. The paper abstract quoted in the README frames the motivation as global sensemaking: naive retrieval-augmented generation fails on questions like "What are the main themes in the dataset?" because that is query-focused summarization rather than explicit retrieval.

That pipeline assumes you can pay for hosted model calls, and it makes a lot of them. This repository is an adaptation of Microsoft's GraphRAG "tailored to support local models downloaded using Ollama", in the README's words, for both LLM and embedding extraction. The audience is therefore narrow and specific: engineers who already accept the graph-index approach but cannot route document text through a hosted API, whether for cost, for data residency, or because the corpus is internal. If you were never sold on GraphRAG's two-stage design in the first place, this fork does not change that argument. It changes who pays for the tokens.

The indexing and query mechanism, and where LazyGraphRAG cuts the cost

The pipeline is Microsoft's, and the repository layout reflects that: the graphrag package sits at the top level alongside settings.yaml, an input directory and an output directory. The Dockerfile's default command is the indexing entry point, python -m graphrag.index --root /app, and querying is a separate invocation of python -m graphrag.query with a --method flag. Documents go in, an index is built under the root, and queries read that index.

The interesting design choice is LazyGraphRAG. Setting lazy_graph_rag: true in settings.yaml skips community summarization at index time and generates summaries on demand during queries. The README puts indexing at roughly 99 percent faster as a result. That number is the project's own claim, not an independent measurement, and it describes index time only. The documented trade-off is explicit: first-query responses are slower than standard global search because the summaries are computed then. For a corpus you index once and query many times, that is a bad trade. For a corpus you re-index often, it is the whole reason to use this fork.

The five query modes map onto different amounts of that precomputation. Global and Local search work against the standard index. Basic search is vector similarity only: embed the query with your Ollama embedding model, retrieve the top-10 most similar entities, build context from those entities plus their relationships and source text units, and answer. It needs no community reports at all. DRIFT runs a primer phase that decomposes the question into scored sub-questions using community reports, a priority-queue search loop that answers each sub-question and generates follow-ups, and a reduce phase that synthesizes the intermediate answers. Lazy is the on-demand variant. The dependency list in requirements.txt explains the shape of the thing: lancedb for the vector store, networkx and graspologic for graph work, tiktoken and nltk for text, and openai pinned at 1.35.7, which is how the pipeline talks to an Ollama endpoint that speaks the OpenAI-compatible API.

Installing GraphRAG Local Ollama with docker-compose

The README calls Docker the fastest path because it needs no Python environment. Copy the example environment file, then bring the stack up. The compose file defines two services, ollama on port 11434 and graphrag, with a healthcheck on the Ollama service and depends_on waiting for it to report healthy.

bash
cp .env.example .env
docker-compose up --build

What you should see is Ollama starting first, the graphrag container waiting on the healthcheck, and then indexing running automatically on container start. The compose file mounts ./input into /app/input, ./output into /app/output, and settings.yaml into /app/settings.yaml, and it sets OLLAMA_HOST=http://ollama:11434 with GRAPHRAG_API_KEY=ollama. Drop .txt documents into ./input before or during startup.

bash
docker-compose run graphrag python -m graphrag.query \
  --root /app --method global "What are the main themes?"

That command runs a query against the index the container built. Swap --method for local, drift, basic or lazy to change the retrieval strategy. To try the lazy path, set the key in settings.yaml first:

yaml
lazy_graph_rag: true

Then query with --method lazy. Note the ordering problem: lazy mode changes what indexing produces, so the setting has to be in place before the index is built, not after.

The alternative entry point is the web UI, which the README documents for a local Python install rather than the container.

bash
pip install -r requirements-ui.txt
python app.py

The README says to open http://localhost:7860. The UI has four tabs: Index for uploading .txt files and watching live log output, Query for the five search modes, Graph for an interactive knowledge-graph visualizer that requires pyvis, and Settings for editing model names, chunk size and the LazyGraphRAG toggle. The Graph tab is the one feature here that Microsoft's CLI does not give you.

Model choice is the real failure mode, not the code

Everything in this fork routes entity extraction, community summarization, embeddings and answer generation through local models. The README names llama3, mistral, gemma2 and phi3 in the repository description. That list is a description, not a compatibility guarantee, and the README does not publish extraction quality figures for any of them. Entity extraction is the step where a weak model hurts most: if it misses entities or invents relationships, the graph is wrong before any query runs, and no query mode repairs that. A 7B model that produces serviceable chat answers is not automatically a serviceable graph extractor.

The second failure mode is hardware. The compose file ships without GPU reservations, and the README tells you to add deploy.resources.reservations.devices to the ollama service yourself for acceleration. Without that, indexing a non-trivial corpus runs on CPU. Combine that with a large model and the practical result is an indexing job measured in hours, which is exactly the situation LazyGraphRAG exists to avoid.

The third is query-mode mismatch. DRIFT's primer phase depends on community reports, and the README is direct that if lazy_graph_rag: true was used during indexing, the primer step will have limited context. So the two headline features work against each other: enabling lazy indexing to save hours leaves DRIFT with less to reason over. Pick one based on whether you re-index often or query deeply.

Finally, JSON and JSONL input is supported alongside .txt and .csv. The README gives the JSON array shape with id, title and text keys, and the example shows title as optional. If your documents use different key names, the README does not document a mapping option, so plan on converting.

How this compares to Microsoft GraphRAG and to plain vector RAG

Against upstream Microsoft GraphRAG, the difference is the model backend and the surface area. Upstream expects hosted model access and gives you a CLI and a Python library. This fork points the same pipeline at an Ollama endpoint through the OpenAI-compatible client, and adds a Gradio app, a graph visualizer and the lazy indexing path. The cost difference is the point; the capability difference is that a local model is generally weaker at the extraction and summarization steps the pipeline leans on. Upstream also carries Microsoft's release process. This repository has no releases, and pyproject.toml still declares version 0.1.1 with a note that the version is managed by scripts/release.sh.

Against plain vector RAG, the difference is architectural rather than operational. Basic search here is essentially vector RAG: embed, retrieve the top-10 similar entities, answer. If that is all you need, the graph machinery, community reports and DRIFT loop are overhead you are paying for in index time. The GraphRAG paper's own argument is that vector retrieval fails on corpus-level questions, and that is the case this project is built for. If your questions are "who is Alan Turing" rather than "what are the main causes of the conflict", --method basic is the honest choice and you should question whether you need this repository at all.

One gap worth naming: the README's contributing section opens with "Need support for llama integration." Whatever that refers to in the maintainers' minds, it is an admission that integration work remains open, and it sits at the top of the document.

Licence, maintenance and what an upgrade costs

The repository is MIT licensed, and pyproject.toml declares license = "MIT" for the graphrag package. MIT is permissive: you can use, modify and redistribute the code, including commercially, provided the copyright notice and licence text travel with it. That is a description of the licence terms, not legal advice, and it says nothing about the licence of the models you pull through Ollama, which is a separate question you have to answer per model.

The maintenance picture is mixed. The repository is not archived and the last push was on 2026-05-08. There are no releases, so there is no changelog to read and no tagged version to pin. The version string in pyproject.toml, 0.1.1, is generated by poetry-dynamic-versioning from git tags, so it reflects tags rather than a published artifact.

Upgrade cost is dominated by the pinned dependency set. requirements.txt pins exact versions across the board: lancedb 0.9.0, numpy 1.25.2, scipy 1.12.0 with a comment in pyproject.toml calling 1.13.0 a footgun, numba 0.60.0, nltk 3.8.1, openai 1.35.7. That is a snapshot of a working environment, which is good for reproducibility and bad for moving forward. Python is constrained to >=3.10,<3.13, so 3.13 is out. The openai pin at 1.35.7 is the one to watch: it is the client used to reach Ollama, and it is old relative to that library's release cadence. Bumping it means re-testing the whole query path, and there is no test matrix in the repository to tell you what would break.

Editorial conclusion

Adopt it if you need Microsoft's two-stage graph index over private documents and cannot send text to a hosted API, or if you want a browser UI around indexing and graph inspection without writing the plumbing yourself. Do not adopt it if you need a supported, versioned release: pyproject.toml still reads 0.1.1, there are no releases, and the README's own contributing section asks for help with llama integration. Before committing, run the Docker path on a small corpus, confirm the Ollama model names in settings.yaml match what your machine can actually pull, and check whether community reports exist in the output directory, since DRIFT degrades without them.

Frequently asked questions

What is GraphRAG Local Ollama?

It is an adaptation of Microsoft's GraphRAG that supports local models downloaded with Ollama for both LLM and embedding extraction, so the indexing and query pipeline runs without hosted model calls. It adds a web UI and extra query modes such as DRIFT, Basic and Lazy on top of the upstream design.

How do I install GraphRAG Local Ollama?

The README's fastest path is Docker: copy .env.example to .env and run docker-compose up --build, then drop .txt documents into ./input. For the web UI, install requirements-ui.txt with pip and run python app.py, then open http://localhost:7860.

What is the default Ollama URL used by the Docker setup?

The docker-compose.yml sets OLLAMA_HOST to http://ollama:11434 for the graphrag service, and the ollama container publishes port 11434 on the host. It also sets GRAPHRAG_API_KEY to ollama, since the pipeline talks to the local endpoint through an OpenAI-compatible client.

Can GraphRAG Local Ollama run without community reports?

Basic search does not need them: it embeds the query, retrieves the top-10 most similar entities and answers from those entities, their relationships and source text units. DRIFT is different, and the README notes that if lazy_graph_rag was enabled during indexing, the primer step will have limited context.

Official sources

  1. Issues
  2. License: MIT
  3. README
  4. TheAiSingularity/graphrag-local-ollama on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/theaisingularity-graphrag-local-ollama.svg)](https://hysenlabs.com/projects/theaisingularity-graphrag-local-ollama)