# Kwipu: Graph RAG for Obsidian Vaults and Markdown Notes on Ollama

> Kwipu is a local Graph RAG engine that builds a property graph from a folder of Markdown, PDF, and DOCX files, then answers questions about them using hybrid search and citations, all through Ollama. It targets knowledge workers with Obsidian vaults or large collections of personal notes who want to query across them without sending data to a cloud API.

**benmaster82/Kwipu** — Ask questions across your Markdown notes using a fully local Graph RAG engine. Built for Obsidian vaults, works with any folder of Markdown files. Extracts entity-relation triples from wikilinks & YAML frontmatter, retrieves answers via hybrid search (vector + BM25 + temporal). Multilingual. No cloud. Runs on Ollama.

- Repository: https://github.com/benmaster82/Kwipu
- Stars: 605 · Forks: 61
- Language: Python
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/benmaster82-kwipu

## Querying Personal Notes Without Sending Them to the Cloud

The problem Kwipu solves is specific: a knowledge worker with hundreds or thousands of Markdown notes wants to ask questions across them, but does not want to upload the documents to a third-party service. Standard RAG pipelines built on cloud LLM APIs send document chunks and query context to the provider with every request. Kwipu routes all LLM inference through Ollama, which can run on the local machine.

The intended users are Obsidian vault owners and engineers with large note collections. The README explicitly states Kwipu works with 'ordinary knowledge folders' and Obsidian-style vaults interchangeably. It accepts .md, .txt, .pdf, and .docx files. The tool then provides three interfaces for querying: a terminal, a 3D web interface, and an MCP server that any MCP-compatible AI client can connect to.

## Important: The Default Model Is a Cloud Model

The README contains a caveat worth reading before starting. The default LLM is `gpt-oss:20b-cloud`, which the README describes as a model that 'can send document chunks, retrieved context, and questions to its provider.' Running Kwipu with the default configuration does not keep data on the local machine. To use a local model, the environment variable KWIPU_LLM_MODEL must be set to a local Ollama model such as `qwen2.5:7b` before indexing.

The embedding model defaults to nomic-embed-text, which runs locally. A remote Ollama endpoint sends data to the host it connects to. Kwipu requires HTTPS for non-loopback Ollama endpoints unless insecure HTTP is explicitly enabled. The README recommends against indexing sensitive documents until the intended execution mode has been confirmed.

## Installation and First Run on Windows PowerShell

The README's quick start targets Windows PowerShell. Clone the repository, create a Python virtual environment, and install the dependencies:

```powershell
git clone https://github.com/benmaster82/Kwipu.git
Set-Location .\Kwipu

py -3.12 -m venv .venv
.\.venv\Scripts\Activate.ps1
python -m pip install --upgrade pip==25.1.1
python -m pip install -r .\bridge\requirements.txt
```

If PowerShell blocks virtual-environment activation, use this process-only policy and retry:

```powershell
Set-ExecutionPolicy -Scope Process -ExecutionPolicy Bypass
.\.venv\Scripts\Activate.ps1
```

The web interface requires Node.js separately:

```powershell
npm --prefix .\frontend ci
```

Linux and macOS equivalents are documented in the README. Python 3.12 or later, Ollama, and (for the web UI) Node.js with npm are the prerequisites. The README notes that if a command is not found, reinstall the prerequisite and open a new PowerShell window before rechecking.

## Pulling Models and Setting the Environment

With Ollama running, pull the embedding model and one LLM. To run locally, pull a local model and set the LLM variable to it:

```powershell
ollama pull nomic-embed-text

# Cloud default:
ollama pull gpt-oss:20b-cloud

# OR local-only example:
# ollama pull qwen2.5:7b
```

Settings are read at process start. The README instructs users to set these environment variables in both the indexer terminal and the bridge terminal:

```powershell
$env:KWIPU_ROOT_DIR = (Get-Location).Path
$env:KWIPU_KNOWLEDGE_DIR = "knowledge_base"
$env:KWIPU_STORAGE_DIR = "storage_graph"
$env:KWIPU_EMBED_MODEL = "nomic-embed-text"
$env:KWIPU_OLLAMA_BASE_URL = "http://localhost:11434"
$env:KWIPU_QUERY_MAX_LENGTH = "4000"

# Run exactly one of these two lines:
$env:KWIPU_LLM_MODEL = "gpt-oss:20b-cloud" # cloud default
# $env:KWIPU_LLM_MODEL = "qwen2.5:7b"       # local example
```

Documents go in the knowledge_base/ folder before starting the indexer. The repository includes a small demo under knowledge_base/examples. Kwipu reads source documents but does not rewrite them.

## How Kwipu Builds the Property Graph

Kwipu extracts two types of relations from the source documents. Structural relations come from Obsidian-style wikilinks (the [[note-name]] syntax) and YAML frontmatter fields. Semantic relations are extracted by the LLM, which identifies entity-relation triples from the document text.

The combined graph stores entities, documents, and their connections in a property graph. Entity types, document metadata, and temporal information are all stored as node or edge properties. The README shows that the web interface visualizes this graph in 3D, allowing a user to explore entities, see how they connect, and trace back from an answer to its source documents.

The graph persists to the KWIPU_STORAGE_DIR directory. Kwipu watches the source folder with a file-system watcher and updates the persisted graph when documents change, without reindexing everything from scratch. The README notes this happens 'safely', meaning it preserves existing graph state while incorporating changes.

## Hybrid Search: Vector, BM25, Temporal, and MCP

Answering a query in Kwipu combines four retrieval signals. Vector search finds semantically similar chunks using the nomic-embed-text embedding model. BM25 adds keyword relevance on top of semantic similarity. Temporal metadata weights documents by recency or by date references in the query. Optional synonym retrieval is also listed as part of the hybrid pipeline.

Answers include citations that link back to the source chunks, so a user can verify what documents an answer is drawn from. The README shows the web interface surfacing these citations and making the source documents openable directly.

Beyond the terminal and web interfaces, Kwipu exposes the same knowledge through an MCP server. Any MCP-compatible client can connect to it and query the graph. The README lists an MCP server setup as one of the primary use paths. The requirements.txt includes mcp==1.27.1 as a core dependency alongside llama-index-core, watchdog, and PyYAML.

## Limitations and When Kwipu Is the Wrong Tool

Kwipu is a single-user local tool. There is no multi-user access control, no web-accessible hosted endpoint, and no managed deployment story. Teams who need multiple people to query the same knowledge base from different machines will need to set up their own infrastructure around Kwipu or use a different tool.

The quality of semantic relation extraction depends on the LLM chosen. A smaller local model like qwen2.5:7b will extract fewer and less accurate relations than a larger cloud model, which affects the depth of graph-based reasoning. For collections with little wikilink structure and thin YAML frontmatter, the graph will be sparsely connected and the advantage over a simpler flat RAG approach will be smaller.

The repository has no GitHub releases. The last push was on 2026-09-07. The README documents multilingual support for English, Italian, French, German, Spanish, and Portuguese patterns, but the quality of relation extraction in non-English notes depends on the multilingual capability of the chosen LLM.

## Comparison with AnythingLLM and Similar Tools

AnythingLLM is a commonly cited local RAG solution that also supports Ollama as a backend. The principal difference is architecture: AnythingLLM uses flat vector search over document chunks, while Kwipu builds a property graph and combines graph traversal with hybrid search. Kwipu's graph approach is better suited for note collections with rich interlinking between topics; flat vector search is often sufficient for a simpler question-answer use case with unstructured documents.

Private-GPT is another self-hosted alternative with Ollama support. It focuses on document ingestion and conversational querying without the property graph layer. Neither AnythingLLM nor Private-GPT extracts wikilinks or YAML frontmatter as structural graph edges, which is the feature that makes Kwipu specifically well-suited for Obsidian vaults.

## Conclusion

Kwipu is the right fit for individuals with an Obsidian vault or a structured folder of Markdown notes who want to query across them in natural language without relying on a cloud API, and who are willing to run Ollama locally and configure environment variables. The default model is a cloud model, not a local one; anyone who treats the 'fully local' framing as a privacy guarantee must change KWIPU_LLM_MODEL to a local Ollama model before indexing sensitive notes. Teams or organizations that need multi-user access, a managed deployment, or a vendor support contract should look at purpose-built enterprise RAG platforms instead.

## FAQ

### Does Kwipu work with Obsidian vaults?

Yes. The README states Kwipu works with Obsidian-style vaults and also with any ordinary folder of Markdown files. It extracts structural relations from Obsidian's wikilink syntax and YAML frontmatter, which makes it particularly useful for densely interlinked Obsidian vaults.

### Can Kwipu run without an internet connection?

Kwipu can run offline when configured with a local Ollama LLM and embedding model. The default LLM is gpt-oss:20b-cloud, which sends data to a cloud provider; to keep everything local, set KWIPU_LLM_MODEL to a local model such as qwen2.5:7b before indexing.

### What file types does Kwipu support?

Kwipu indexes .md, .txt, .pdf, and .docx files. It reads source documents but does not modify them. Documents are placed in the knowledge_base/ folder before starting the indexer.

## Sources

- [benmaster82/Kwipu on GitHub](https://github.com/benmaster82/Kwipu)
- [Issues](https://github.com/benmaster82/Kwipu/issues)
- [License: MIT](https://github.com/benmaster82/Kwipu/blob/main/LICENSE)
- [README](https://github.com/benmaster82/Kwipu/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/benmaster82-kwipu
