Model or dataset
ggozad/haiku.rag avatar
ggozad/haiku.rag

haiku.rag: local-first agentic RAG with citations, on embedded LanceDB

Agentic RAG for local and self-hosted document search: hybrid retrieval, reranking and multimodal RAG on embedded LanceDB, with Docling parsing and an MCP server

615 stars51 forksPythonMIT

At a glance

What is it?
haiku.rag is a Python library and CLI that indexes your own documents into an embedded LanceDB store and answers questions with page-level citations. It is local-first, MIT licensed, and ships an MCP server, but the search questions Google returns for its name are about an unrelated Anthropic model.
Who is it for?
Adopt haiku.rag if you want agentic RAG over your own documents without running a database server, and you are on Python 3.12 or newer with an embedding provider such as Ollama or OpenAI already available. Skip it if you need a hosted, multi-tenant retrieval service with an SLA, or if you will not commit to the provider setup the quick start requires.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository received new commits within the last day.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The problem haiku.rag targets: citations over your own files, no database server

Most RAG setups ask you to stand up a vector database, wire an ingestion pipeline to it, and then reconcile the retrieved chunks with the original document so an answer can point somewhere. haiku.rag collapses that into a Python package. The README states the goal plainly: answer questions about your own documents with citations to page numbers and section headings, running locally on an embedded database with no server required.

The audience is narrow and identifiable. It is a developer who already has an embedding provider available (Ollama, OpenAI, LM Studio, vLLM, VoyageAI or Cohere) and wants retrieval over PDFs and other documents without operating infrastructure. The package requires Python 3.12 or newer, and the classifiers list Windows 10/11, macOS and Linux. It is published on PyPI as haiku.rag, with a slimmer haiku.rag-slim base that the full package depends on.

One caveat about the name: the search questions Google associates with "haiku" here are about an unrelated Anthropic model, not this library. The project's own README and pyproject are the reliable source for what it does.

How retrieval works: Docling parsing, LanceDB storage, RRF fusion, optional reranking

The pipeline has three visible stages. Docling parses documents, and the README says the full DoclingDocument is stored, which is what makes structure-aware context expansion possible. Chunks and their vectors land in LanceDB, an embedded store that also backs the optional S3, GCS, Azure and LanceDB Cloud targets.

Retrieval is hybrid: vector search plus full-text search, combined with Reciprocal Rank Fusion. That fusion step is the reason a query like "attention mechanism" can match both semantically similar passages and exact term occurrences, rather than forcing you to pick one. Reranking is a separate, pluggable stage: local cross-encoders, Cohere, Zero Entropy or vLLM.

On top of retrieval sit two capabilities worth naming. Vision QA passes figure bytes to vision-capable models alongside chunk text, and multimodal embedders (vLLM, VoyageAI, Cohere, enabled with multimodal: true) place picture vectors in the same space as text, so a text query can return figure hits. The analysis capability runs sandboxed Python for aggregation and computation across documents. Two optional capabilities address long conversations: evidence compaction replaces earlier search results with the evidence they cited, and a citation policy requires every answer to declare what grounds it, including declaring that nothing does.

Installing haiku.rag and running a first indexed query

The README requires Python 3.12 or newer. The full package is the recommended install because it includes document processing, all embedding providers and rerankers. With uv, the README gives the pip-compatible form.

bash
pip install haiku.rag

If you prefer a smaller footprint, install haiku.rag-slim instead and add only the extras you need. The installation documentation lists the available options; the pyproject shows extras named tui, s3, cross-encoder and ingester.

bash
pip install haiku.rag-slim

Before the first query you need an embedding provider configured. The README points to the tutorial for provider setup and notes that Ollama or OpenAI are typical choices. Once that is in place, index a document and search it.

bash
haiku-rag add-src paper.pdf
haiku-rag search "attention mechanism"

The add-src command parses the PDF and writes chunks into the LanceDB store. The search command returns matching chunks; because fusion is in play, expect a mix of semantic and keyword matches rather than a single ranking signal. For an answer with provenance, use ask, which returns text plus citations to page numbers and section headings.

bash
haiku-rag ask "What datasets were used for evaluation?"

From Python the same flow is available through an async context manager, which is the shape to copy if you are composing your own agent.

python
from haiku.rag.client import HaikuRAG

async with HaikuRAG("knowledge.lancedb", create=True) as rag:
    await rag.create_document_from_source("paper.pdf")
    results = await rag.search("self-attention")
    answer, citations = await rag.ask("What is the complexity of self-attention?")

The client prints scores alongside page numbers and content, which is the practical way to sanity-check that retrieval is finding the right passages before you trust the answers.

Where haiku.rag is the wrong choice

The embedded design is the selling point and also the constraint. LanceDB runs in-process, so there is no separate service to scale horizontally, and the README does not describe a multi-tenant access model. If several teams need isolated corpora behind one endpoint with quotas and audit logs, you are building that layer yourself.

Provider setup is a hard prerequisite, not an optional extra. The quick start explicitly notes that an embedding provider is required and points to the tutorial. There is no documented path to a working index without one, and no bundled local model is described. If your environment cannot reach Ollama, OpenAI or an equivalent, the tool does nothing useful.

Multimodal search is narrower than the feature list suggests. Cross-modal retrieval depends on multimodal embedders, and the README names only vLLM, VoyageAI and Cohere for that path. Choose a text-only embedder and image-as-query is off the table.

The project's own classifiers mark it as Development Status 4 - Beta, and the version in pyproject is 0.85.0. That is a pre-1.0 surface, and no API stability guarantee across releases is stated. The README also does not document rollback beyond the tags feature, which names database states with haiku-rag tag, so treating tags as your undo mechanism is an assumption worth testing on a copy before you rely on it.

haiku.rag compared with a hosted RAG service or a hand-rolled stack

The obvious alternative is a managed retrieval service, where you upload documents to someone else's index and query it over HTTP. The difference is where the data and the failure modes live. With haiku.rag the index is a local LanceDB directory, the parsing happens on your machine through Docling, and the answer is produced by whatever model you configure. Nothing leaves your environment except the calls to the embedding and QA providers you chose. A hosted service inverts that: less setup, but your documents leave your network and your retrieval quality is whatever the vendor ships.

The second alternative is assembling the stack yourself: Docling for parsing, LanceDB for storage, a Pydantic AI agent for the loop, plus your own fusion and reranking code. That gives you full control and full responsibility for the glue. haiku.rag's contribution is precisely that glue, plus the pieces that are easy to get wrong: Reciprocal Rank Fusion, citation plumbing that carries page numbers and section headings, evidence compaction for long conversations, and the MCP server that exposes search and document reading as tools to assistants such as Claude Desktop, Claude Code and Codex. If you have already written and validated that glue, the library's value drops sharply. If you have not, it is the part you would otherwise spend weeks on.

Maintenance, releases and what the MIT licence means here

The repository is not archived, and the last push was on 2026-09-10. Releases are frequent: 0.82.1 on 2026-09-03, 0.83.0 on 2026-09-09 and 0.84.0 on 2026-09-10, with pyproject carrying version 0.85.0. The cadence means upgrades are a real cost. The full haiku.rag package pins its slim dependency to an exact version, so a minor bump moves both packages together, and the extras in pyproject are pinned the same way. Check the changelog link in pyproject before upgrading a production index.

Licensing is MIT, declared in pyproject as license text and present as a LICENSE file at the repository root. That is permissive and imposes no copyleft obligation on your own code. It says nothing about the terms of the embedding and QA providers you point it at, and nothing about the licences of the models you run through Ollama, vLLM or a hosted API. Those are separate agreements, and this article gives no legal advice on them.

The operational surface is larger than a library usually implies. The README describes a production ingester, haiku-ingester, as a long-lived service with a persistent SQLite queue, an async worker pool with retries, a dead-letter queue, FS, HTTP, S3 and WebDAV source adapters, a FastAPI control plane and a browser dashboard. Running that in production means owning a queue, a database file and a service process, which is a different maintenance profile from calling a library from a script.

Exposing haiku.rag to an AI assistant over MCP

The MCP server is the feature that changes how the tool gets used day to day. The README shows a stdio server started from the CLI.

bash
haiku-rag mcp --stdio

For Claude Desktop, the configuration is a JSON block naming the command and arguments. The README gives this example.

json
{
  "mcpServers": {
    "haiku-rag": {
      "command": "haiku-rag",
      "args": ["mcp", "--stdio"]
    }
  }
}

Claude Code and Codex install a plugin from the project's marketplace instead, which the README says registers the server and a skill.

bash
claude plugin marketplace add ggozad/haiku.rag
claude plugin install haiku-rag

The server provides search, document reading and analysis tools inside the assistant. That is a meaningful difference from pasting context into a chat window: the assistant retrieves from your indexed corpus on demand, and the analysis tool can run code rather than guess at arithmetic. The trade-off is that the assistant now depends on your local index being populated and your provider being reachable, and the README does not describe what the tools return when the index is empty.

Editorial conclusion

Adopt haiku.rag if you want agentic RAG over your own documents without running a database server, and you are on Python 3.12 or newer with an embedding provider such as Ollama or OpenAI already available. Skip it if you need a hosted, multi-tenant retrieval service with an SLA, or if you will not commit to the provider setup the quick start requires. Before you commit, verify that your chosen embedder is supported for the modality you care about (multimodal requires vLLM, VoyageAI or Cohere), and check the installation docs for the extras list, since the slim package installs only what you name.

Frequently asked questions

What is haiku.rag best used for?

It is built for question answering over your own documents with citations to page numbers and section headings, running locally on an embedded LanceDB database with no server required. The README also lists hybrid search, multimodal retrieval, an analysis capability that runs sandboxed Python, and an MCP server for AI assistants.

Does haiku.rag need a separate vector database server?

No. The README states it runs locally on an embedded database, LanceDB, with no server required, and it also supports S3, GCS, Azure and LanceDB Cloud as storage targets.

What Python version does haiku.rag require?

The README states Python 3.12 or newer is required, and pyproject sets requires-python to >=3.12 with classifiers for Python 3.12 and 3.13.

Can haiku.rag answer questions about images as well as text?

Yes, with the right models. Vision-capable models receive figure bytes alongside chunk text, and multimodal embedders on vLLM, VoyageAI or Cohere, enabled with multimodal: true, put picture vectors in the same space as text so an image can be used as a query.

How do I use haiku.rag with Claude Desktop?

Run haiku-rag mcp --stdio and add an mcpServers entry to your Claude Desktop configuration naming the haiku-rag command with args ["mcp", "--stdio"], as shown in the README. The README also describes installing a plugin for Claude Code and Codex from the project's marketplace.

Is haiku.rag the same as Anthropic's Claude Haiku model?

No. haiku.rag is a Python library for local and self-hosted document search, published on PyPI and licensed MIT. The questions Google returns about Claude Haiku refer to an unrelated Anthropic model.

Official sources

  1. ggozad/haiku.rag on GitHub
  2. License: MIT
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/ggozad-haiku-rag.svg)](https://hysenlabs.com/projects/ggozad-haiku-rag)