tobi/qmd: a local hybrid search engine for markdown notes and agent workflows
mini cli search engine for your docs, knowledge bases, meeting notes, whatever. Tracking current sota approaches while being all local
At a glance
- What is it?
- QMD indexes markdown, meeting transcripts and documentation on your own machine, combining BM25, vector search and LLM reranking behind one CLI. It is MIT licensed and installs from npm, but its semantic layer needs a local model and a separate embedding pass.
- Who is it for?
- Adopt tobi/qmd if your notes, meeting transcripts and docs already live as markdown on one machine and you want keyword, semantic and reranked search from a single CLI that an agent can also call over MCP. Do not adopt it if you need a hosted multi-user search service, or if you cannot run a local GGUF model for embeddings and reranking.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 20 days ago.
- What is it written in?
- Mainly TypeScript, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 28, 2026, and from our analysis. They are not legal advice.
DEEP OPEN-SOURCE ANALYSIS
The problem tobi/qmd solves, and who actually has it
Personal knowledge bases rot in a predictable way. You accumulate markdown notes, meeting transcripts and project documentation across a handful of directories, and the only search you have is whatever your editor ships or `grep`. Keyword search finds exact strings and fails on paraphrase. Semantic search finds paraphrase and fails on identifiers, error codes and file paths.
QMD's README frames the target user directly: people who want an on-device search engine for "everything you need to remember," and specifically people wiring search into agentic flows. That second audience explains a lot of the design. The CLI emits `--json` and `--files` output, and the project ships an MCP server so an agent can call search as a tool rather than shelling out and parsing text.
It is not a general document management system. There is no web UI, no multi-user access control, and no hosted component. The unit of work is a directory of markdown files on the machine where you run the command.
How the hybrid pipeline is wired
The README's flowchart is the clearest statement of the architecture. A query enters and splits three ways. The original query goes to both backends. A query expansion step produces typed sub-queries: `lex` goes to BM25 full-text search, while `vec` and `hyde` go to vector search. HyDE generates a hypothetical document to search against; the `vec` path embeds dense sentences; the `lex` path extracts BM25 keywords.
The two result sets then meet at Reciprocal Rank Fusion, which combines ranked lists without needing comparable scores. A final LLM reranker reorders the fused list before output. The README is explicit that routing is exclusive: typed expansions go only to their matching backend, and the untransformed original query is the piece that hits both.
That is a heavier pipeline than most local search tools, and the cost is visible in the command surface. `qmd search` is keyword only. `qmd vsearch` is semantic only. `qmd query` runs the full hybrid path with reranking, which the README calls "best quality." The three commands exist because the trade-off is real: the full path loads models, and the cheap path does not.
Installing tobi/qmd and running a first hybrid query
The package is published as `@tobilu/qmd` and installs globally with either npm or Bun. The README also documents running it without a global install via `npx` or `bunx`.
npm install -g @tobilu/qmd
# or
bun install -g @tobilu/qmdNext, register the directories you want searchable. Each collection gets a name, and the README's example uses `~/notes`, `~/Documents/meetings` and `~/work/docs`.
qmd collection add ~/notes --name notes
qmd collection add ~/Documents/meetings --name meetings
qmd collection add ~/work/docs --name docsThe `context add` step is the one the README singles out, describing it as "the key feature of QMD." Context attaches a description to a collection URI using the `qmd://` scheme, and the description is returned alongside matching sub-documents so a model has something to reason about beyond raw text.
qmd context add qmd://notes "Personal notes and ideas"
qmd context add qmd://meetings "Meeting transcripts and notes"
qmd context add qmd://docs "Work documentation"Before semantic search works, embeddings have to be generated. This is the step that pulls down and runs a local model, so expect it to take time on a first run.
qmd embedWith the index built, the three search modes are separate commands. `qmd search` does fast keyword matching, `qmd vsearch` does semantic matching, and `qmd query` runs the fused and reranked path.
qmd search "project timeline"
qmd vsearch "how to deploy"
qmd query "quarterly planning process"Retrieval is by path, docid or glob. Search results show a docid prefixed with `#`, which `qmd get` accepts directly.
qmd get "meetings/2024-01-15.md"
qmd get "#abc123"
qmd multi-get "journals/2025-05*.md"For agent use, `--json` returns structured results and `--files` with `--min-score` returns a file list above a threshold. The README's example pairs `--all` with `--files --min-score 0.3` for keyword search and `0.4` for the hybrid path.
The MCP server, and why the HTTP transport exists
QMD exposes four MCP tools: `query`, `get`, `multi_get` and `status`. The README positions the CLI as sufficient on its own, with MCP as the tighter-integration option. Configuration for Claude Desktop is a standard `mcpServers` block pointing at `qmd` with the `mcp` argument; Claude Code gets a plugin install path instead.
The more interesting piece is the HTTP transport. By default the MCP server runs over stdio and each client launches its own subprocess, which means each client loads models separately. `qmd mcp --http` runs a shared server on port 8181 by default, and `--daemon` backgrounds it with a PID file at `~/.cache/qmd/mcp.pid`. The README states the reason plainly: LLM models stay loaded in VRAM across requests, while embedding and reranking contexts are disposed after five minutes idle and recreated on the next request with roughly a one second penalty.
The HTTP server also exposes `POST /query` (aliased as `/search`) for structured search without speaking MCP, plus `GET /health` for liveness. Filter validation is strict: the README says invalid filters return `400`.
Origin and Host validation is the part worth reading carefully. Requests carrying an `Origin` header that is not a loopback address get `403`, and a `Host` header naming something other than the bind address is rejected too. The README explains the motive: loopback binding alone does not stop DNS rebinding, because the browser request originates from your own machine. Requests with no `Origin` header, which covers curl, MCP clients and editors, pass through. If you bind `0.0.0.0`, the host check is skipped with a startup warning, and the README notes the endpoints are unauthenticated, so authentication belongs in front of anything reachable off-host.
Where tobi/qmd gets awkward
The first constraint is that semantic search is not free. `qmd embed` is a separate pass over your collections, and the reranking path in `qmd query` runs a local model. On a laptop without a capable GPU, the full hybrid path is the slow one, which is precisely why the README offers `qmd search` as the fast alternative. If your queries are mostly exact identifiers, you are paying for machinery you will not use.
The second is operational. The HTTP transport keeps models resident in VRAM, which is the point, but it also means a long-lived process holding GPU memory. The README does not document rollback or an uninstall path for the index, so if you want to reclaim that state you are working from the repository layout rather than from instructions.
The third is a documentation gap that will bite agent integrators. The MCP parameter table notes that `query` takes `collections` as an array and that "singular `collection` is silently ignored." Silent ignore on a wrong parameter name is a poor failure mode for a tool whose main consumer is a model that may guess the singular form. The README also does not document what happens when a collection name in that array does not exist.
Finally, the metadata `filter` parameter is described as a recursive operator-discriminated JSON AST. That is a real capability, but the README excerpt does not spell out the operator set, so you will be reading the source or the linked docs section to use it.
How tobi/qmd differs from grep and from hosted search
The honest comparison is against the tools already on the machine. `grep` and ripgrep are exact-match, streaming, and have no index to build or maintain. They are faster to start and they never download a model. QMD's `qmd search` sits closest to that end of the spectrum, but it still requires collections to be registered and an index to exist.
The other direction is a hosted search service with a vector store. Those handle multiple users, remote access and persistence as a service, and they do not require you to own the hardware running the model. QMD makes the opposite bet: everything runs locally through node-llama-cpp with GGUF models, which the README lists as the runtime. The payoff is that your notes never leave the machine and there is no per-query cost. The price is that index freshness, model downloads and GPU memory become your problem.
A separate comparison is against agent frameworks that bundle retrieval. QMD stays a CLI with an MCP surface, so it composes with whatever agent you already run rather than replacing it. The `--json` and `--files` output formats are the integration contract, and the README's agent examples are built on exactly those two flags.
Licence, maintenance and upgrade cost
QMD is MIT licensed, which places few restrictions on use, modification or redistribution. The repository includes a LICENSE file at the top level. Nothing in the README suggests a separate commercial tier or a licence key, and the package is published publicly as `@tobilu/qmd`. This is a description of what the licence identifier says, not legal advice; if you are redistributing a modified build, read the LICENSE file yourself.
On maintenance: the repository is not archived, and the last push was on 2026-09-09. The most recent release listed is v2.8.3 from 2026-08-16, following v2.5.3 in May 2026 and v2.5.2 a week before that. The version history shows a project that has been moving, and the README points readers at CHANGELOG.md for progress notes.
Upgrade cost has one sharp edge. The repository contains a top-level `migrate-schema.ts`, which implies the index schema can change between versions and that upgrades may require a migration step. The README does not document that migration as part of a normal upgrade, so before bumping a version on an index you care about, check CHANGELOG.md for schema notes. Rebuilding an index is also an option, but it means re-running `qmd embed` and paying the model cost again.
Editorial conclusion
Adopt tobi/qmd if your notes, meeting transcripts and docs already live as markdown on one machine and you want keyword, semantic and reranked search from a single CLI that an agent can also call over MCP. Do not adopt it if you need a hosted multi-user search service, or if you cannot run a local GGUF model for embeddings and reranking. Before committing, verify that `qmd embed` completes on your hardware, check `qmd status` reports the collections you expect, and confirm the model download size fits your disk and VRAM budget.
Frequently asked questions
What does tobi/qmd do?
It is an on-device search engine for markdown notes, meeting transcripts, documentation and knowledge bases. It combines BM25 full-text search, vector semantic search and LLM reranking, all running locally through node-llama-cpp with GGUF models.
What are QMD files?
QMD does not introduce a file format of its own. It indexes existing markdown files, which is why the README's examples point collections at directories such as ~/notes and ~/Documents/meetings.
How do I install tobi/qmd?
Install it globally with npm install -g @tobilu/qmd or bun install -g @tobilu/qmd. The README also documents running it directly through npx @tobilu/qmd or bunx @tobilu/qmd without a global install.
Does tobi/qmd work with Claude and other AI agents?
Yes. The README documents an MCP server exposing query, get, multi_get and status tools, with configuration examples for Claude Desktop and a plugin install path for Claude Code. The --json and --files output formats are also designed for agentic workflows.
Can I run the tobi/qmd MCP server as a shared background process?
Yes. qmd mcp --http runs a shared server on port 8181 by default, and qmd mcp --http --daemon backgrounds it with a PID file at ~/.cache/qmd/mcp.pid. The README states that models stay loaded in VRAM across requests this way.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/tobi-qmd)
Community notes