ChunkHound: local-first codebase intelligence for agents that need citations
Your entire engineering context, deeply understood
At a glance
- What is it?
- ChunkHound indexes a repository into DuckDB with Tree-sitter parsing and answers code, git history and web research questions with citations. It is aimed at teams whose agents edit code without context, but the semantic and research paths depend on external providers.
- Who is it for?
- Adopt ChunkHound if your agents or reviewers need cited context across current code and git history and you are willing to configure an embedding provider and an LLM provider. Do not adopt it if you need a stable release: pyproject.toml classifies it as Development Status 3 - Alpha, and the README does not document rollback or index migration.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository received new commits within the last day.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The context gap ChunkHound targets
The README opens with a claim rather than a feature list: agents can generate code but miss "how behavior flows across files, what changed across a branch or release, and which external constraints matter." That is the problem statement. The tool is for engineers who already run a coding agent and have noticed that the agent edits a function without knowing which other modules call it, or reviews a diff without knowing which release introduced the behaviour.
The README names four jobs: research before editing, understanding large PRs and releases, tracing bugs and incidents, and reconciling code with external documentation. Each maps to a command surface. A reviewer gets `--commit-range`, a support engineer gets `--last-n`, and someone chasing a webhook failure gets `websearch`. The audience is narrow on purpose. This is not a general search engine for a monorepo; it is a context supplier for people who are about to change, review or explain software.
How indexing and retrieval actually work
The mechanism visible in the repository is a pipeline, not a model. Files are parsed with Tree-sitter, which is why the requirements list a long tail of per-language grammars: tree-sitter-go, tree-sitter-rust, tree-sitter-hcl, tree-sitter-zig and others. Parsing produces structural chunks rather than fixed-size text windows, which is the usual reason a code search tool can answer a question about a function instead of a paragraph.
Those chunks land in DuckDB. The Python dependency pins `duckdb==1.5.4`, and there is a Rust extension, `chunkhound_native`, that links DuckDB dynamically. The Makefile explains the linkage: `DUCKDB_DOWNLOAD_LIB=1` plus an RPATH baked in at link time (`$ORIGIN` on Linux, `@loader_path` on macOS), with `scripts/copy_duckdb_runtime.py` placing the downloaded library next to the compiled extension. That is a deliberate build complexity, and the Makefile documents an air-gapped fallback using `DUCKDB_LIB_DIR` and `DUCKDB_STATIC=1`.
Retrieval has layers. Regex search works with no provider at all. Semantic search requires an embedding provider. Deep research requires both an LLM provider and an embedding provider "with reranking support." Web research uses the same provider stack as local research, so the external-docs workflow is not a separate subsystem. The pyproject keywords mention a "CAST algorithm" and vector search, but the README does not explain either, so the ranking internals are not documented at the level a reader could audit.
Installing ChunkHound and running a first question
The README requires Python 3.10 or newer and `uv`, installed with the project's own one-liner. The package itself installs as a uv tool. Note that pyproject.toml constrains Python to `>=3.10,<3.15`, so the newest interpreter is not automatically supported.
curl -LsSf https://astral.sh/uv/install.sh | sh
uv tool install chunkhoundAfter install, the README's try-it sequence is two commands: index the current directory, then ask an architecture question. The README states that regex search works without any API key, so this first pass is useful even before you configure a provider.
chunkhound index .
chunkhound research "How does authentication work?"For semantic and deep research, create `.chunkhound.json` in the project root. The README gives this example with VoyageAI embeddings and the Claude Code CLI as the LLM, which needs no key of its own.
{
"embedding": { "provider": "voyageai", "api_key": "your-key" },
"llm": { "provider": "claude-code-cli" }
}Git history questions use the same binary with a range or count. The README shows `--last-n`, `--commit-range` and `--commit-hash` as the selectors, and `websearch` as the separate command for external sources.
chunkhound research "What changed in auth recently?" --last-n 20
chunkhound research "Draft changelog bullets for billing since v2.4" --commit-range v2.4..HEAD
chunkhound search "database migration" --commit-hash abc1234Provider dependency is the real constraint
The local-first label is accurate about indexing and overstated about answering. The README says plainly that semantic search requires an embedding provider and deep research requires an LLM provider plus an embedding provider with reranking. If you want zero code egress, the README points to Ollama for embeddings, but that is a local model you now have to run and size yourself. There is no statement in the README about how large an index gets, how long indexing takes, or what hardware the local path needs.
The reranking requirement is the sharpest edge. An embedding provider that works fine for semantic search may not satisfy the deep research path, and the README does not list which providers qualify. A team that picks a provider on price and then discovers reranking is unsupported has to re-index with a different embedding model, because embeddings from different models are not interchangeable. The README does not document rollback or re-indexing procedure, so that cost is invisible until it is incurred.
The project also carries a `Development Status :: 3 - Alpha` classifier in pyproject.toml, and the README badge states the codebase is 100% AI generated. Both are worth reading as signals about API stability rather than quality. The presence of MIGRATION_GUIDE.md at the repository root suggests breaking changes have already happened at least once.
ChunkHound versus a general codebase indexer
The closest comparison in the search data is against Sourcegraph-style codebase indexing, and the difference is architectural. A hosted code search product indexes a repository once on its servers and serves results to a browser; ChunkHound runs on the developer's machine, writes to a local DuckDB file, and exposes itself over MCP so an agent can call it as a tool. The trade is reach against locality. You give up organisation-wide search across every repository and you get an index that never leaves the machine and a query interface an agent can drive.
The second alternative is plain grep plus an agent's own file reading. Regex search in ChunkHound works without providers, which puts it in the same class as grep, but the git-history selectors (`--commit-range`, `--last-n`, `--commit-hash`) are the part grep does not offer at all. If your questions are always about the current working tree, grep is simpler and has no index to maintain. If they are about what changed between a tag and HEAD, ChunkHound is doing something grep cannot.
Maintenance, upgrades and licence
The last push to the default branch was on 2026-09-09, and the most recent release is v5.2.1 from 2026-07-12, following v5.2.0 the same day and v5.1.0 on 2026-05-20. The repository is not archived. The version cadence in the repository shows two minor releases between May and July 2026, which is a fast enough pace that pinning a version matters more than usual for a tool that writes an index you may have to rebuild.
Upgrade cost concentrates in two places. The DuckDB pin is exact (`duckdb==1.5.4`), and the Rust extension links DuckDB dynamically, so a DuckDB upgrade is a build change, not a pip bump. The embedding model behind your index is the second: changing it invalidates stored vectors. MIGRATION_GUIDE.md exists at the root, and reading it before a major upgrade is the cheapest step available.
The licence is MIT, declared both in the README badge and in pyproject.toml classifiers. That permits commercial and closed-source use. It does not cover the providers you connect to: your embedding and LLM calls are governed by VoyageAI, OpenAI, Anthropic or whichever service you configure, and those terms are separate from the MIT grant on the code. This is not legal advice; check the provider terms for your own situation.
Editorial conclusion
Adopt ChunkHound if your agents or reviewers need cited context across current code and git history and you are willing to configure an embedding provider and an LLM provider. Do not adopt it if you need a stable release: pyproject.toml classifies it as Development Status 3 - Alpha, and the README does not document rollback or index migration. Verify first that your embedding provider supports the reranking the deep research path needs, and check the MIGRATION_GUIDE.md before upgrading across a major version.
Frequently asked questions
What is the ChunkHound MCP server?
ChunkHound exposes its codebase intelligence over the Model Context Protocol so that agents can call it as a tool. The pyproject.toml describes the package as local-first codebase intelligence for AI assistants via MCP, and the MCP dependency is declared directly rather than only pulled in transitively.
Does ChunkHound work with Claude Code?
The README's configuration example sets the LLM provider to claude-code-cli, and the requirements section lists Claude Code CLI and Codex CLI as LLM options that need no API key. You still need an embedding provider for semantic search.
Can I use ChunkHound without any API keys?
Partly. The README states that regex search works without any providers, so `chunkhound index .` and plain search are usable with no keys. Semantic search needs an embedding provider, and deep research needs an LLM provider plus an embedding provider with reranking support.
Which languages does ChunkHound support?
Parsing is done with Tree-sitter, and requirements.txt lists grammars for Python, JavaScript, TypeScript, Java, Go, Rust, C, C++, C#, PHP, HCL, Bash, Kotlin, Lua and Zig among others. The README summarizes this as Python, JavaScript, TypeScript, Java, Go, Rust, C/C++ and more.
Does ChunkHound search git history as well as current code?
Yes. The README shows `chunkhound search` and `chunkhound research` accepting `--last-n` for the last N commits, `--commit-hash` for a specific commit, and `--commit-range` for a range such as `v2.4..HEAD` or `main..HEAD`.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/chunkhound-chunkhound)