Model or dataset
CodeBendKit/codeseek avatar
CodeBendKit/codeseek

CodeSeek: a Rust call-graph and hybrid search CLI for Claude Code and Codex

Rust-powered code intelligence CLI for AI coding agents. Builds call graphs and hybrid semantic search indexes (Dense + Sparse + RRF + Reranker) across 7 languages. Ships as native MCP tools for Claude Code and Codex CLI.

768 stars44 forksRustMIT

At a glance

What is it?
CodeSeek indexes a repository with Tree-sitter, LanceDB, Tantivy and a cross-encoder reranker, then exposes the result as MCP tools. It is a small, MIT-licensed project aimed at agent workflows, and the setup wizard is the first thing you meet.
Who is it for?
Adopt CodeSeek if you already drive Claude Code or Codex CLI over a repository and want symbol search plus caller/callee tracing available as MCP tools without writing the indexer yourself. Skip it if you need Windows support, a fully offline embedding stack, or a project whose release cadence you can plan against: the last push was on 2026-08-02 and the latest release, v0.1.31, dates from 2026-07-29.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 60 days ago.
What is it written in?
Mainly Rust, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What CodeSeek solves for agent-driven codebases

An agent asked to change a function usually greps for the name, reads a few files, and guesses at the blast radius. CodeSeek replaces that guess with two indexes built ahead of time: a call graph extracted from the AST, and a hybrid search index over function and class definitions. The README frames the tool as a "code intelligence CLI tool for Claude Code", and the command set follows that framing. `codeseek callers main` lists the functions that call `main`; `codeseek callees process_data` lists what `process_data` calls; `codeseek callgraph apply_rerank` walks both directions with a configurable depth.

The audience is narrow on purpose. This is for engineers who already run an agent against a repository and want the agent to answer structural questions without reading the whole tree. It is not an editor plugin, not a language server, and not a replacement for `grep` when you already know the exact string you want. The value shows up when the question is semantic ("how does the code embedding work") or structural ("who calls this before I change its signature").

How CodeSeek builds its index and fuses search results

Indexing runs through a fixed pipeline. Source files go through a Tree-sitter AST parse covering 7 languages, functions, classes and methods are extracted, texts are embedded in batches of 20 per API call with a SQLite cache in front, vectors land in LanceDB, a BM25 index is built in Tantivy, and the call graph is serialized as something the README calls PetCodeGraph. Everything is written under `~/.codeseek/<project_hash>/`. The process is idempotent in the specific sense that the first run is a full build and later runs compare MD5 hashes and reprocess only changed files.

Query time fans out to three retrievers: LanceDB approximate nearest neighbour search over dense vectors, Tantivy BM25 over tokens, and a graph lookup. The three result lists are merged with Reciprocal Rank Fusion into a Top-20 candidate set, and a cross-encoder reranker (the README names Qwen3-Reranker) scores each query/code pair to produce the final Top-5. The README's own table marks dense and sparse search as fast and does not put a figure on the reranker stage, which is the part that adds a model call per candidate.

The `search` command is described as falling back from vector search to graph name matching, so a query that embeds poorly can still resolve if the symbol name matches something in the graph. That fallback is worth knowing about because it means a poor embedding configuration degrades search rather than breaking it outright.

Installing CodeSeek and running a first search

The npm package is the documented path. It ships a JavaScript wrapper that runs the setup wizard, downloads the matching Rust binary from GitHub Releases, and forwards every subcommand to that binary. Node 20 or later is required per package.json.

bash
npm install -g codeseek

Running the bare command starts the interactive wizard, which asks for an embedding API token, a model, and a base URL. Nothing can be indexed until that wizard has run, because the embedding endpoint is the only source of dense vectors.

bash
codeseek
codeseek init

The first `codeseek init` is a full build; later runs are MD5-incremental. When it finishes, `codeseek status` reports functions, files and the last update time, which is the quickest way to confirm the parse actually found your code rather than an empty tree.

bash
codeseek status
codeseek search main --limit 10
codeseek callers main

The README shows a natural-language query returning ranked symbols with scores and file paths, for example `get_embedding (0.7973)` followed by its path in `rust-core/src/services/embedding_service.rs`. Note the `:0` line numbers in that output: the sample shows line 0 for every hit, so do not expect precise line positions from search results.

To wire it into an agent, run the install command, which writes MCP configuration for Claude Code and Codex CLI.

bash
codeseek install
codeseek install-hooks

The second command installs post-commit and post-merge git hooks that call `codeseek init`, so the index follows your commits. `codeseek uninstall` removes the MCP integration and `codeseek uninit` deletes the current project index. Every query command accepts `--json`.

Where CodeSeek is the wrong tool

The embedding dependency is the sharpest constraint. The wizard wants an API token and a base URL, and the pipeline batches 20 texts per call. There is no documented local embedding option, so a repository under an air-gapped or no-egress policy cannot be indexed with the documented setup. The SQLite cache reduces repeat calls across incremental runs but does not remove the dependency.

Platform coverage is limited. The README lists macOS arm64, macOS x64 and Linux x64. Windows is absent, and the Homebrew tap is a second option rather than a fallback for unsupported platforms. If your team develops on Windows, the documented install paths do not cover you.

There is also a scale question the README does not answer. Incremental indexing is MD5-based per file, which is cheap for edits, but a rebase, a branch switch or a large merge can invalidate many files at once and push the next `init` back toward a full rebuild with its API cost. The `install-hooks` command triggers on post-commit and post-merge, which is exactly the moment a large merge lands. The README does not document a cost estimate, a rate limit, or a dry-run mode for `init`.

Finally, if your question is lexical and precise, this is more machinery than you need. Searching for a string literal or a config key is `grep`'s job; CodeSeek indexes definitions, not arbitrary text.

CodeSeek compared with plain ripgrep plus an agent

The realistic alternative is not another code intelligence server but the baseline most teams already run: ripgrep for lookup, plus an agent that reads files on demand. That combination has no index, no embedding endpoint, no daemon and no staleness problem. It also has no call graph, so an agent asked "what breaks if I change this signature" has to reconstruct the answer by reading and reasoning, which is where it burns context and where it guesses.

The difference in approach is retrieval versus reasoning. CodeSeek precomputes structure and similarity so the agent receives a ranked list of candidates with paths; ripgrep returns exact line matches and leaves the ranking to the model. A hybrid search stack with RRF and a cross-encoder is a meaningful amount of machinery over a plain index, and the README's sample output is the honest illustration of what it buys: a query phrased as a question resolves to `get_embedding` at 0.7973, well clear of the second hit at 0.2855. If your queries are already symbol names, that gap does not matter and the simpler tool wins.

MCP registration, licence and upgrade cost

`codeseek install` writes MCP server configuration to `~/.claude.json` (global, all projects) or `./.mcp.json` (project-local) for Claude Code, and to `~/.codex/config.toml` for Codex CLI. The README states that Claude Code auto-discovers the tools after a restart and lists five: `codeseek_search`, `codeseek_callers`, `codeseek_callees`, `codeseek_callgraph` and `codeseek_status`. The server itself runs as `codeseek serve --mcp` over stdio JSON-RPC, which is what Claude Code invokes internally. Writing to the global `~/.claude.json` affects every project on the machine, so the project-local `./.mcp.json` is the lower-blast-radius choice when you are evaluating.

Upgrade cost is shaped by the release history: v0.1.29 on 2026-07-08, v0.1.30 on 2026-07-10, v0.1.31 on 2026-07-29, and the last push on 2026-08-02. Three patch releases in three weeks, then no pushes for roughly seven weeks as of today. That is a young project at a 0.1.x version, and the npm wrapper downloads a platform binary from GitHub Releases on install, so an upgrade re-fetches the binary. The npm package declares a `preuninstall` script that only prints a reminder to run `codeseek uninstall-hooks`, which means removing the package does not remove the git hooks for you.

On licensing: the repository is MIT, and package.json declares `"license": "MIT"`. MIT is permissive and places few obligations on internal use, but the embedding model and reranker the wizard configures are separate components with their own terms, and the README does not discuss those terms. That is a question for whoever owns your model procurement, not something this article can settle.

Editorial conclusion

Adopt CodeSeek if you already drive Claude Code or Codex CLI over a repository and want symbol search plus caller/callee tracing available as MCP tools without writing the indexer yourself. Skip it if you need Windows support, a fully offline embedding stack, or a project whose release cadence you can plan against: the last push was on 2026-08-02 and the latest release, v0.1.31, dates from 2026-07-29. Before committing, verify that your embedding endpoint answers the wizard, that `codeseek status` reports a plausible function count for your repository, and that `codeseek install` wrote the config file you expect rather than the global one.

Frequently asked questions

What is CodeSeek and what does it index?

CodeSeek is a Rust CLI that builds an AST-based call graph and a hybrid semantic search index over your repository, then exposes them as MCP tools for Claude Code and Codex CLI. It parses 7 languages with Tree-sitter and stores vectors in LanceDB and a BM25 index in Tantivy.

How do I install CodeSeek?

The documented path is `npm install -g codeseek`, which requires Node 20 or later. The package runs a setup wizard, downloads the correct Rust binary for your platform, and forwards all subcommands to it. Homebrew and a source build via `./build.sh --release` are also documented.

Which platforms does CodeSeek support?

The README lists macOS arm64, macOS x64 and Linux x64. Windows is not listed among the supported platforms.

Does CodeSeek need an API key to work?

Yes. The first-run wizard prompts for an embedding API token, a model and a base URL, and indexing batches texts to that endpoint 20 at a time with a SQLite cache in front. The README does not document a local or offline embedding option.

How do I remove CodeSeek's git hooks?

Run `codeseek uninstall-hooks`. The npm package's preuninstall script only prints a reminder to do this, so uninstalling the package does not clean up the hooks that `codeseek install-hooks` added.

Official sources

  1. CodeBendKit/codeseek on GitHub
  2. Issues
  3. License: MIT
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/codebendkit-codeseek.svg)](https://hysenlabs.com/projects/codebendkit-codeseek)