Probe: ripgrep speed with a tree-sitter reading of the file
AI-friendly semantic code search engine for large codebases. Combines ripgrep speed with tree-sitter AST parsing. Powers AI coding assistants with precise, context-aware code understanding.
At a glance
- What is it?
- Probe is a Rust code and markdown context engine that skips embeddings entirely and answers Elasticsearch-style boolean queries against tree-sitter parse trees. Its README is confident about the approach and vague about itself in three places: two conflicting licenses, a provider list that disagrees with the agent options table, and a build that copies a binary it never compiles.
- Who is it for?
- Probe is worth trying on a large repository where an agent keeps re-reading the same files, because retrieval quality is the part of that loop it actually changes, and it changes it without asking anyone to host an embedding service. What to check first is the license, since the Cargo.toml package block and the license recorded for the repository say different things, and nothing on the page resolves it.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 2 days ago.
- What is it written in?
- Mainly Rust, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 3, 2026, and from our analysis. They are not legal advice.
Editorial analysis
Three names for one tool, and two different licenses
The package manifest is where the naming starts to diverge:
name = "probe-code"
version = "0.6.0"
rust-version = "1.88.0"
license = "MIT"
default-run = "probe"The crate is `probe-code`, the npm package the quick start installs is `@probelabs/probe`, and the binary users type is `probe`. The `default-run` line exists precisely because the first two do not match the third, so a bare `cargo run` needs telling which of the targets to launch.
The license is a harder problem, because two sources disagree. That manifest block says MIT. The license recorded for the repository says Apache-2.0. Both are written down, neither page comments on the difference, and the repository has a `LICENSE` file at its root. Anyone shipping Probe inside a product needs to resolve that before release rather than after.
The manifest also fixes the toolchain floor at Rust 1.88.0 with edition 2021, which is a useful detail given the recent release tags all sit on the 0.6.0 release candidate line.
Twenty tree-sitter grammars, four of them missing from the language list
Structural search is the whole premise, so the grammar list is the architecture. The dependency block carries twenty tree-sitter crates: rust, javascript, typescript, python, go, c, cpp, java, ruby, php, swift, c-sharp, html, md, yaml, bash, qmljs, solidity, crystal and haskell.
The language bullet on the landing page names a shorter set: Rust, Python, JavaScript, TypeScript, Go, C/C++, Java, Ruby, PHP, Swift, Solidity, Crystal, C#, Bash and QML, followed by the word more. That list leaves out Haskell, HTML, Markdown and YAML, all of which are compiled in. For a tool whose selling point is reading code as structure, the four grammars it silently includes are the ones worth knowing about.
Three of the pins are also stricter than the rest:
tree-sitter-c-sharp = { version = "=0.23.1" }
tree-sitter-solidity = "=1.2.10"
tree-sitter-crystal = { git = "https://github.com/crystal-lang-tools/tree-sitter-crystal", rev = "f71f4ca62ac0" }C# and Solidity are held to an exact patch, and Crystal does not come from a registry at all: it is a git dependency pinned to one revision of a community toolchain. That last one has a practical consequence, since a build of the repository depends on a third-party URL resolving to exactly that commit.
The agent options table ends inside the --model row
The page runs out mid-table. The agent options section introduces the available flags and reaches three rows before the value cell of `--model` stops after its opening backtick. What survives is `--path <dir>` for the search directory, defaulting to the current one, and `--provider <name>` with the values `anthropic`, `openai` and `google`. Everything the page promised after the model flag, including the rest of the table of contents entries for the script mode, installation, supported languages, documentation and environment variables, sits beyond that point.
The provider list is where the visible part contradicts itself. The features section calls the agent multi-provider across Anthropic, OpenAI, Google and Bedrock. The options table names three, and Bedrock is not one of them.
Two smaller mismatches sit in the same neighbourhood. The quick start says the agent works with any model through your own key, giving `GOOGLE_API_KEY` as the example. The comparison table names the agent integration as a full agent loop plus MCP plus the Vercel AI SDK, while the table of contents calls the same section a Node.js SDK, and the repository has an `npm/` directory at its root to go with it.
The container image copies a binary it never builds
The Dockerfile is worth reading as a release artifact rather than a build recipe. It starts `FROM gcr.io/distroless/cc-debian12`, takes build arguments for version, build date, VCS ref and target architecture, stamps a block of OpenContainers labels, and then does the only interesting thing:
COPY binaries/${TARGETARCH}/probe /usr/local/bin/probeThere is no `cargo build` in the file. The image expects a binary that something else already produced and staged under `binaries/<arch>/`, which means a bare `docker build` has nothing to copy. `TARGETARCH` is supplied by docker buildx, so a build that does not go through buildx has no architecture to interpolate either.
The rest of the file is written for a stripped base. Distroless means no shell and no package manager, the image runs as non-root by default, and the health check is therefore the binary itself, `HEALTHCHECK` invoking `/usr/local/bin/probe --version` on a thirty second interval. The compose file mounts the repository read-only at `/workspace` and sets the container command to `--help`, so the default `docker compose up` starts a service that prints usage text and exits rather than serving anything.
Compose forwards two API keys and one variable the page never explains
The compose file defines three services. The first is the search binary described above. The other two are the chat example from `examples/chat`, one as a CLI with `stdin_open` and `tty` enabled and one as a web service publishing port 3000 and running with `--web`, checked by a health probe that curls `http://localhost:3000/health` every thirty seconds with a five second start period.
Both chat services receive the same three variables: `ANTHROPIC_API_KEY`, `OPENAI_API_KEY` and `ALLOWED_FOLDERS`. The first two are self-explanatory. The third is not, because nothing visible on the page names it, explains its format, or says what happens when it is empty, and the compose default is exactly that, an empty string. A filesystem allowlist wired into an agent that can read a mounted repository deserves a documented answer before anyone runs it.
The same gap appears in the other direction. The features section promises a Bedrock provider and the quick start suggests a Google key, but neither service is given a variable for either. Whatever provider support exists beyond Anthropic and OpenAI is not reachable through the compose file as written.
The Makefile still defaults its version to v0.1.0
Release plumbing is declared in a Makefile at the root, and its version default has not moved in a long time:
VERSION ?= v0.1.0
RELEASE_DIR := release/$(VERSION)
BINARY_NAME := probe
release: clean-release version linux macos windowsThe package version is 0.6.0 and the newest tags are on the 0.6.0 release candidate line, so anyone running `make release` without overriding `VERSION` produces artifacts labelled v0.1.0 from 0.6.0 sources, in a directory named after the old number. The override is one variable, but nothing in the repository warns about it.
The rest of the target set is more careful. Four platform targets are named, with macOS split into an x86 target and an arm64 target so both slices ship, and the Linux target prints a reminder that the cross target may need adding through rustup first. Windows is handled by detecting `OS=Windows_NT` and prefixing commands through cmd.exe so `RUST_BACKTRACE=1` actually takes effect. There is also an `install` target that installs from the working directory rather than the registry, and a separate `build-windows.bat` at the root for the same platform from the other direction.
The 10x claim meets an empty benches directory
The opening lines make two numbers claims. Code is read ten times more often than it is written, and one Probe call captures what other tools need ten or more agentic loops for. Neither figure is supported anywhere on the page, and there is no measurement section behind them.
The repository does have a `benches/` directory at its root, and the workspace manifest also carries `examples/reranker` and `examples/reranker/rust_bert_test` as members, so the machinery for measuring exists in the tree. What is missing is the output. Nothing on the page says what is benchmarked against what, on which repository, or with which query set, which leaves the two numbers as framing rather than as evidence.
The ranking story has the same shape. A comparison row credits BM25, TF-IDF and hybrid scoring with SIMD acceleration, and the feature list adds optional BERT reranking on top. That reranker is the one part of the ranking path that is a model rather than arithmetic, and it ships as two workspace crates rather than as an optional feature flag. Worth knowing before treating the deterministic claim, same query returning the same results, as covering the whole retrieval stack.
The query language promises NOT, then lists +required and -excluded
The worked example is the clearest statement of intent on the page:
User: "find the authentication logic"
-> LLM generates: probe search "verify_credentials OR authenticate OR login OR auth_handler"
-> Probe returns complete AST blocks in millisecondsThe argument for this design is stated directly: embedding tools solve vocabulary mismatch, finding authentication when the code says `verify_credentials`, but an agent is already the thing that handles that, so it can write the disjunction itself. The agent runs three or four searches with session dedup rather than one embedding query.
The operator vocabulary is where the prose disagrees with itself. The feature list advertises Elasticsearch-style queries with `AND`, `OR`, `NOT`, phrases and filters. The sentence that actually enumerates the language gives `AND`, `OR`, `+required`, `-excluded`, exact phrases, `ext:rs` and `lang:python`, with `NOT` absent from it. So the one operator advertised in the summary is the one never demonstrated. The file filters do appear in the commands, as an extension qualifier and as a separate `--language rust` flag on the pattern query:
npx -y @probelabs/probe search "authentication AND login" ./src
npx -y @probelabs/probe extract src/main.rs:42
npx -y @probelabs/probe query "fn $NAME($$$) -> Result<$RET>" --language rustThat last line is the structural matcher, with metavariables standing in for a name, an argument list and a return type, which is why the AST dependency exists alongside tree-sitter.
Editorial conclusion
Probe is worth trying on a large repository where an agent keeps re-reading the same files, because retrieval quality is the part of that loop it actually changes, and it changes it without asking anyone to host an embedding service. What to check first is the license, since the Cargo.toml package block and the license recorded for the repository say different things, and nothing on the page resolves it. Second, check the build: the container image copies a prebuilt binary out of binaries/ instead of compiling, so a plain docker build will not produce anything usable unless the release tooling has already staged it. Third, if you plan to use the agent rather than the search, note that its documented options table is incomplete and that its provider list does not match the feature list. For search alone, the deterministic BM25 path is the part that is fully documented.
Frequently asked questions
What is Probe, and does it need an index or an embedding model?
Probe is a Rust code and markdown context engine that answers Elasticsearch-style boolean queries with BM25 ranking against tree-sitter parse trees, using ripgrep for scanning and ast-grep for pattern matching. The comparison table sets setup time at none, external dependencies at none, and works offline at always, because search makes no API call and the code never leaves the machine.
Which languages can Probe parse, and is that list complete?
The manifest compiles twenty tree-sitter grammars: rust, javascript, typescript, python, go, c, cpp, java, ruby, php, swift, c-sharp, html, md, yaml, bash, qmljs, solidity, crystal and haskell. The language list on the page names fifteen and omits Haskell, HTML, Markdown and YAML, even though those grammars are present.
What license is Probe released under?
Two different answers are written down. The Cargo.toml package block declares the license as MIT, while the license recorded for the repository says Apache-2.0, and a LICENSE file sits at the root. The page does not explain which one covers the distributed binary.
How do I install Probe and run a search from the terminal?
Through npm, as @probelabs/probe, either registered as an MCP server with the args agent --mcp or mcp, or invoked directly with commands such as npx -y @probelabs/probe search "authentication AND login" ./src, extract src/main.rs:42, or query with a pattern and --language. A Rust build from source needs Rust 1.88.0 per the manifest.
Does Probe's built-in agent need its own API key?
Not when it piggybacks on Claude Code or Codex authentication, which is the recommended route and is why no extra key is needed for that setup. Otherwise the agent accepts your own key, with GOOGLE_API_KEY given as the example, and supports Anthropic, OpenAI, Google and Bedrock, with retry, fallback and context compaction.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/probelabs-probe)