Open-source project
lajosdeme/mole avatar
lajosdeme/mole

mole reserves every model call against a budget it keeps in the database schema

A deep-research agent with an enforced budget, verified quotes, and a privacy boundary for local data.

314 stars18 forksGoApache-2.0

At a glance

What is it?
A single Go binary that decomposes a question, reads sources, keeps only quotes it can find verbatim, and answers with citations. The parts that distinguish it from a chat box with web search are the ledger, the quote check, and the local data boundary.
Who is it for?
mole earns its place when you need a claim you can defend, not a summary you can skim: the quote check and the cross-command crossings log make both the evidence and the leakage inspectable after the fact. Skip it for casual lookups, because a budget is mandatory and every run costs real calls.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 51 days ago.
What is it written in?
Mainly Go, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 2, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The budget is enforced in the schema, not estimated in the prompt

Every model call is reserved before it happens and settled afterwards, against a ledger whose non-negative constraints live in the database schema itself. So `--usd 0.50` is a ceiling the run stops at, not a hint passed to a model. The README reports measured overshoot across the test corpus as 0%.

A budget is required, and the two units cannot be mixed on one command line. Only dollar mode can price a search call, because pricing a search needs a per-call rate. Only token mode can bound a model whose rates mole does not know, which is the case for most self-hosted endpoints.

There is a wrinkle worth internalizing before you choose: a model served from `localhost` is priced at zero and still counted in tokens. For a self-hosted run that costs no money, `--tokens` is therefore the unit that actually bounds it, because the dollar ceiling would never be reached.

A quote that is not verbatim in its source is discarded at extraction

Each extracted claim carries a quote, and a claim whose quote does not appear verbatim in the page it was mined from is thrown away at extraction, before it can reach an answer. The check is on the page text, not on the model's confidence.

Claims that survive can be re-read against their source afterwards. That is what makes the output auditable rather than merely plausible, and it is why mole also looks for contradictions between the claims it collected instead of merging them into one smooth paragraph.

A claim that turns out not to be supported is marked as such in the report rather than quietly dropped. So the failure mode is visible: you get a list that includes its own weak entries, flagged, instead of an answer that hides them.

The concurrency guarantees live in internal/budget and the race detector runs twice

The Makefile says why the race target exists. The concurrency guarantees in internal/budget are the point of the work, and they are only meaningfully tested under `-race`, so the target runs the whole package set with the race detector and a repeat count of two. The aggregate `check` target is `vet` followed by that same race run, which means the default build gate is the strict one rather than a plain test pass.

The build is `CGO_ENABLED=0` for both binaries, which is what makes the single static binary claim hold with no runtime dependencies. The version string is stamped at link time from `git describe --tags --always --dirty`, and the release pipeline has a `release-check` target that runs goreleaser in snapshot mode to build every platform and write to dist/ without publishing.

There is also a `seed` target that writes a synthetic session so `mole trace` has something to show, which tells you trace output is a debugging surface rather than a reporting one.

Local files stay local because only aggregates with five records or more come back

You register a single file or a whole directory, then run a research question with `--actors local_compute`. The privacy boundary is structural rather than a promise in a prompt: the model never sees a row and never writes SQL. It picks a hypothesis template and column names, mole renders the statement, and only aggregates are allowed back. Those are counts, means, test results, and buckets covering at least five records.

That five-record floor is the interesting number, because it is what stops a distinct group from being identified by its own count. It also means a small table may return nothing useful rather than a leaky summary.

Formats are CSV, TSV, JSON, and JSONL. Parquet is explicitly not supported, so a warehouse export needs converting first. And `mole crossings <session-id>` prints exactly what left the machine, which is the command to run before you trust the boundary rather than after.

Toolkit mode splits the reasoning from the parts that are not model calls

mole speaks MCP, and the documented integration has two shapes. You can hand mole a question and collect the answer, which is the batch path behind `mole research`. Or you can run toolkit mode, where your coding agent does the reasoning with its own model and mole supplies the pieces that are not model calls.

That split matters for cost accounting. In toolkit mode the agent's model is not the model mole reserves against, so the ledger in the database covers a smaller share of what you actually spend. Nothing in the README reconciles the two.

`mole serve` is the server side of this. It listens on a unix socket, mode 0600, in a private directory, rather than a TCP port, which keeps the whole MCP surface inside the machine's filesystem permissions.

Two unrelated projects already own the name mole on Homebrew and the AUR

The install section is longer than it needs to be because of name collisions. On Homebrew the tap is fully qualified as `lajosdeme/mole/mole`, and it has to be: an unrelated `mole`, a macOS cleanup tool, lives in homebrew/core, so `brew install mole` will always resolve to that one instead.

On Arch the packages are `mole-research-bin` for the prebuilt release binaries and `mole-research` to build from source. The plain `mole` name and `mole-bin` on the AUR belong to an SSH tunnelling tool that has held them since 2020. The research package installs `/usr/bin/mole` and declares the conflict, so pacman tells you instead of overwriting.

Both Homebrew and the AUR install a binary called `mole`, so only one of them can be linked at a time. If you are moving between them, the binary is what collides, not the package manager.

Keys go in a 0600 config file and never in the MCP client config

You need two providers, one for search and one for the model, and both keys live in `~/.config/mole/config.json` with mode 0600. The explicit prohibition is worth repeating: never in environment variables that leak into process listings, and never in `.mcp.json`.

bash
mole config set search.provider tavily          # or: brave
mole config set search.tavily-key tvly-...

mole config set llm.provider anthropic          # or: openai-compatible
mole config set llm.api-key sk-...
mole config set llm.model claude-sonnet-5
mole config set llm.cheap-model claude-haiku-4-5

mole doctor                                     # verify everything above

Two model slots, a main one and a cheap one, which is how a single run keeps its small classification calls off the expensive model. Any OpenAI-compatible endpoint works as well: DeepSeek, Ollama, llama.cpp, vLLM, or a proxy, set through `llm.provider openai-compatible` plus a `llm.base-url`. `mole doctor` is the check that both providers resolve before you spend anything.

The database is SQLite, created on first use under your XDG data directory, which means the first command that touches state also creates the file. Nothing in the README describes a way to point that database elsewhere, so an XDG_DATA_HOME override is the only lever visible from the documentation.

Dataset mode merges rows by fuzzy key and admits where sources disagree

The same research loop can emit a dataset instead of prose. You declare a schema, and an exclamation mark marks the field that identifies a row.

bash
mole research "largest UK supermarket chains and their revenue" \
  --mode dataset \
  --schema 'company:text!,revenue:number=annual revenue in GBP,employees:number' \
  --usd 0.50

Rows are merged across sources by fuzzy key rather than exact match, so `Aldi` and `Aldi UK` collapse into one row carrying two sources. That merge is what makes the disagreement visible later.

Export has two shapes. CSV holds one value per cell and says so: it carries a source count and a `contested` column naming the fields the sources disagree about. JSON carries every disagreeing value along with the sources behind each one. Choosing CSV means losing the minority values; choosing JSON means carrying the mess.

Export is separate from research, so the same session can be re-exported in both shapes without spending again. The follow-up command is separate too: `mole ask <session-id>` answers from the claims that session already collected, with no new searching and no new spending beyond the single call that phrases the answer. That is the cheapest way to interrogate a run you are not happy with.

Editorial conclusion

mole earns its place when you need a claim you can defend, not a summary you can skim: the quote check and the cross-command crossings log make both the evidence and the leakage inspectable after the fact. Skip it for casual lookups, because a budget is mandatory and every run costs real calls. Before your first run, run mole doctor and confirm both a search provider and a model provider are set, decide dollar budget or token budget deliberately since the two are mutually exclusive, and if you point it at local files, read the output of mole crossings once before you trust the privacy claim. Build from source with make install rather than go install if you want mole version to report the release tag instead of dev.

Frequently asked questions

What is mole and how do I install it?

mole is a deep-research agent that ships as a single static binary built with CGO_ENABLED=0 and no runtime dependencies. The install script is piped into a shell and verifies a SHA-256 checksum, Homebrew uses the fully qualified `lajosdeme/mole/mole`, Arch uses the AUR packages `mole-research-bin` or `mole-research`, and Debian and Ubuntu have a `.deb` on the releases page.

Does mole need a budget flag to run?

Yes, and `--usd` and `--tokens` are mutually exclusive. Only dollar mode can price a search call, and only token mode can bound a model whose rates mole does not know. A model served from localhost is priced at zero and still counted in tokens.

How does mole check that a claim is real?

Each claim carries a quote, and a claim whose quote does not appear verbatim in the page it was mined from is discarded at extraction. Surviving claims can be re-read against their source, and one that turns out not to be supported is marked in the report instead of being quietly dropped.

Can mole analyse my local CSV without uploading it?

Register a file or folder with `mole connect add` and pass `--actors local_compute`. The model never sees a row and never writes SQL; it picks a hypothesis template and column names, mole renders the statement, and only aggregates come back: counts, means, test results, and buckets covering at least five records. `mole crossings` prints what left the machine.

Where should mole keep my API keys?

In `~/.config/mole/config.json` at mode 0600. The documentation explicitly rules out environment variables, which leak into process listings, and `.mcp.json`. Set them with `mole config set` and verify them with `mole doctor`.

Official sources

  1. Issues
  2. lajosdeme/mole on GitHub
  3. License: Apache-2.0
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/lajosdeme-mole.svg)](https://hysenlabs.com/projects/lajosdeme-mole)