# ketch: one binary that hides eleven search backends

> ketch is a MIT-licensed Go CLI for web search, code search, version-aware library documentation, scraping and crawling, with no daemon and no API server. One configuration call chooses the backend, every command has JSON output and documented exit codes, and the module file retracts a release that shipped broken.

**1broseidon/ketch** — Fast, stateless CLI for web search and scrape. Built for AI agents.

- Repository: https://github.com/1broseidon/ketch
- Website: https://chain.sh/ketch/
- Stars: 682 · Forks: 38
- Language: Go
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/1broseidon-ketch

## One binary, three research surfaces, no daemon

The premise is that research tooling for agents usually means wiring up several provider SDKs, each with its own authentication and its own response shape, and ketch collapses that into a single static binary. Three surfaces cover most of what an agent needs. `ketch search` does web search with no API key required, across eleven named providers. `ketch code` greps real open-source source across public repositories through three different code search backends. `ketch docs` returns curated, version-aware library documentation from one source. Two more commands turn pages and text-based PDFs into clean markdown: `ketch scrape` for a single page and `ketch crawl` for a site. There is no daemon and no API server to run, which is the property that makes it usable from a container or a one-shot agent invocation. The project states two audiences explicitly: humans who want the same job a browser tab or a pipe through a converter would do, and agents that want predictable structured output, JSON on every command, and documented exit codes they can branch on.

## auto picks a backend in a fixed order and reports which one answered

The backend decision is the design centre of the tool, and it is a preference order rather than a random choice. The default backend tries providers in a fixed order and returns the first that answers, then reports which provider served in a `backend:` field in the output. What it prefers first is whatever you have actually configured: your own search instance before anything hosted, then any provider for which you have set a key, and only then the keyless hosted providers. That produces a slightly counter-intuitive rule worth internalising: setting a key is how you get a specific provider and higher limits, and you do not also have to set the backend, because the key alone changes the order. An operator therefore configures the backend once and every later invocation is just a search command with no idea which provider is behind it, which is the whole point for an agent caller. You can still override per call with an explicit backend flag, and you can register a key and pick a provider in the same breath:```sh
ketch config set brave_api_key <key>   # auto now prefers Brave
ketch search "golang error handling" -b ddg   # or pick a provider explicitly
```

## --multi fuses rankings and tags each result with the engines that found it

Two flags change the shape of a search rather than the provider. `--multi` queries several backends at once and fuses their rankings with Reciprocal Rank Fusion, which is the mechanism that makes agreement between engines matter: a page several engines rank highly floats to the top. Results are deduplicated by URL and each one is tagged with the engines that returned it, so the output tells you not just what was found but how many independent indexes agreed on it. `--random` shuffles the backend list, tries one, and falls back to the rest, and the stated use is not wanting to burn the rate limits of every provider to get one provider's results. Both accept either a bare form meaning every usable backend, or an explicit list naming the ones you want. Both are mutually exclusive with the single-backend flag and with each other, so a script has to pick a mode rather than combine them. The three ordinary search forms sit alongside them: a plain query, a variant that fetches and extracts full content for every result, and the multi and random modes above.

## Extraction trusts the page's declared structure before falling back

The extraction pipeline reads what the page says about itself before it guesses. It looks for the content landmark the page declares, such as a main or an article element, then for a document assembled from uniform sections, then for the smallest element holding the prose. Site furniture is removed by what it is rather than by a list of selectors, which covers navigation, hidden and collapsed controls, link rails and tables of contents. A readability implementation is the fallback for a page that declares no structure at all, and it is a real dependency rather than a hope. One configuration key decides how aggressive the pruning may be. The default mode also drops blocks by name and phrase, which catches related-post rails, comment threads, share bars and was-this-helpful boxes, and a complete mode keeps everything the structure does not explicitly condemn, trading precision for the last fraction of recall. Pages served in a legacy encoding are decoded before any of this happens, which is the kind of detail that only appears after someone has been bitten by mojibake. The same pipeline works on stdin, so you can pipe a page you already have:```sh
curl -L https://chain.sh/ketch | ketch extract
cat page.html | ketch extract --select article --max-chars 4000
```

## PDF output is never binary, and an external converter becomes authoritative

PDF handling has the clearest failure policy in the project. Detection is done from the response MIME type or the file signature, and extraction uses a built-in pure-Go parser, so a text PDF needs nothing installed. A scanned or image-only PDF needs optical character recognition, and the built-in path returns a precondition error carrying a hint about the OCR converter rather than emitting an empty document. Operators can point the tool at an external converter that writes markdown to standard output, capped at ten megabytes, and the command is parsed with a shell-lexer library and must contain exactly one input placeholder:```sh
ketch config set external_pdf_to_md_converter_command 'pdftotext "{input}" -'
ketch config set external_pdf_to_md_converter_timeout_sec 300
```The policy that matters is what happens when the external converter fails. When one is configured it is authoritative, so a failure is returned to the caller instead of silently falling back to the built-in parser, which means a broken converter produces a visible error rather than quietly different output. Binary output is refused outright: the raw and selector flags reject PDFs as validation errors. And even with the forced-browser flag, a PDF still goes through text extraction and never opens a browser, which keeps the heavyweight path away from the most common case.

## The install script verifies a checksum, and the pinning example is one release behind

Installation offers five routes, and the script route is the one with the most detail about what it does:```sh
# macOS / Linux, any of x86_64 or arm64
curl -fsSL https://ketch.run/install | sh

# Homebrew
brew install ketch

# npm
npm install -g ketch-cli

# go install
go install github.com/1broseidon/ketch@latest

# Or download a prebuilt binary (linux/darwin/windows, amd64/arm64)
# from https://github.com/1broseidon/ketch/releases
```The script picks the build matching your operating system and architecture, verifies it against a checksum file published with the release, and puts the binary in a system directory when that is writable and in a per-user directory otherwise. The readme links to the script itself and tells you to read it first if you would rather not pipe a download into a shell, which is the correct instruction given what the script does. Versions and destinations can be passed through to the same script with a shell-style argument, for example `sh -s -- --version v0.17.0 --bin-dir ~/bin`. One detail to note is that the example pins a release two minors behind the newest tag in the project's release list, so treat the example as illustrative rather than as a recommendation. A fresh readme is not always refreshed in the same commit as a release, and this is the place it shows.

## go.mod retracts a release and explains why in a comment

The module file is unusually candid about its own history. It declares a Go version and then retracts a released version, with the reason written directly above the directive: that version was published and withdrawn on the same day because its documentation backends were not ready, and the following patch release was the previous version plus fixes. Retracting is a formal mechanism that tells tooling not to select that version, and documenting the reason in the manifest means a user hitting it does not have to guess. The dependency list is thirteen entries, each pinned exactly, and it reads like a map of what the tool actually does. A readability implementation for the extraction fallback, a markdown converter for the output format, a query and CSS library for element selection, a browser driver for the forced-browser mode, a shell lexer for parsing the external converter command, a pure-Go PDF parser, the Model Context Protocol SDK, a command framework, an embedded key-value store for the cache directory, and a standard library extension set. No provider SDK appears, which is the claim the whole project rests on.

## The command surface is also the package layout

The repository root is laid out one directory per capability, and the names match the commands a user types: code, config, cookies, crawl, doctor, extract, health, scrape, search, plus cache for the on-disk state, updatecheck for version notices, and a URL rewriting package that decides what a rewritten address means to a fetcher. Beyond those sit directories for the MCP server, a plugin manifest, bundled skills, an example site, the npm packaging, a design directory and the documentation. That layout is worth noting because it makes the claim about being agent-friendly checkable: if the tool were secretly a thin wrapper over provider libraries, the root would not have its own extraction, PDF and cache packages. Two top-level files describe the manual and the release process, and there is a hooks directory, a lint configuration, a release configuration and agent-facing context files, so the repository is set up to be worked on by coding agents as well as by people.

## Conclusion

ketch is worth installing if you are writing an agent that needs research and you would rather not maintain four provider SDKs, because the backend decision moves to an operator and the agent-facing surface becomes three commands plus JSON. Two cautions to keep in mind. Output quality is the union of its fallbacks, so the keyless hosted providers are what you get by default and their limits and rate behaviour are your rate behaviour. And the PDF path has two tiers, a built-in parser that refuses scanned documents with a precondition error, and an external converter that becomes authoritative when configured, so read which one is active before you trust a PDF result in a pipeline.

## FAQ

### Does ketch require an API key for web search?

No. The default backend tries providers in a fixed order and returns the first that answers, reporting which one served in the output. Setting a key changes the preference order and raises limits, without needing to set the backend separately.

### What does the --multi flag do in ketch search?

It queries several backends at once and fuses their rankings with Reciprocal Rank Fusion, so a page several engines rank highly floats upward. Results are deduplicated by URL and each is tagged with the engines that returned it.

### How does ketch extract a PDF, and what happens with scanned documents?

A PDF is detected from the response type or file signature and read with a built-in pure-Go parser. A scanned or image-only document needs optical character recognition, and the built-in path returns a precondition error with a hint instead of an empty document.

### What happens if the external PDF converter configured in ketch fails?

The failure is returned. When an external converter is configured it is authoritative, so ketch does not silently fall back to the built-in parser. Its output is capped at ten megabytes and the command must contain exactly one input placeholder.

### Which research surfaces does the ketch CLI expose?

Web search across eleven providers, code search across public repositories through three backends, version-aware library documentation, plus scrape and crawl for turning pages and text-based PDFs into markdown. Every command takes JSON output.

### Was a release of ketch withdrawn?

Yes. The module file retracts one version, with a comment explaining that it was published and withdrawn on the same day because its documentation backends were not ready, and that the following patch was the previous release plus fixes.

## Sources

- [1broseidon/ketch on GitHub](https://github.com/1broseidon/ketch)
- [License: MIT](https://github.com/1broseidon/ketch/blob/main/LICENSE)
- [Project website](https://chain.sh/ketch/)
- [README](https://github.com/1broseidon/ketch/blob/main/README.md)
- [Releases](https://github.com/1broseidon/ketch/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/1broseidon-ketch
