CLI tool
1broseidon/ketch avatar
1broseidon/ketch

ketch: a stateless Go CLI that gives agents one binary for search, code, docs and scraping

Fast, stateless CLI for web search and scrape. Built for AI agents.

675 stars39 forksGoMIT

At a glance

What is it?
ketch wraps nine web search backends, three code search backends and a Context7 docs surface behind a single Go binary with --json on every command. It is a good fit for agent pipelines that want one config call instead of per-provider glue code, and a poor fit for anyone who needs JavaScript-heavy pages without installing a browser.
Who is it for?
Adopt ketch if your agents call search or scrape from a shell and you want one binary with --json, documented exit codes and a single config surface, and if you are willing to run ketch doctor against every backend you enable. Do not adopt it if you need JavaScript-rendered pages out of the box, since the browser is a separate install step, or if you need a hosted service with per-tenant quotas.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 2 days ago.
What is it written in?
Mainly Go, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The problem ketch solves: provider glue code in agent pipelines

An agent that needs to research something usually ends up with three or four SDKs: one for web search, one for code search, one for library documentation, one for fetching and cleaning pages. Each has its own authentication, its own response shape and its own way of failing. The README frames ketch as the collapse of that surface into one binary with three research commands, search, code and docs, plus scrape and crawl for turning HTML and text-based PDFs into markdown. The stated audience is split deliberately: humans who would otherwise use a browser tab or curl piped into pandoc, and agents that want predictable output. The agent-facing part is the design constraint that matters. Every command accepts --json, exit codes are documented for control flow, and ketch config returns effective configuration as JSON so a caller can discover which backends are active instead of probing the environment. An operator sets the backend once with ketch config set backend searxng, and later invocations just call ketch search without knowing which provider answers. That is the whole pitch, and it is a narrow one: ketch is not trying to be a search engine or a crawler framework, it is trying to be the stable process boundary between an agent and whichever provider is configured.

How the backend abstraction and rank fusion actually work

The architecture is a command layer over swappable backends. The search surface lists Brave, DuckDuckGo, SearXNG, Exa, Firecrawl, Keenable, Tavily, Parallel and SerpBase, with brave as the default. The code surface offers Grep, Sourcegraph and GitHub Code Search, with Grep as the default and no configuration needed. The docs surface is Context7 alone, described as curated and version-aware snippets. Scraping is handled in-process: the go.mod file pulls in codeberg.org/readeck/go-readability/v2 and github.com/JohannesKaufmann/html-to-markdown/v2, so the readability-then-markdown pipeline is a library call, not a subprocess. Two flags change the search fan-out. --multi queries several backends at once and fuses their rankings with Reciprocal Rank Fusion, so a page that several engines rank highly floats upward, results are deduplicated by URL and each result is tagged with the engines that returned it. --random shuffles the backend list, tries one, and falls back to the rest, which the README describes as the way to get one provider's results without spending rate limits on all of them. Both accept a bare form covering all usable backends or an explicit list such as =brave,exa, and both are mutually exclusive with --backend and with each other. The caching layer is bbolt, and there is a separate cache command to show stats or clear it. The repository also carries a mcp directory and the go.mod requires github.com/modelcontextprotocol/go-sdk, so ketch mcp serve exposes the five research surfaces as MCP tools over stdio.

Installing ketch and running a first search

Three install paths are documented: a Homebrew tap, go install, and prebuilt binaries for linux, darwin and windows on amd64 and arm64. The Go module requires go 1.25.7, so building from source needs a toolchain at least that new. The repository Makefile builds the binary with go build -o ketch . and runs the suite with go test ./...; golangci-lint run is the lint target.

bash
brew install 1broseidon/tap/ketch

Alternatively, install straight from the module path:

bash
go install github.com/1broseidon/ketch@latest

Scraping needs no configuration at all, which makes it the honest first test. Pointing ketch scrape at a page prints a small YAML-style front matter block with url, title and a word count, followed by the extracted markdown and a numbered list of links.

bash
ketch scrape https://go.dev/doc/effective_go

Code search is likewise zero-configuration, because the default backend is Grep. The example in the README searches a Go symbol and limits output to two results; the output header echoes the query, the language, the backend name and the result count, then lists file paths with line numbers and a GitHub URL.

bash
ketch code "http.NewRequestWithContext" --lang go --limit 2

Web search is the one surface that needs setup. The default backend is brave, and the README states that Brave, Tavily and SerpBase require a free API key. Once the key is set, the same query can be run three ways, plain, with --scrape to fetch and extract full content per result, or with --multi to federate across every usable backend.

bash
ketch config set brave_api_key <key>
ketch search "golang error handling"
ketch search "golang error handling" --multi

If you already have HTML and do not want a fetch, ketch extract runs the readability and markdown pipeline on stdin. The README shows both a curl pipe and a file with a CSS selector and a character cap.

bash
curl -L https://chain.sh/ketch | ketch extract
cat page.html | ketch extract --select article --max-chars 4000

PDF handling and the external converter trade-off

ketch scrape detects a PDF from the response MIME type or from the %PDF- signature and extracts text with a built-in pure-Go parser, which the go.mod confirms as github.com/ledongthuc/pdf. The limitation is stated plainly: scanned or image-only PDFs need OCR and return a precondition error with an OCR-converter hint. Operators who need better output can configure an external PDF-to-Markdown converter that writes markdown to stdout, capped at 10 MiB, with a shlex-parsed command that must contain exactly one {input} placeholder. The README gives pdftotext as the example and a 300 second default timeout. The design choice worth noticing is that once the external converter is configured it becomes authoritative: failures are returned rather than silently falling back to the built-in parser. That is the right call for reproducibility, because a silent fallback would make output quality depend on an invisible failure, but it also means a broken converter command turns every PDF scrape into an error. Two more constraints are documented: --raw and --select reject PDFs as validation errors, so PDF binary output is never emitted, and even with --force-browser a PDF still goes through text extraction rather than opening Chromium's PDF viewer.

Where ketch is the wrong tool

The stateless design is the source of its main weakness. JavaScript-rendered pages are not handled by the base binary; there is a browser command group with install and status subcommands that manages headless Chrome, and the go.mod depends on github.com/go-rod/rod, but the browser is something you install rather than something that works on first run. If your scrape targets are single-page applications, that install step is a real part of your deployment, not a footnote. The second boundary is authentication. Nine search backends are listed, and the README states that Brave, Tavily and SerpBase need a free key, while ddg, searxng, exa, firecrawl, keenable and the remaining providers are described in the backend table whose setup column is truncated in the available material. If you need a provider that is not on the list, ketch has nothing to say to you. Third, ketch is a client, not a service: there is no daemon and no API server to run, which is the point, but it also means there is no shared rate-limit budget, no per-tenant accounting and no central place to enforce a quota across a fleet of agents. Each process talks to the provider directly. Finally, the version history shows a withdrawn release: go.mod carries a retract directive for v0.16.0, noting it was published and withdrawn the same day because its docs backends were not ready, with v0.16.1 described as v0.15.0 plus fixes. Anyone pinning versions should read that comment before assuming the newest tag is the safest.

Alternatives and how their approach differs

The closest alternative in spirit is a small script that calls each provider's SDK directly and normalizes the output itself. That gives you full control over retries, caching and response shape, and it costs you exactly what ketch claims to remove: per-provider auth code, per-provider response parsing, and a different failure mode for each one. If your pipeline uses only one provider and will never use a second, the SDK plus a thin wrapper is less machinery. The other direction is a hosted search API that bundles fetching and extraction server-side. Those remove the install and the key management from your side, but they put a network hop and a vendor between your agent and the page, and they cannot run in an air-gapped environment. ketch sits between the two: local process, local cache in bbolt, local PDF parser, but still dependent on an external provider for the actual search index. For the code search surface specifically, ketch is a client over Grep, Sourcegraph or GitHub Code Search rather than an index of its own, so it inherits those services' coverage and rate limits. The docs surface is Context7 only, with no second option listed, which makes it the least substitutable part of the tool.

Maintenance, upgrade cost and licence

The repository is not archived, and the last push was on 2026-09-10, the same day v0.16.2 was released, with v0.16.1 two days earlier and v0.15.0 on 2026-09-07. That is a rapid release cadence, which cuts both ways: fixes arrive quickly, and so do surface changes. The retract directive for v0.16.0 is the concrete upgrade hazard, because a withdrawn version can still be present in a module cache or a lockfile, and Go's retract mechanism only warns. Pinning to a specific tag and reading CHANGELOG.md before moving is the cheap mitigation. ketch is MIT licensed, which permits commercial and closed-source use, modification and redistribution provided the copyright notice and licence text are retained; that is a summary of the licence identifier, not legal advice, and anyone embedding ketch in a distributed product should read the LICENSE file in the repository root. The dependency list is moderate and mostly Go-native, with go-rod, bbolt, cobra, the MCP Go SDK and two HTML processing libraries as the notable ones, so the supply-chain surface is smaller than a typical Node-based equivalent. The site directory and the published reference at 1broseidon.github.io/ketch are where the full flag list lives, and the README defers to them rather than duplicating every flag.

Editorial conclusion

Adopt ketch if your agents call search or scrape from a shell and you want one binary with --json, documented exit codes and a single config surface, and if you are willing to run ketch doctor against every backend you enable. Do not adopt it if you need JavaScript-rendered pages out of the box, since the browser is a separate install step, or if you need a hosted service with per-tenant quotas. Before committing, verify that your chosen search backend works without a key, that the external_pdf_to_md_converter_command you set actually writes Markdown to stdout, and that your Go toolchain satisfies go 1.25.7 in go.mod if you build from source.

Frequently asked questions

Does ketch need an API key to work?

Not for everything. ketch scrape, ketch extract and ketch code work without configuration, and the default code backend is Grep. Web search defaults to brave, and the README states that Brave, Tavily and SerpBase require a free API key set through ketch config set.

Can ketch handle JavaScript-heavy pages?

Only with the browser installed. There is a browser command group with install and status subcommands that manages headless Chrome, and go.mod depends on go-rod, so rendering is available but it is a separate setup step rather than default behaviour.

What does ketch's --multi flag do?

It queries several search backends at once and fuses their rankings with Reciprocal Rank Fusion, deduplicating by URL and tagging each result with the engines that returned it. It is mutually exclusive with --backend and with --random.

How does ketch extract text from PDFs?

ketch scrape detects PDFs from the response MIME type or the %PDF- signature and uses a built-in pure-Go parser. Scanned or image-only PDFs need OCR and return a precondition error, and an external converter can be configured through external_pdf_to_md_converter_command if the built-in output is not enough.

Can ketch run as an MCP server?

Yes. The README lists an mcp command that runs ketch as an MCP server over stdio, exposing the five research surfaces as tools, and go.mod requires the Model Context Protocol Go SDK.

Official sources

  1. 1broseidon/ketch on GitHub
  2. License: MIT
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/1broseidon-ketch.svg)](https://hysenlabs.com/projects/1broseidon-ketch)