Hound (master-fetch): a keyless MCP web-fetch server for AI agents
MCP server for web fetching with Cloudflare bypass, Trafilatura extraction, and smart routing. Free, self-hosted, no API keys.
At a glance
- What is it?
- Hound is an MIT-licensed MCP server that gives an agent fetch, crawl, PDF and keyless search over one local process. It installs from PyPI in two commands, and its own README is candid about where the stealth tier stops working.
- Who is it for?
- Adopt Hound if you run a local MCP client and want fetch, crawl and keyless search without a scraper account: the two-command install and the structured response fields are the real draw. Skip it if you need a supported SLA, a stable tool contract across majors, or a single static binary, because the project is self-described as Beta and the README lists its own failure modes.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 68 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 27, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The gap Hound fills for agents that need the web
An agent that can call tools but cannot read a page is limited to whatever you paste into the prompt. The usual fix is a hosted scraping or search API, which means an account, a key, and per-request billing. Hound takes the other route: one MCP server process on your machine, no accounts, no keys, MIT licensed. It is aimed at people running Claude Code, Cursor, OpenCode, Hermes or Pi locally and who want the agent to fetch, crawl, read PDFs and search without wiring a vendor into the loop. The README frames it as "Hound is for the agent itself. You install it once; the agent calls it whenever it needs the web." That framing matters for scope: this is not a scraping framework you script against, it is a tool surface an agent decides when to use. The project also ships a pi-extension directory and a package.json entry pointing at ./pi-extension/extensions/hound.ts, so Pi users get a native extension rather than a generic MCP hookup.
Six tools, a warm browser, and structured fields the agent can branch on
The mechanism is a single local process that keeps one browser warm and exposes six tools over MCP. Fetching goes through primp for HTTP with TLS impersonation; when a page needs a real browser, patchright drives Chromium, and trafilatura plus markdownify turn the result into markdown. Search is keyless by default and can be upgraded with your own keys for Serper, Tavily, Exa, Firecrawl or TinyFish via hound keys add, which the README says become the primary source with key stacking and automatic fallback to the local engines when keys run out. The more interesting design choice is the response shape. Instead of returning text and letting the agent guess, Hound returns fields such as content_ok, next_action, summary, page_type, content_age_days, is_stale, source_type, is_official, relevance_score and fetch_relevance, so the agent branches on structured values rather than parsing error strings. Hard blocks (404, bot wall, auth) are supposed to return clean errors instead of fake content. Tool definitions are hand-written rather than generated from Pydantic schemas, which the README puts at roughly 2.9K tokens across the six tools. That is a real constraint on context budget, and it is the kind of number worth re-measuring yourself rather than taking on faith.
Installing Hound and pointing an MCP client at it
The README gives a two-command install. The [all] extra pulls fetch, crawl, keyless search, PDF, OCR and neural rerank; plain hound-mcp is the lean install with no browser dependencies. The second command installs the Chromium build that patchright drives.
pip install hound-mcp[all]
playwright install chromiumAfter that, point any MCP client at the hound command. The README states there are no arguments, no keys and no environment variables required for the default setup. Before trusting it in a client, verify the install itself:
hound -v # version + update status
hound --doctor # health check + fix adviceIf the launcher breaks after an update, the README gives a recovery path rather than a reinstall:
python ~/.hound/repair.pyDocker is the other supported route, and the compose file documents the port and transport. The container serves streamable HTTP at http://0.0.0.0:8765/mcp, which is what an Open WebUI 0.6.31 or later client would point at.
services:
hound:
build:
context: .
image: hound-mcp:latest
shm_size: "1gb"
init: true
ports:
- "0.0.0.0:8765:8765"The compose file is explicit that the HTTP endpoint has no authentication and that the port binding should be changed to "8765:8765" only on a trusted network. It also explains why shm_size is 1gb: Chromium renderers use /dev/shm heavily and Docker's 64MB default causes silent tab crashes on media-heavy pages. It adds dns entries (1.1.1.1 and 9.9.9.9) because network-level ad-blocking was filtering browser subresources and producing error pages. Those three lines are the difference between a container that works and one that fails intermittently, and they are worth reading before you trim the file.
Where Hound breaks, by its own account
The README has a Known Gotchas section, which is more than most projects of this kind offer. Two entries stand out. First, stealthy fetches can return anti-adblock overlays or "CSS onerror" error pages when something on the network is filtering the browser's subresources, which is why the Docker setup ships its own DNS servers. That is not a bug you fix in Hound; it is an environment problem that presents as a fetch failure. Second, the compose comments note that Hound's stealthy tier passes --disable-dev-shm-usage but the dynamic tier does not, so the shared-memory requirement is tier-dependent. There is also a platform boundary: the pyproject description says Hound degrades gracefully to HTTP-only mode when browser dependencies are unavailable, naming Termux and Android. In that mode the anti-bot work is gone, so any target that needs a real browser will fail. The pyproject classifier is Development Status :: 4 - Beta, and the same file documents a self-healing update path with hound --rollback and a repair script. Those exist because updates have broken installs before. Treat automatic self-update as a convenience with a failure history, not as a guarantee.
Hound versus a hosted scraping API, and when to pick which
The closest alternative is a paid scraping or search API such as Firecrawl, Tavily or Exa. The difference is not quality, it is where the work runs and who pays for it. A hosted API terminates the fetch on someone else's infrastructure, handles proxy rotation and CAPTCHA as a service, and bills per request. Hound runs everything locally, which means the cost is your CPU, your IP address and your maintenance time. The README even acknowledges the hosted camp by letting you bring those same providers' keys in through hound keys add, with key stacking and fallback to the local engines. That is a sensible hedge: if the keyless engines are not returning what you need, you are not locked out. But the trade is real. A hosted API gives you a support channel, a versioned contract and someone else to blame when a target changes. Hound gives you a local process, a browser that needs to be installed and kept current, and an update mechanism that has needed a repair script. If your agent runs on a laptop and your targets are ordinary pages, local wins on cost and on data never leaving the machine. If you are fetching at volume from a server and cannot tolerate an occasional unexplained block, the hosted route is the one with someone on the other end.
Maintenance, licence and upgrade cost
The licence is MIT, which permits commercial and closed-source use and requires only that the copyright notice and permission notice travel with the code. There is one licence caveat worth checking yourself: the repository contains a NOTICE.ddgs.txt file, and the pyproject comments describe a vendored ddgs engine layer as MIT. Vendored third-party code carries its own notice obligations, so read NOTICE.ddgs.txt before redistributing rather than assuming the top-level LICENSE covers everything. On maintenance, the last push to the default branch was on 2026-07-24, and the most recent release listed is v12.4.1, tagged "Crawl proxy rotation", from the same day. The version has moved quickly through the 12.x line, with v12.3.1, v12.4.0 and v12.4.1 all dated 2026-07-24. Rapid minor releases are a cost as well as a signal: the tool contract the agent depends on can shift, the update path has needed repair, and there is no long-term support branch mentioned anywhere in the project's own files. The CLI does provide hound --rollback to undo the last update, which is the practical mitigation. Pin your version, read the CHANGELOG before upgrading, and treat the self-update command as something you run deliberately rather than automatically.
Editorial conclusion
Adopt Hound if you run a local MCP client and want fetch, crawl and keyless search without a scraper account: the two-command install and the structured response fields are the real draw. Skip it if you need a supported SLA, a stable tool contract across majors, or a single static binary, because the project is self-described as Beta and the README lists its own failure modes. Before trusting it on your targets, run hound --doctor and check the Known Gotchas section against the sites you actually fetch, then pin the version and keep hound --rollback in mind.
Frequently asked questions
Does Hound need an API key to search the web?
No. The README describes search as keyless and local by default, with no accounts and no per-request billing. You can optionally add your own keys for Serper, Tavily, Exa, Firecrawl or TinyFish with hound keys add, and Hound falls back to the keyless local engines when those keys are exhausted.
Can Hound run in Docker, and which port does it use?
Yes. The docker-compose.yml builds a hound-mcp:latest image and serves the MCP streamable-HTTP transport at http://0.0.0.0:8765/mcp. The compose file warns that the endpoint has no authentication and that the port binding should be changed to "8765:8765" only on a trusted network.
What happens if Hound's browser dependencies are missing?
The pyproject description states that Hound degrades gracefully to HTTP-only mode when browser dependencies are unavailable, naming Termux and Android. In that mode the browser-based anti-bot work is not available, so targets that require a real browser will not be fetched.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/dondai1234-master-fetch)