# browser-search: A Three-Tier Self-Hosted Web Search and Browse Skill for AI Agents

> Johell1NS/browser-search is an MIT-licensed skill for AI agents that combines SearXNG for multi-source search, Camofox for standard browsing, and CloakBrowser for anti-bot bypass, running entirely on self-hosted Docker and npm infrastructure with no API keys or rate limits.

**Johell1NS/browser-search** — A skill for AI agents: search the web with SearXNG, browse with Camofox, bypass protections with CloakBrowser. Anti-hallucination by design. Self-hosted, free, unlimited.

- Repository: https://github.com/Johell1NS/browser-search
- Website: https://github.com/Johell1NS/browser-search#readme
- Stars: 526 · Forks: 38
- Language: JavaScript
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/johell1ns-browser-search

## Why AI Agents Need a Dedicated Web Access Stack

An AI agent browsing the public web without specialized tooling runs into two problems immediately. The first is accuracy: without live web access, the agent can only draw from its training data, producing answers that may be months or years out of date. Asking it to verify a recent benchmark result, a product's current pricing, or a news event that happened last week produces confident but potentially wrong answers.

The second problem is access. A large fraction of high-value web pages sit behind anti-bot protections: Cloudflare challenge pages, DataDome browser fingerprint checks, Akamai behavioral analysis. Standard browser automation tools fail at these barriers and return a blocked error, forcing the agent to give up on that source.

browser-search addresses both problems with a three-component stack: SearXNG for the initial search phase, Camofox for browsing standard sites, and CloakBrowser for sites that block automated access.

## SearXNG, Camofox, and CloakBrowser: What Each Tool Does

SearXNG is a self-hosted metasearch engine that queries multiple upstream sources (Google, Wikipedia, Bing, DuckDuckGo, and others) simultaneously and returns consolidated JSON results with titles, snippets, and URLs. It runs as a Docker container on localhost:8080. The search phase is fast because SearXNG returns structured results without downloading or rendering the full target pages.

Camofox is a browser navigable via REST API. It handles the majority of standard pages: sites that do not actively block automated browsers. It runs as a Docker container and is the default tool for the browsing phase.

CloakBrowser handles sites where Camofox fails. The README documents that CloakBrowser applies 58 C++ source-level patches to its browser implementation and scores 0.9 on reCAPTCHA v3. It automatically detects and resolves Cloudflare, Turnstile, DataDome, Imperva, PerimeterX, and DDoS-Guard challenges, waiting for them to clear before extracting content.

The three tools are not used in parallel. The typical flow is: SearXNG produces a list of relevant URLs, then the agent browses each URL with Camofox. If Camofox encounters a challenge page, the agent escalates to CloakBrowser for that specific URL. This tiered approach keeps most interactions fast (SearXNG and Camofox are lightweight) while preserving coverage for protected sites.

## Anti-Hallucination by Design: Deterministic Scripts

The README describes the anti-hallucination property as a design outcome, not a side effect. The skill enforces a specific interaction pattern: all three tools are invoked through deterministic scripts rather than natural-language tool descriptions. The model cannot paraphrase the commands or infer alternate invocations. It follows the exact script, or the script does not run.

The example in the README shows a concrete invocation:

```bash
node scripts/searxng/searxng.mjs search "largest llm benchmark 2026"
```

This returns a list of URLs the agent then browses with Camofox or CloakBrowser. Because the agent must retrieve real web content before answering, the skill's Deep Research mode enforces a 'search first, answer second' workflow. The agent is not permitted to answer a factual question based on its training data alone when the skill is active: it must verify the claim against live sources first.

This approach works with low-capability models as well as high-capability ones, because correctness depends on the script execution rather than on the model's ability to recall accurate training data.

## Installation and Skill Configuration

browser-search installs as an npm package. The package.json lists four dependencies: cloakbrowser ^0.5.5, playwright-core 1.62.1, mmdb-lib ^3.0.3, and socks-proxy-agent ^10.1.0. SearXNG and Camofox run as Docker containers whose configuration files live in the docker/ directory of the repository.

The SKILL.md file at the repository root is the agent-facing rulebook. It is plain text and editable. The README notes that you can modify the core rules, add your own, or remove rules that do not apply to your workflow. This makes it straightforward to adjust the Deep Research behavior, add domain-specific instructions, or restrict which sources the agent consults.

The skill works with OpenCode (its primary target), Claude Code, Cursor, and other agents that read SKILL.md-based skills. The README describes the skill as identical across all supported agents, with the only difference being how each agent discovers and activates skills in its environment.

## Running on Minimal Hardware

The README states that browser-search was built and tested on a Raspberry Pi. This is not a performance claim but a statement about the minimum resource floor: if the three-component stack runs on a single-board computer, it runs on any development machine or always-on server. The README positions this as a practical argument against high-infrastructure self-hosted web search stacks.

This constraint does not extend to CloakBrowser's anti-bot bypass work. CloakBrowser's 58 C++ patches and browser fingerprint manipulation are computation that happens within the browser process itself. A Raspberry Pi running concurrent CloakBrowser sessions may experience slower challenge resolution than a desktop machine with more CPU headroom. The Raspberry Pi claim covers the coordination layer (SearXNG, Camofox, and the npm skill logic), not necessarily the full CloakBrowser workload under heavy concurrent use.

Since everything runs locally, there are no per-query charges, rate limits, or API keys to manage. The ongoing operational cost is the hardware and electricity to run three Docker containers and one npm-backed process, plus the time to keep the Docker images updated as SearXNG and Camofox release new versions.

## Compared to Hosted Search APIs

Commercial search APIs for AI agents (Serper, Tavily, Bing Search API, and others) offer a clean integration: one HTTP call, one API key, one monthly bill. The agent sends a query and gets back structured results. The overhead is low but the costs accrue per query, and the results come from a single provider's index.

browser-search's model inverts that: higher setup cost, zero per-query cost, and multi-source results via SearXNG's metasearch. For a team running a high-volume research agent, the per-query cost of a commercial API can exceed the cost of self-hosting within weeks. For a developer running occasional research sessions, the setup overhead may not be justified. The anti-bot capability is the feature that no commercial search API provides: Serper and Tavily return search results pages, not the content behind a Cloudflare wall. CloakBrowser is the only component in this stack that crosses that barrier.

## Conclusion

browser-search fits a developer or team running an AI agent that does regular web research and needs factual answers rather than hallucinated ones. The three-tier escalation handles most real-world sites, including those with Cloudflare and DataDome protection, without requiring API keys or per-query payments. The resource floor is modest: the project is tested on Raspberry Pi hardware. The trade-off is operational complexity: three Docker containers (SearXNG, Camofox) and one npm package (CloakBrowser) must be running and accessible to the agent. Verify that your target sites are within CloakBrowser's supported challenge list before depending on it for critical research tasks.

## FAQ

### How does browser-search prevent AI agent hallucinations during web research?

The skill enforces deterministic script execution: the agent invokes fixed scripts for SearXNG search, Camofox browsing, and CloakBrowser bypass rather than reasoning about what to call. In Deep Research mode, the agent must verify every factual claim against live web sources before answering, eliminating reliance on potentially outdated training data.

### What happens when Camofox encounters a Cloudflare-protected website?

The agent escalates automatically to CloakBrowser for that URL. CloakBrowser detects Cloudflare, Turnstile, DataDome, Akamai, Imperva, PerimeterX, and DDoS-Guard challenges and waits for them to resolve before extracting the page content. No human intervention or configuration change is needed.

### Does browser-search require API keys or subscriptions to operate?

No. SearXNG, Camofox, and CloakBrowser all run on your own hardware via Docker and npm. There are no external API keys, no subscription fees, and no per-query rate limits. The only cost is the infrastructure to run the self-hosted components.

## Sources

- [Johell1NS/browser-search on GitHub](https://github.com/Johell1NS/browser-search)
- [License: MIT](https://github.com/Johell1NS/browser-search/blob/master/LICENSE)
- [Project website](https://github.com/Johell1NS/browser-search#readme)
- [README](https://github.com/Johell1NS/browser-search/blob/master/README.md)
- [Releases](https://github.com/Johell1NS/browser-search/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/johell1ns-browser-search
