# ask-search needs json enabled in SearxNG, and Reddit will still block the deep dive

> A self-hosted search skill that gives an agent a search command backed by your own SearxNG instance, with no API key and no third-party query log. The setup has one step people skip, and the documentation is honest that search works where fetching the page does not.

**ythx-101/ask-search** — Self-hosted web search skill for AI agents (OpenClaw/Claude Code/Antigravity) via SearxNG

- Repository: https://github.com/ythx-101/ask-search
- Stars: 537 · Forks: 49
- Language: Python
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/ythx-101-ask-search

## The problem is four priced APIs and one privacy leak

The case for the tool is made as a list of the alternatives and what each costs you. The Brave Search API is priced at three dollars per thousand queries and rate limited. Google Custom Search is five dollars per thousand with daily caps. The Bing API is paid and described as a complex setup. And the web search built into assistants sends your queries to third-party servers, which is a different problem entirely: not cost, but the fact that the query log lives somewhere you do not control. The stated want is narrow and reasonable, which is that a local agent should be able to search without paying per query or leaking what it is looking for. ask-search answers it by wrapping SearxNG, a self-hosted meta search engine that aggregates Google, Bing, DuckDuckGo, Brave and more than seventy other sources. One command, all the results, no per-query cost, and the search traffic originates from your own instance.

## Enabling json in search.formats is the step people miss

The fast path is three commands when you already have SearxNG running, and the longer path is Docker. What matters in both cases is a single configuration change on the SearxNG side, and it is called out again at the end of the setup notes as a requirement rather than a suggestion. You have to add `json` to the search formats in `settings.yml`:

```yaml
search:
  formats:
    - html
    - json
```

Without it SearxNG answers in HTML and the skill has nothing to parse, so the failure looks like an empty result set rather than an error. The Docker route publishes on a loopback port only, mapping `127.0.0.1:8080` to the container, with the secret passed in as an environment variable, and a compose file is included in the repository for anyone who prefers it. A compose-based install also asks you to generate a secret rather than invent one, using a single `python3` line that prints a 32-byte hex token, and then to copy the example environment file and paste it in.

## Ten flags, and --urls-only is the one an agent wants

The default behaviour is a plain search returning the top ten results, and everything else is a flag on the same command. `--num` sets a limit, `--categories news` restricts to news, `--lang` filters by language, `--json` returns the raw payload, and `-e` takes a comma-separated engine list so you can pin `google,brave` rather than letting SearxNG choose. The one that matters most for agent use is `--urls-only`, which is documented specifically as something to pipe into a fetch step. Output is a numbered list where each entry carries a title, a URL, a snippet, and a bracketed tag naming which engines returned it, which is useful for a reader judging how much weight a result has. The multi-step workflow the documentation sketches is three commands: search with a result limit, hand a promising URL to a fetch tool, and switch to news mode with a category and language filter when the question is time-sensitive.

## SEARCH_PROVIDER can be tavily, which puts the bill back

The configuration surface is three environment variables and one of them exists as an escape hatch you may not need. `SEARXNG_URL` defaults to `http://localhost:8080` and points at your instance. `SEARCH_PROVIDER` defaults to `searxng` and accepts either `searxng` or `tavily`. `TAVILY_API_KEY` has no default and is required only when the provider is switched to Tavily, with keys obtained from the vendor's site. So the tool is not strictly committed to the self-hosted path, which is worth knowing if you want a managed backend for a while, but using it that way reintroduces exactly the per-query cost and third-party logging the project set out to avoid. The two supported shells are shown side by side in the documentation, exporting the URL for the default path and the provider plus key for the opt-in one. The repository also ships a `searxng/` directory holding the compose file and environment template.

## Three MCP tools, one of which needs a Tavily key

The MCP server is a single Python file, `mcp/server.py`, launched by your client with `python3` and configured through a standard `mcpServers` block that carries the SearxNG URL and an optional Tavily key in its environment. It requires one package beyond the script, `pip install mcp`. Three tools are exposed: `web_search` for general search through SearxNG, `web_search_news` for news through SearxNG, and `web_search_tavily` for search through the Tavily API, which is the one that needs a key. Antigravity is listed as an MCP client, Claude Code as a plain command, OpenClaw as a skill installed by copying `SKILL.md` into its skills directory or by pointing its skill loader at the repository URL, and anything with a shell can call the `ask-search` binary directly. The repository itself is small: `scripts/core.py` holds the main logic and CLI entry point, `install.sh` is the installer, and `SKILL.md` is the OpenClaw descriptor.

## Search always works, and the deep dive is where it breaks

The limitations section is the most useful part of the documentation, because it separates two things people conflate. ask-search returns URLs and snippets from search engine indexes, and the search itself is reliable, because SearxNG queries search engines rather than the target sites, and those engines have already indexed the content. The problem appears only when the agent tries to fetch the full page. The table is explicit: most sites work for both search and fetch because they have no aggressive anti-bot, while Reddit blocks datacenter IPs on fetch, Zhihu presents a login wall plus browser fingerprinting that needs JavaScript and an account, and Medium returns partial content behind a paywall. So a deep-dive pipeline built on this will search fine across all three and then fail on fetch for two of them, which is a different failure mode from a search that returns nothing and one you can plan around.

## A SOCKS tunnel to a residential IP is the documented workaround

The first solution offered is to route fetches through a residential network instead of a datacentre one. If you have a machine at home, a home server or a laptop, you open an SSH SOCKS tunnel from the server with dynamic forwarding on a local port, and then point curl at it with the `socks5h` scheme so the hostname resolves on the far side. For Reddit there is a specific shortcut worth knowing: appending `.json` to any post URL returns the full post and its comments as structured data, which removes the HTML scraping problem entirely. A systemd unit is provided as a template, with keepalive options, a forward-failure guard, and a description field, so the tunnel can be made to come back after a reboot. The SearxNG setup notes add two more operational items in the same spirit: bot detection can block requests from some addresses, and the fix is adding your server IP to the allow list in `limiter.toml`, and binding to 127.0.0.1 is the default recommendation unless you actually need remote access.

## Conclusion

ask-search suits anyone whose agent searches the web often enough that per-query pricing adds up, or who does not want their agent's queries landing in someone else's logs, since running SearxNG yourself moves both problems at once. It is the wrong tool if you have no machine to host SearxNG, because the whole design assumes a local instance on port 8080 that you keep patched, and the bot-detection notes are a standing maintenance item rather than a one-time setup. Before you wire it into an agent, do three things: enable `json` in the search formats or every call fails quietly, bind SearxNG to 127.0.0.1 unless you genuinely need remote access, and read the deep-dive table, because it sets your expectation that search results are reliable while page fetches against Reddit, Zhihu and Medium are not.

## FAQ

### What do I need before ask-search works?

A SearxNG instance you host yourself, and one configuration change on it: add json to the search formats in settings.yml so results are machine readable. Then run bash install.sh and call ask-search. The fast path is a clone, the install script and a single query if SearxNG is already running.

### Which agents can use ask-search?

Four integration paths are documented. OpenClaw uses it as a CLI skill installed from SKILL.md, Claude Code as a CLI command, Antigravity through the MCP server, and any shell can run the ask-search binary directly. The MCP server exposes web_search, web_search_news and web_search_tavily, and requires pip install mcp.

### Does ask-search cost anything to run?

Not on the default path, since SearxNG is self-hosted and aggregates engines itself. One paid provider is built in as an option: set SEARCH_PROVIDER=tavily and supply TAVILY_API_KEY, which reintroduces per-query cost. The tool exists specifically to avoid Brave, Google and Bing API pricing.

### Why can my agent search a page but not read it?

Because SearxNG queries search engines rather than the target sites, so search works even where a direct fetch does not. The documentation lists Reddit as blocked on datacenter IPs, Zhihu as behind a login wall and browser fingerprinting, and Medium as partially paywalled. The documented fix for the first two is an SSH SOCKS tunnel through a residential IP, and for Reddit, appending .json to a post URL returns the full post and comments as data.

## Sources

- [Issues](https://github.com/ythx-101/ask-search/issues)
- [License: MIT](https://github.com/ythx-101/ask-search/blob/master/LICENSE)
- [README](https://github.com/ythx-101/ask-search/blob/master/README.md)
- [Releases](https://github.com/ythx-101/ask-search/releases)
- [ythx-101/ask-search on GitHub](https://github.com/ythx-101/ask-search)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/ythx-101-ask-search
