# AutoHunter puts an LLM between FOFA and nmap, nuclei and sqlmap

> AutoHunter is a self-hosted SRC hunting platform that feeds FOFA assets to per-target LLM workers running a fixed scanner toolchain, then asks a human to adjudicate. The scanning is real; the scope handling around the mass-verification stage is the part an operator has to think about.

**StanleyNull/AutoHunter** — Automated SRC vulnerability-hunting system combining FOFA asset mapping with multiple LLM workers for autonomous discovery, review and intelligence accumulation. Powered by StanleyNull.

- Repository: https://github.com/StanleyNull/AutoHunter
- Stars: 469 · Forks: 69
- Language: Python
- License: Apache-2.0
- Published: 2026-09-17 · Updated: 2026-09-17 · Language: en
- Canonical page: https://hysenlabs.com/projects/stanleynull-autohunter

## One Worker per target, with the model choosing what runs

AutoHunter splits the work across four roles that pass targets along a fixed path. A Collector pulls assets from FOFA or from manual entry, probes whether they are alive, scores them, tags the owner, and queues them. Each queued target gets its own Worker, a one to one mapping the project insists on, and that Worker hands the job to a language model that decides what to run next. Inside the container the model drives nmap, nuclei, sqlmap, httpx and whatweb, so the requests come from those binaries rather than from generated text. A Reviewer sits between the Worker and a person: it sends work back for deeper digging, or drops it as a false positive. What reaches the review screen is a finding with evidence attached.

The README is written in Chinese, and the contributor table at the top credits named people for the engines, the agent UX, the LLM pool, the FOFA integration, workdir cleanup, a DNS probe and a proxy pool. Ownership of the scanning behaviour is spread across people who each own one layer, so judging the output means judging that layering. The last push to the main branch was on 2026-10-01 and the repository publishes no GitHub releases, which leaves no versioned artifact to pin. You build from the branch tip.

## Why running nuclei on its own is a different job

Run nuclei directly and the division of labour changes. ProjectDiscovery nuclei is a template runner: you pick which templates execute against which targets, and the YAML decides what gets sent. AutoHunter keeps that binary in the image at version 3.3.7 and inserts a model above it. A model reads a target, works out what to probe, and invokes nuclei, nmap, sqlmap and httpx as tools instead of following a fixed template list. A Reviewer filters that output afterwards and returns anything it judges incomplete to the Worker.

That arrangement is the point of the project and also its cost. Template driven scanning is reproducible: the same template against the same target produces the same requests, and you can read the file to know in advance what happens. A model in the loop removes that property. Nothing in the repository fixes the probe sequence, so two runs against one target can diverge. FOFA on its own covers the other half differently, since it is the asset mapping service and AutoHunter wraps it in a Collector that adds liveness probing, scoring and owner tagging before a target ever reaches a Worker.

## LLM_API_KEY is the only value the wizard insists on

One value blocks startup, and the rest of the channel settings decide how much of the platform works:

```bash
LLM_BASE_URL=https://api.deepseek.com/v1
LLM_API_KEY=
LLM_MODEL=deepseek-chat
LLM_TEMPERATURE=0.3
LLM_MAX_TOKENS=4096
AUTOHUNTER_TOOL_COMPAT=auto
LLM_PROVIDER_MODE=single
LLM_PROTOCOL=auto
LLM_PROVIDERS_JSON=[]
LLM_PROVIDER_FAIL_THRESHOLD=5
```

The .env.example header calls LLM_API_KEY the minimum needed to boot, with FOFA_KEY and the access tokens marked as strongly recommended rather than required. Skip FOFA and the platform still runs, but the Collector has no asset source, so targets must be entered by hand.

LLM_PROTOCOL accepts auto, openai_chat or anthropic_messages, and the Worker, the Reviewer and the report assistant share that one channel. AUTOHUNTER_TOOL_COMPAT matters most behind a gateway. Auto prefers native tool calling and falls back to a prompt simulation only when the endpoint rejects the tools parameter, which covers models that accept a tools argument and then never call a tool. LLM_PROVIDER_MODE switches between one endpoint and a pool declared in LLM_PROVIDERS_JSON, where each entry carries name, base_url, api_key, model, protocol, temperature, weight and enabled. The visible file stops mid key at LLM_PROVIDER_BEHAVI, so the behaviour tuning after that point is not readable here.

## The Dockerfile pins the scanners while the Python libraries float

Two variables decide which scanner binaries land in the image:

```dockerfile
    NUCLEI_VER=3.3.7; HTTPX_VER=1.6.9; \
```

Both are fetched as zip archives from upstream release URLs, with ghfast.top and ghproxy.net tried before a direct fetch. A helper runs unzip -tq on the archive and fails the build when the download is not a valid zip, so a blocked mirror leaves a build error rather than an image quietly missing a tool. TARGETARCH is injected by buildkit, which resolves the archive name to linux_amd64 or linux_arm64 and nothing else. sqlmap comes from the monthly PyPI release instead of a clone of its repository, a choice the Dockerfile attributes to build reliability.

requirements.txt goes the other way. openai>=1.40.0, fastapi>=0.110.0, sqlalchemy>=2.0.0 and pydantic>=2.0.0 are floors with no ceilings, and httpx is the only capped dependency at httpx[socks]>=0.27.0,<0.29. A rebuild months from now resolves those floors to whatever has since been released, while the scanners stay where they are.

## websockets>=13.0 fences browser automation out of the container

One line in requirements.txt carries an operational warning. websockets>=13.0 has a floor because uvicorn 0.51 and later need it. Immediately above it sits a comment naming the danger: packages such as pyppeteer and undetected-chromedriver pull older versions of websockets without a constraint, so installing either inside the container drags the dependency below 13 and uvicorn stops working.

One line further down, httpx tells the same story from the other side. Its comment records that the proxy pool depends on SOCKS5 support, which arrives through the socks extra, and asks that it not be reverted to a plain httpx.

Together those lines put browser automation outside the runtime, in a project whose whole purpose is automated reconnaissance. A reviewer who wants the workers to drive a headless browser has nowhere to put it: the fix that would enable it is the same fix that takes down the server. Three lines of a dependency file therefore hold the container's real constraints, the version floor for uvicorn, the prohibition on browser packages, and the extra that has to stay.

## The authors say the model may delete files on its own initiative

The bluntest statement in the project about its own behaviour is the one that should stop a reader. A warning block near the top says delete restrictions are in place, but that under some external conditions, including the model's own inclination, it may still modify or delete a small amount of data, and the author asks for issue reports rather than blame.

Set that beside the compose file. The worker workspace is a named volume mounted at /work, driven by WORKER_WORK_ROOT, described as the temporary working area for workers. A temporary area held in a Docker volume outlives the container, so anything removed there stays removed after a restart. SQLite lives in a second volume at /app/data instead, and the compose file states that restart recovery depends on it, which is why the settings page can download and restore an online consistent snapshot.

The consequence for an operator is that permission to scan one asset and permission for an agent to write to disk are two separate grants. The README asks for the first and admits uncertainty about the second.

## Mass verification reaches past the single authorized target

After a finding clears review, AutoHunter runs a stage the README calls the Hunter killer. It analyses whether one vulnerability generalises, then actively verifies the same issue against multiple sites running the same stack, and only then hands the result to a person. The same verified credentials, endpoints and fingerprints also go into a shared intelligence store that later Workers read from.

This is where scope grows past the target in front of you. Written authorization is the project's own precondition, and it is normally scoped to a host or a program boundary. A stage whose purpose is to hit several other hosts in order to prove the bug repeats leaves that boundary by design, and it runs unattended overnight. Nothing in the visible text asks for per host permission before the mass verification step, and the cautions section that the table of contents lists is not readable in this copy of the file.

The decision is not whether the technique is sound. It is whether the authorization you hold covers the extra hosts that stage will touch, answered before the first overnight run rather than after.

## scripts/install.sh builds from the branch tip and 18800 only moves on the host

Installation is two commands once Docker is present:

```bash
# 1. 装 Docker（官方一键脚本，已装可跳过）
curl -fsSL https://get.docker.com | sh && sudo systemctl enable --now docker
sudo usermod -aG docker $USER && newgrp docker      # 免 sudo，重登生效

# 2. 拉代码 + 一键部署（交互式引导：填 LLM/FOFA Key → 自动生成令牌 → 构建启动）
git clone https://github.com/StanleyNull/AutoHunter.git autohunter && cd autohunter
bash scripts/install.sh
```

The wizard checks Docker, asks for the LLM key, treats the FOFA key as optional, generates a high strength access token, builds and starts the container, then prints the address and the token. Building for the first time compiles the frontend and installs the scanning tools, which the README puts at 5 to 15 minutes.

Port 18800 is the only published port, and only half of it moves. Only half of that port moves: the compose file parameterises the host side through AUTOHUNTER_HOST_PORT, defaulting to 18800, while the container side is a literal 18800. Change the host port and the firewall lines printed in the README stop matching what the container publishes:

```bash
sudo ufw allow 18800/tcp                                              # Ubuntu/Debian
sudo firewall-cmd --permanent --add-port=18800/tcp && sudo firewall-cmd --reload   # CentOS/RHEL
```

Cloud instances need the same port opened in the vendor security group as well.

## Conclusion

AutoHunter fits a researcher who already holds written authorization for a defined target list and wants a model to drive a fixed scanner toolchain overnight instead of watching alerts. It does not fit an operator after reproducible coverage, since the model picks the probe order each run. Before the first run, confirm the authorization covers every host the mass-verification stage will touch, and confirm your endpoint accepts native tool calling, because AUTOHUNTER_TOOL_COMPAT=auto drops to prompt simulation when it does not.

## FAQ

### What does AutoHunter need before it will start?

One value blocks startup, LLM_API_KEY, which the .env.example header calls the minimum needed to boot. FOFA_KEY and the access tokens are marked as strongly recommended instead. The wizard behind scripts/install.sh generates the access tokens itself during the guided setup.

### Which scanning tools does AutoHunter ship in the container?

The image installs nmap, sqlmap, nuclei and whatweb, and it installs the ProjectDiscovery httpx binary from a release archive rather than through pip. The nuclei and httpx binaries are pinned by the NUCLEI_VER and HTTPX_VER variables, set to 3.3.7 and 1.6.9 in the Dockerfile, while the Python httpx library is constrained separately in requirements.txt.

### Does AutoHunter publish releases I can pin a deployment to?

The repository has no GitHub releases, so there is no versioned artifact to install. The compose file builds from the local clone with build: . and tags the result autohunter:latest, which means each rebuild or pull resolves to whatever the main branch currently holds. The last push was on 2026-10-01.

### Can AutoHunter run on Windows?

The README describes a Windows path through Docker Desktop with WSL2, and it requires an installed Linux distribution rather than only the bundled docker-desktop one, with wsl --install named as the command. That section of the file ends mid sentence at a link to Docker Desktop, so the remaining Windows steps are not readable.

## Sources

- [Issues](https://github.com/StanleyNull/AutoHunter/issues)
- [License: Apache-2.0](https://github.com/StanleyNull/AutoHunter/blob/main/LICENSE)
- [README](https://github.com/StanleyNull/AutoHunter/blob/main/README.md)
- [StanleyNull/AutoHunter on GitHub](https://github.com/StanleyNull/AutoHunter)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/stanleynull-autohunter
