Model or dataset
Netw0rkNoob/VulnClaw avatar
Netw0rkNoob/VulnClaw

VulnClaw: an LLM-driven pentest CLI with MCP tools and evidence gating

基于 AI Agent + MCP 工具链 + 渗透 Skill 编排, 配合大语言模型, 自然语言输入 → 自动完成「信息收集 → 漏洞发现 → 漏洞利用 → 报告生成」全流程。

3,462 stars479 forksPythonMIT

At a glance

What is it?
VulnClaw turns a natural-language target description into a recon-to-report loop, driven by an OpenAI-compatible model over MCP tool servers. The interesting part is not the loop, it is the evidence gate that refuses to accept a flag the tools never printed.
Who is it for?
VulnClaw fits authorized engagements, CTF rooms and lab teaching where a model can be given a tool budget and its claims are checked against raw output. It does not fit teams that need a signed, repeatable scan report, because the model chooses the next step and two runs will not match.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 7 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What VulnClaw actually automates, and for whom

The README describes VulnClaw as a standalone AI penetration testing agent that takes a natural-language instruction such as a target URL and runs information gathering, vulnerability discovery, exploitation and report generation in sequence. The intended users are named explicitly: authorized penetration tests, CTF competitions, security teaching and red team exercises. The project ships a security scope badge reading Authorized Only, and the README carries a security statement section, so the author is not pretending this is a general-purpose scanner you point at anything.

The practical shift is in the interface. Instead of writing a command per phase, you write one sentence and the model decides which tools to call. That is the same bet Claude Code and Codex make for software work, applied to a domain where a wrong tool call has consequences. VulnClaw's answer to that risk is not a fixed playbook. The README states that the model-led solver engine is the default and that the model itself decides when to call a tool and when to stop, while the older stage planner is not restored.

The evidence gate, AgentState and the stall guard

The mechanism worth understanding is how tool output is stored and how the model is allowed to use it. Every tool result is written into AgentState.evidence with the raw text preserved, and the active context by default receives only a high-signal preview. The model has to call evidence_search or evidence_view to pull the original back. That keeps the context window from filling with HTML, but it also means the model can lose track of what it already read.

The countermeasure is a light correction layer. Repeated reads of the same evidence range are suppressed, and consecutive rounds that produce no new evidence trigger a stall guard. The README is explicit that this layer does not resurrect the old stage planner, so it is a brake, not a steering wheel.

The strongest claim in the README is the anti-hallucination gate: a claimed flag or conclusion is only accepted if it appears character for character in real tool output. For CTF work that is the difference between a useful tool and a machine that invents a flag and stops. A second gate, described as the NO_PATH gate, refuses to let the model declare a dead end while high-signal material such as a source sink, a form parameter, a request surface, a local proof or a response difference is still unexplored. Both gates are policy, not proof. They constrain what the model may assert, and they depend on the tools actually capturing the relevant bytes.

Installing VulnClaw and running a first check

The README gives two install paths, PyPI which it recommends, and a source checkout installed in editable mode. Both assume Python 3.10 or newer, which matches the requires-python field in pyproject.toml. Run the install command, then confirm the binary resolves before configuring anything.

bash
pip install vulnclaw

Configuration is a four-step sequence in the README. First pick a provider, which auto-fills the base URL and model name. Then set the API key. Then launch the CLI or the TUI. The provider command accepts names including minimax, openai, anthropic, deepseek, zhipu, moonshot, qwen, siliconflow and ollama.

bash
vulnclaw config provider minimax
vulnclaw config set llm.api_key sk-your-key-here
vulnclaw doctor

The doctor command is the step to run before anything else, because it reports Python, Node.js, npx and nmap availability alongside the resolved provider, auth mode, base URL and model, then lists MCP server status. The README's sample output shows fetch and memory as enabled at priority P0. If those two are not enabled, the model has no working tool surface.

For a container setup, the compose file publishes the Web UI on 127.0.0.1:7788 and persists state to a named volume mounted at /data. The README warns that localhost inside the container refers to the container, so host services must be reached through host.docker.internal.

bash
cp .env.example .env
docker compose up --build

Where the design breaks down

The README is candid about one component: python_execute is described as high-risk experimental and explicitly not a strong isolation sandbox. It runs payload construction and response parsing with the libraries listed in pyproject.toml, which include pycryptodome, lxml and beautifulsoup4. If you run VulnClaw on a workstation holding cloud credentials, that tool is the one to think about first.

The second limitation is reproducibility. Because the default engine is model-led with no fixed round count, the same target can produce different tool sequences on different runs. That is fine for a CTF where you want the flag, and awkward for an engagement where the client expects a report that another engineer can regenerate.

The third is dependency pinning. pyproject.toml pins the MCP client to mcp>=1.0,<2.0 with a comment explaining that the lifecycle and probe code uses the 1.x symbol streamablehttp_client, renamed in mcp 2.0. The pin is deliberate, but it means an upstream major version is a migration, not an upgrade.

Finally, the knowledge base. The README says the knowledge base module and seed data exist and that retrieval augmentation is being wired into the main flow gradually. Treat the kb extra as present but not yet load-bearing.

VulnClaw against a scripted scanner

The obvious alternative is a scripted scanner plus a human operator, or a framework such as an nmap and Nuclei pipeline driven by a Makefile. The difference is where the decisions live. In a scripted pipeline the sequence is fixed in advance and the output is deterministic; a template either matches or it does not. In VulnClaw the sequence is chosen at runtime by the model, guided by evidence it has read and by the two gates. You trade determinism for the ability to follow an unexpected thread, such as a filter that rejects a payload the parser still accepts.

That trade is not free in either direction. A scripted scanner will not notice that a response body changed shape between two probes unless someone wrote a check for it. VulnClaw has runtime_diff_probe for exactly that class of problem, per the README, and it flags PHP serialization candidates as requiring remote verification rather than trusting a local PHP version. A scripted pipeline would encode that rule once and apply it every time. VulnClaw asks the model to reach for the tool.

Licence, maintenance and the cost of upgrading

VulnClaw is MIT licensed, stated in both the README badge and the pyproject.toml license field. MIT permits commercial use and modification with attribution and without warranty. The repository also ships a SECURITY.md and a CODE_OF_CONDUCT.md, and the project has a Discord community linked from the README. None of that is a support contract, and the licence text itself is the only thing that governs what you may do with the code.

On maintenance, the last push was on 2026-09-05, and the most recent release in the list is v0.3.9 dated the same day, following v0.3.8 on 2026-08-09 and v0.3.7 on 2026-08-04. The repository is not archived. The release cadence in that window is roughly weekly to monthly, which tells you the author is still shipping, but the version number sitting below 0.4 tells you the interfaces are still moving.

Upgrade cost concentrates in two places. The MCP client pin means an mcp 2.0 release requires code changes rather than a version bump. Provider configuration is the other: .env.example warns that setting VULNCLAW_LLM_PROVIDER alone through environment variables does not auto-fill the base URL or model, unlike the interactive config command, so a mismatched key and endpoint produces an invalid api key error rather than a clear provider error. Anyone deploying through Docker should set base URL and model explicitly.

Editorial conclusion

VulnClaw fits authorized engagements, CTF rooms and lab teaching where a model can be given a tool budget and its claims are checked against raw output. It does not fit teams that need a signed, repeatable scan report, because the model chooses the next step and two runs will not match. Before adopting it, run vulnclaw doctor on the machine that will hold the API key, then check that the fetch and memory MCP servers start, since those two are the ones the README calls ready out of the box. The python_execute tool is the one to disable first if you cannot accept arbitrary code running with your own credentials.

Frequently asked questions

Which tool is commonly used for vulnerability assessment?

VulnClaw itself is positioned as a vulnerability assessment and exploitation CLI, and its README lists nmap among the tools the environment check looks for. Beyond that, the repository does not name a specific external scanner as its assessment engine.

What is VulnClaw?

It is an AI penetration testing CLI that takes natural-language input and runs information gathering, vulnerability discovery, exploitation and report generation. It drives the loop through an OpenAI-compatible model and MCP tool servers, with a separate Web UI mode started by vulnclaw web on 127.0.0.1:7788.

How do I install VulnClaw?

The README recommends pip install vulnclaw, or cloning the repository and running pip install -e . from the checkout. Python 3.10 or newer is required, and vulnclaw doctor should be run afterwards to confirm the model configuration and MCP servers.

Can VulnClaw run without an API key?

The README documents a vulnclaw login flow for ChatGPT subscription sign-in that needs no API key, and points to docs/keyless-auth.md while noting ToS risk. It also states that a local Ollama model works through the OpenAI-compatible endpoint with any placeholder key, provided the model supports tool calling.

Official sources

  1. License: MIT
  2. Netw0rkNoob/VulnClaw on GitHub
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/netw0rknoob-vulnclaw.svg)](https://hysenlabs.com/projects/netw0rknoob-vulnclaw)