# PentestAgent: AI-Driven Black-Box Security Testing with Multi-Agent Orchestration

> PentestAgent is a Python framework that uses large language models to drive black-box penetration testing workflows. It supports Anthropic, OpenAI, and any LiteLLM-compatible provider, runs through a terminal UI, and allows a parent agent to spawn isolated child agents over stdio for parallel reconnaissance and exploitation tasks.

**GH05TCREW/pentestagent** — PentestAgent is an AI agent framework for black-box security testing, supporting bug bounty, red-team, and penetration testing workflows.

- Repository: https://github.com/GH05TCREW/pentestagent
- Stars: 3,114 · Forks: 617
- Language: Python
- License: MIT
- Published: 2026-09-09 · Updated: 2026-09-09 · Language: en
- Canonical page: https://hysenlabs.com/projects/gh05tcrew-pentestagent

## What Problem PentestAgent Addresses

Traditional penetration testing requires a practitioner to manually choose tools, run them in sequence, interpret output, and decide what to probe next. That decision loop is time-consuming and depends heavily on individual expertise. PentestAgent replaces the manual orchestration layer with an LLM that reads tool output, decides what to run next, and delegates subtasks to child agents when the scope justifies parallel work.

The intended users, as described in the pyproject.toml keywords, are security professionals conducting penetration tests, bug bounty hunters, and red teamers working against targets they have explicit authorization to test. The framework operates in black-box mode: it does not assume prior knowledge of the target's internals and works from what its tools discover at runtime.

## Four Modes from Single Instruction to Multi-Agent Crew

PentestAgent exposes four operating modes through commands typed into its terminal UI. The Assist mode (`/assist <task>`) sends a single instruction with tool execution and returns a result. The Agent mode (`/agent <task>`) runs autonomously on a single task from start to finish. The Crew mode (`/crew <task>`) activates a multi-agent setup where an orchestrator spawns specialized worker agents. The Interact mode (`/interact <task>`) opens a guided conversation where the agent provides context and asks follow-up questions during the engagement.

Additional TUI commands include `/target <host>` to set the active target, `/tools` to list available tools, `/notes` to display saved session notes, `/report` to generate a report from the session, `/conversations` to browse and restore previous sessions, and `/mcp` to list or add MCP servers. Press `Esc` to stop a running agent and `Ctrl+Q` to quit.

## Installing PentestAgent and Pointing It at an LLM

The project requires Python 3.10 or later. Clone the repository and run the setup script for your platform:

```bash
git clone https://github.com/GH05TCREW/pentestagent.git
cd pentestagent
./scripts/setup.sh
```

For Windows, use `scripts\setup.ps1` instead. The script creates a virtual environment and installs all dependencies. To install manually:

```bash
python -m venv venv
source venv/bin/activate
pip install -e ".[all]"
playwright install chromium
```

Chromium is required for the browser tool. After install, create a `.env` file in the project root. For Anthropic:

```
ANTHROPIC_API_KEY=sk-ant-...
PENTESTAGENT_MODEL=claude-sonnet-4-20250514
```

For OpenAI:

```
OPENAI_API_KEY=sk-...
PENTESTAGENT_MODEL=gpt-5
```

Any LiteLLM-supported model string works, including local Ollama models (`ollama/qwen2.5:7b-instruct`) and custom relay endpoints via `OPENAI_API_BASE`. See `.env.example` in the repository root for the full list of supported environment variables. Once configured, launch the TUI:

```bash
pentestagent
```

To start with a target already set:

```bash
pentestagent -t 192.168.1.1
```

## Running Tools in Docker: Base Image and Kali Image

The repository ships two Docker images. The base image, available as `ghcr.io/gh05tcrew/pentestagent:latest`, includes nmap, netcat, curl, openvpn, and wireguard-tools. The Kali image at `ghcr.io/gh05tcrew/pentestagent:kali` adds heavier offensive tools including metasploit, sqlmap, and hydra.

To pull and run the base image:

```bash
docker run -it --rm \
  -e ANTHROPIC_API_KEY=your-key \
  -e PENTESTAGENT_MODEL=claude-sonnet-4-20250514 \
  ghcr.io/gh05tcrew/pentestagent:latest
```

For the Kali image:

```bash
docker run -it --rm \
  -e ANTHROPIC_API_KEY=your-key \
  ghcr.io/gh05tcrew/pentestagent:kali
```

Or build locally with docker-compose:

```bash
docker compose build
docker compose run --rm pentestagent
```

The docker-compose.yml maps a `./loot` volume for output and exposes port 8080. The Kali image configuration in docker-compose.yml requires `privileged: true` with `NET_ADMIN` and `SYS_ADMIN` capabilities. The file notes explicitly: "this is risky on shared hosts; prefer running inside a disposable VM."

## How spawn_mcp_agent Enables Hierarchical Pentesting

The `spawn_mcp_agent` tool is the mechanism behind the Crew mode and the `/spawn` TUI command. When called, it starts a child copy of PentestAgent as a subordinate MCP server over stdio. The child process has its own runtime, LLM client, conversation history, and notes store. After the child starts, its tool set becomes available to the parent agent on the next tool call. This allows the parent to delegate scoped subtasks to isolated children without external orchestration infrastructure.

The README documents a concrete pattern where two child agents run parallel reconnaissance on different network ranges:

```
spawn_mcp_agent  target="10.0.1.0/24"  scope=["10.0.1.0/24"]
spawn_mcp_agent  target="10.0.2.0/24"  scope=["10.0.2.0/24"]
```

After both children start, the parent calls async task methods on each child, then collects results with `await_tasks` and a timeout. The `/despawn` TUI command terminates a child manually and removes its tools from the session. The `no_mcp` parameter defaults to `true` for child agents, which the README recommends to prevent children from also connecting to external MCP servers.

## Playbooks for Structured, Repeatable Assessments

Beyond free-form agent sessions, PentestAgent includes prebuilt attack playbooks for structured black-box assessments. A playbook defines the approach for a specific type of security evaluation. Run one with:

```bash
pentestagent run -t example.com --playbook thp3_web
```

Playbooks provide a more controlled alternative to open-ended Crew or Agent sessions when the engagement scope maps to a known assessment pattern. The README documents `thp3_web` as an example; the full list of available playbooks is not documented in the README.

## Limitations and Comparison with Script-Based Automation

PentestAgent has several real constraints. The Kali Docker image requires privileged mode, which is inappropriate on shared infrastructure. Web search functionality requires a Tavily API key, an external service that must be provisioned separately. The project carries a Development Status: 3 - Alpha classifier in pyproject.toml, meaning the interface is not yet stable and breaking changes should be expected. The version is 0.2.0 and the repository has no GitHub releases, so there are no versioned release artifacts to pin to.

A natural comparison is with Metasploit, the widely used open-source penetration testing framework maintained by Rapid7. Metasploit provides a large library of exploits, payloads, and post-exploitation modules that practitioners run manually through msfconsole. The key difference in approach is that Metasploit requires the tester to select modules, set options, and chain tasks explicitly. PentestAgent delegates those decisions to an LLM, which selects and sequences tools dynamically based on what prior tool runs return. This makes PentestAgent faster for exploration but less predictable for controlled, reproducible test runs where each step must be documented precisely before execution.

## Development Status and License

The last push was on 2026-09-21. The repository has no GitHub releases. The project is MIT-licensed. The built-in tools at the time of this writing are: `terminal`, `browser`, `notes`, `web_search` (requires a Tavily API key), and `spawn_mcp_agent`. MCP server extensibility is available via the `/mcp` command in the TUI. The `loot/` directory in the repository root is the default output volume for Docker runs, providing a persistent location for session artifacts and tool output across container restarts.

## Conclusion

PentestAgent is suited to security professionals who want an LLM to orchestrate tool selection and task delegation during engagements, rather than scripting that logic manually. It is not suited to production environments where uncontrolled outbound network activity is a concern, nor to teams without legal authorization for the target, since the framework executes real tools against real systems. The Kali Docker image uses privileged mode, which the docker-compose.yml notes is risky on shared hosts and should run inside a disposable VM. Version 0.2.0 carries a Development Status: 3 - Alpha classifier, so the API surface should be treated as unstable.

## FAQ

### Does PentestAgent require an internet connection to function?

The LLM backend requires connectivity to your chosen API provider. The web_search tool additionally requires a Tavily API key. The terminal tool and local Docker-based tools can run without external connectivity once the container is built, but the LLM calls themselves will always reach out to the configured provider endpoint.

### Can PentestAgent use a local LLM instead of a cloud provider?

Yes. The .env.example shows `PENTESTAGENT_MODEL=ollama/qwen2.5:7b-instruct` as a supported model string, and any LiteLLM-supported model string can be used. The quality of the agent's decisions will depend on the capability of the local model.

### Is PentestAgent safe to run against live production systems?

The README does not address this question. The framework executes real tools against whatever target you specify. Using it against systems without explicit authorization is illegal in most jurisdictions, and the Kali image's privileged Docker mode adds additional infrastructure risk on shared hosts.

## Sources

- [GH05TCREW/pentestagent on GitHub](https://github.com/GH05TCREW/pentestagent)
- [Issues](https://github.com/GH05TCREW/pentestagent/issues)
- [License: MIT](https://github.com/GH05TCREW/pentestagent/blob/main/LICENSE)
- [README](https://github.com/GH05TCREW/pentestagent/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/gh05tcrew-pentestagent
