PentestAgent: AI-Driven Black-Box Security Testing with Multi-Agent Orchestration
PentestAgent is an AI agent framework for black-box security testing, supporting bug bounty, red-team, and penetration testing workflows.
At a glance
- What is it?
- PentestAgent is a Python framework that uses large language models to drive black-box penetration testing workflows. It supports Anthropic, OpenAI, and any LiteLLM-compatible provider, runs through a terminal UI, and allows a parent agent to spawn isolated child agents over stdio for parallel reconnaissance and exploitation tasks.
- Who is it for?
- PentestAgent is suited to security professionals who want an LLM to orchestrate tool selection and task delegation during engagements, rather than scripting that logic manually. It is not suited to production environments where uncontrolled outbound network activity is a concern, nor to teams without legal authorization for the target, since the framework executes real tools against real systems.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 10 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 25, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What Problem PentestAgent Addresses
Traditional penetration testing requires a practitioner to manually choose tools, run them in sequence, interpret output, and decide what to probe next. That decision loop is time-consuming and depends heavily on individual expertise. PentestAgent replaces the manual orchestration layer with an LLM that reads tool output, decides what to run next, and delegates subtasks to child agents when the scope justifies parallel work.
The intended users, as described in the pyproject.toml keywords, are security professionals conducting penetration tests, bug bounty hunters, and red teamers working against targets they have explicit authorization to test. The framework operates in black-box mode: it does not assume prior knowledge of the target's internals and works from what its tools discover at runtime.
Four Modes from Single Instruction to Multi-Agent Crew
PentestAgent exposes four operating modes through commands typed into its terminal UI. The Assist mode (`/assist <task>`) sends a single instruction with tool execution and returns a result. The Agent mode (`/agent <task>`) runs autonomously on a single task from start to finish. The Crew mode (`/crew <task>`) activates a multi-agent setup where an orchestrator spawns specialized worker agents. The Interact mode (`/interact <task>`) opens a guided conversation where the agent provides context and asks follow-up questions during the engagement.
Additional TUI commands include `/target <host>` to set the active target, `/tools` to list available tools, `/notes` to display saved session notes, `/report` to generate a report from the session, `/conversations` to browse and restore previous sessions, and `/mcp` to list or add MCP servers. Press `Esc` to stop a running agent and `Ctrl+Q` to quit.
Installing PentestAgent and Pointing It at an LLM
The project requires Python 3.10 or later. Clone the repository and run the setup script for your platform:
git clone https://github.com/GH05TCREW/pentestagent.git
cd pentestagent
./scripts/setup.shFor Windows, use `scripts\setup.ps1` instead. The script creates a virtual environment and installs all dependencies. To install manually:
python -m venv venv
source venv/bin/activate
pip install -e ".[all]"
playwright install chromiumChromium is required for the browser tool. After install, create a `.env` file in the project root. For Anthropic:
ANTHROPIC_API_KEY=sk-ant-...
PENTESTAGENT_MODEL=claude-sonnet-4-20250514For OpenAI:
OPENAI_API_KEY=sk-...
PENTESTAGENT_MODEL=gpt-5Any LiteLLM-supported model string works, including local Ollama models (`ollama/qwen2.5:7b-instruct`) and custom relay endpoints via `OPENAI_API_BASE`. See `.env.example` in the repository root for the full list of supported environment variables. Once configured, launch the TUI:
pentestagentTo start with a target already set:
pentestagent -t 192.168.1.1Running Tools in Docker: Base Image and Kali Image
The repository ships two Docker images. The base image, available as `ghcr.io/gh05tcrew/pentestagent:latest`, includes nmap, netcat, curl, openvpn, and wireguard-tools. The Kali image at `ghcr.io/gh05tcrew/pentestagent:kali` adds heavier offensive tools including metasploit, sqlmap, and hydra.
To pull and run the base image:
docker run -it --rm \
-e ANTHROPIC_API_KEY=your-key \
-e PENTESTAGENT_MODEL=claude-sonnet-4-20250514 \
ghcr.io/gh05tcrew/pentestagent:latestFor the Kali image:
docker run -it --rm \
-e ANTHROPIC_API_KEY=your-key \
ghcr.io/gh05tcrew/pentestagent:kaliOr build locally with docker-compose:
docker compose build
docker compose run --rm pentestagentThe docker-compose.yml maps a `./loot` volume for output and exposes port 8080. The Kali image configuration in docker-compose.yml requires `privileged: true` with `NET_ADMIN` and `SYS_ADMIN` capabilities. The file notes explicitly: "this is risky on shared hosts; prefer running inside a disposable VM."
How spawn_mcp_agent Enables Hierarchical Pentesting
The `spawn_mcp_agent` tool is the mechanism behind the Crew mode and the `/spawn` TUI command. When called, it starts a child copy of PentestAgent as a subordinate MCP server over stdio. The child process has its own runtime, LLM client, conversation history, and notes store. After the child starts, its tool set becomes available to the parent agent on the next tool call. This allows the parent to delegate scoped subtasks to isolated children without external orchestration infrastructure.
The README documents a concrete pattern where two child agents run parallel reconnaissance on different network ranges:
spawn_mcp_agent target="10.0.1.0/24" scope=["10.0.1.0/24"]
spawn_mcp_agent target="10.0.2.0/24" scope=["10.0.2.0/24"]After both children start, the parent calls async task methods on each child, then collects results with `await_tasks` and a timeout. The `/despawn` TUI command terminates a child manually and removes its tools from the session. The `no_mcp` parameter defaults to `true` for child agents, which the README recommends to prevent children from also connecting to external MCP servers.
Playbooks for Structured, Repeatable Assessments
Beyond free-form agent sessions, PentestAgent includes prebuilt attack playbooks for structured black-box assessments. A playbook defines the approach for a specific type of security evaluation. Run one with:
pentestagent run -t example.com --playbook thp3_webPlaybooks provide a more controlled alternative to open-ended Crew or Agent sessions when the engagement scope maps to a known assessment pattern. The README documents `thp3_web` as an example; the full list of available playbooks is not documented in the README.
Limitations and Comparison with Script-Based Automation
PentestAgent has several real constraints. The Kali Docker image requires privileged mode, which is inappropriate on shared infrastructure. Web search functionality requires a Tavily API key, an external service that must be provisioned separately. The project carries a Development Status: 3 - Alpha classifier in pyproject.toml, meaning the interface is not yet stable and breaking changes should be expected. The version is 0.2.0 and the repository has no GitHub releases, so there are no versioned release artifacts to pin to.
A natural comparison is with Metasploit, the widely used open-source penetration testing framework maintained by Rapid7. Metasploit provides a large library of exploits, payloads, and post-exploitation modules that practitioners run manually through msfconsole. The key difference in approach is that Metasploit requires the tester to select modules, set options, and chain tasks explicitly. PentestAgent delegates those decisions to an LLM, which selects and sequences tools dynamically based on what prior tool runs return. This makes PentestAgent faster for exploration but less predictable for controlled, reproducible test runs where each step must be documented precisely before execution.
Development Status and License
The last push was on 2026-09-21. The repository has no GitHub releases. The project is MIT-licensed. The built-in tools at the time of this writing are: `terminal`, `browser`, `notes`, `web_search` (requires a Tavily API key), and `spawn_mcp_agent`. MCP server extensibility is available via the `/mcp` command in the TUI. The `loot/` directory in the repository root is the default output volume for Docker runs, providing a persistent location for session artifacts and tool output across container restarts.
Editorial conclusion
PentestAgent is suited to security professionals who want an LLM to orchestrate tool selection and task delegation during engagements, rather than scripting that logic manually. It is not suited to production environments where uncontrolled outbound network activity is a concern, nor to teams without legal authorization for the target, since the framework executes real tools against real systems. The Kali Docker image uses privileged mode, which the docker-compose.yml notes is risky on shared hosts and should run inside a disposable VM. Version 0.2.0 carries a Development Status: 3 - Alpha classifier, so the API surface should be treated as unstable.
Frequently asked questions
Does PentestAgent require an internet connection to function?
The LLM backend requires connectivity to your chosen API provider. The web_search tool additionally requires a Tavily API key. The terminal tool and local Docker-based tools can run without external connectivity once the container is built, but the LLM calls themselves will always reach out to the configured provider endpoint.
Can PentestAgent use a local LLM instead of a cloud provider?
Yes. The .env.example shows `PENTESTAGENT_MODEL=ollama/qwen2.5:7b-instruct` as a supported model string, and any LiteLLM-supported model string can be used. The quality of the agent's decisions will depend on the capability of the local model.
Is PentestAgent safe to run against live production systems?
The README does not address this question. The framework executes real tools against whatever target you specify. Using it against systems without explicit authorization is illegal in most jurisdictions, and the Kali image's privileged Docker mode adds additional infrastructure risk on shared hosts.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/gh05tcrew-pentestagent)