# PentestGPT: what the USENIX-published agent actually installs and runs

> PentestGPT is an LLM-driven penetration testing framework from GreyDGL, published at USENIX Security 2024. Its v1.0 branch drives Claude Code or Codex as autonomous backends, while a modernized legacy mode keeps the human-in-the-loop workflow with many providers.

**GreyDGL/PentestGPT** — Automated Penetration Testing Agentic Framework Powered by Large Language Models

- Repository: https://github.com/GreyDGL/PentestGPT
- Stars: 15,644 · Forks: 2,721
- Language: Python
- License: MIT
- Published: 2026-09-09 · Updated: 2026-09-09 · Language: en
- Canonical page: https://hysenlabs.com/projects/greydgl-pentestgpt

## The gap PentestGPT fills between a scanner and a human operator

A conventional scanner enumerates known signatures. A human tester forms hypotheses, runs a tool, reads the output, and changes direction. PentestGPT sits in the second category: it is an agentic framework that uses a large language model to decide which tool to run next and how to interpret what comes back. The README describes it as an "AI-Powered Autonomous Penetration Testing Agent" and says the work was published at USENIX Security 2024, with the paper linked from the repository.

The intended users are CTF players and penetration testers working through a defined target. The README lists multi-category support covering Web, Crypto, Reversing, Forensics, PWN and Privilege Escalation, and shows targets in the 10.10.11.x range, which is the HackTheBox address space. The Dockerfile installs nmap, gobuster, dirb, netcat-openbsd, openvpn and tmux, so the container is meant to be a disposable working environment rather than a library you import. That framing matters: PentestGPT is not a reporting tool bolted onto a scan, it is an operator that happens to write a report at the end of one of its modes.

## Two engines in one repository: the agentic pipeline and pentestgpt-legacy

Version 1.0 introduced what the README calls the agentic upgrade, and the repository now ships two distinct entry points. The first is `pentestgpt`, an autonomous agent that works through a multi-stage pipeline. In CTF mode the stages are recon, exploit and walkthrough. In pentest mode they are asset discovery, vulnerability identification and report. Each stage's findings feed the next, and sessions can be saved and resumed.

The second entry point is `pentestgpt-legacy`, described as the modernized legacy mode. This preserves the human-in-the-loop design from the paper. It runs three cooperating LLM sessions, labelled reasoning, generation and parsing, and maintains a Pentesting Task Tree while the operator drives with commands such as `next`, `more`, `todo` and `discuss`. The split is not cosmetic. The autonomous pipeline is backend-pluggable for Claude Code and Codex only. The legacy mode talks natively to OpenAI, Anthropic, Google Gemini, DeepSeek, xAI, Qwen, Moonshot and local Ollama through their official SDKs. If your provider is not Claude or Codex, the legacy path is the only one available to you.

That architectural choice has a cost. The pyproject.toml carries a comment calling the root `unified_agent` directory an "obsolete root unified_agent compatibility copy" with transitional dependencies, and says the maintained framework resolves unified-agent from UnifedAgentWrapper instead. A repository that documents its own directory as obsolete is a repository mid-refactor.

## Installing PentestGPT and running a first target

The README lists three prerequisites: Python 3.12 or newer, the uv package manager, and an authenticated Claude Code CLI or Codex CLI. The Docker route bundles both CLIs, so it is the shorter path if you do not already have either installed locally. The local route is three commands, and `make install` is documented as running `uv sync`.

```bash
git clone https://github.com/GreyDGL/PentestGPT.git
cd PentestGPT
make install
```

After that, the README's first usage example runs the agent against a target in CTF mode, which it says is the default. Expect the agent to start a session, work the recon stage, and stream activity as it goes.

```bash
pentestgpt --target 10.10.11.234
```

The pentest mode swaps the stage list for asset discovery, vulnerability identification and a report. The README also shows passing free-text context, which is how you steer the model toward a specific attack surface.

```bash
pentestgpt --target 10.10.11.50 --instruction "WordPress site, focus on plugin vulnerabilities"
pentestgpt --target 10.10.11.234 --mode pentest
pentestgpt --list-sessions
```

For the container route, the Makefile drives everything. `make docker-build` builds the image, and `make docker-login` is documented as a one-time idempotent step that checks existing logins and only logs in what is missing. The README notes that Claude's token is stored in a volume while Codex uses its own in-container `codex login` with an OAuth callback forwarded through socat, and explains that the Codex token is not seeded because ChatGPT refresh tokens are single-use. `make docker-auth-status` reports whether both are logged in, and adding ROUNDTRIP=1 performs a live one-token check.

```bash
make docker-build
make docker-login
make docker-auth-status
make docker-run TARGET=http://127.0.0.1:8000 BACKEND=codex MODEL=gpt-5.5 MODE=ctf
```

Login state survives container recreation. The README states that `make docker-down` keeps the volumes and `make docker-nuke` removes them to force a fresh login or rotate a token. The compose file mounts `./workspace` into `/workspace`, adds NET_ADMIN capability and `/dev/net/tun` for OpenVPN, and sets CPU and memory limits of 4 cores and 8 GB by default. Those limits are adjustable, and the file comments say more CPUs and memory improve performance on complex tasks.

## Provider keys, model selection and the legacy smoke test

The legacy mode reads API keys from your environment or a `.env` file copied from `.env.example`. The example file lists OPENAI_API_KEY, ANTHROPIC_API_KEY, GEMINI_API_KEY (or GOOGLE_API_KEY), DEEPSEEK_API_KEY, GROK_API_KEY, QWEN_API_KEY and KIMI_API_KEY, plus commented base-URL overrides for Ollama, OpenAI and DeepSeek. Only providers you configure are enabled.

Model selection is per session. The README shows a reasoning model and a parsing model being chosen independently, and shows a local Ollama model addressed with the `ollama:` prefix and an OpenAI-compatible base URL. Running the command with no arguments auto-picks the best available models.

```bash
pentestgpt-legacy --reasoning-model claude-opus-4-8 --parsing-model gemini-3.5-flash
pentestgpt-legacy --reasoning-model ollama:qwen3 --base-url http://localhost:11434/v1
pentestgpt-legacy --list-models
```

Two diagnostic commands deserve attention. `--list-models` prints every supported model and marks which providers are configured. The `.env.example` comments describe `--smoke-test` as verifying that each model actually responds, and the README's truncated tail mentions a live round-trip that prints a pass/fail matrix. Note the naming inconsistency: the README's command table and the environment file do not use the same flag name for that check, so confirm which one your installed version accepts before scripting around it. The autonomous framework, by contrast, reads no API key from `.env` at all. It uses the authenticated CLI through unified-agent, and provider, model and effort are selected with CLI flags.

## Where PentestGPT is the wrong tool

The most obvious limitation is that the agent runs real offensive tools against a real address. There is no dry-run mode documented in the README, and no scope file, allowlist or confirmation prompt described anywhere in the repository documentation. If you point it at the wrong target, the tools in that Dockerfile execute. Everything about authorization, rules of engagement and target selection is your responsibility, and the README does not discuss any of it.

A second constraint is platform. The prerequisites are Python 3.12, uv and a CLI. There is no Windows installer, no Android package, no hosted service described in the README, and the homepage field in the repository metadata is empty even though the README points to pentestgpt.com. Anyone searching for a download or an app will not find one here.

The third is maturity signalling. The pyproject.toml classifies the project as "Development Status :: 4 - Beta" at version 1.0.0, and the repository root contains two migration and production-hardening reports alongside a `unified_agent` directory that the build file itself calls obsolete. That is not disqualifying, but it means the internal interfaces are still moving, and code that imports from those modules is likely to break on upgrade. Finally, the staged pipeline is fixed. If your methodology requires a specific, repeatable scan profile that produces identical output every run, an LLM choosing tools will not give you that, and a scripted scanner will.

## How this differs from a scripted scanner or a general coding agent

The closest comparison is a scripted framework such as Metasploit's resource scripts or a custom nmap and gobuster pipeline. Those tools do exactly what you wrote, in the order you wrote it, and produce the same output twice. PentestGPT inverts that: the model reads tool output and decides what comes next, which is what lets it follow an unexpected lead but also what makes its path non-reproducible. The README's Pentesting Task Tree in legacy mode is the middle ground, since it keeps a structured record of open tasks while the operator stays in control.

The second comparison is a general-purpose coding agent. PentestGPT's autonomous pipeline is built on top of Claude Code and Codex, so at the substrate level it is the same class of tool. What the project adds is the domain layer: the fixed stage lists for CTF and pentest work, session persistence, the Pentesting Task Tree, and a container image preloaded with nmap, gobuster, dirb and OpenVPN support. If you already have Claude Code installed and are comfortable prompting it yourself, the value you get from PentestGPT is that scaffolding rather than the model access, and that is worth weighing honestly.

## Licence, maintenance and upgrade cost

PentestGPT is MIT licensed, both in the repository metadata and in the pyproject.toml licence field, with an OSI Approved :: MIT License classifier. MIT is permissive: it allows commercial use, modification and redistribution provided the copyright notice and permission notice are retained. That is a statement about the licence text, not advice about your situation. Two things sit outside the licence and are worth checking yourself. The Dockerfile installs third-party tools including nmap, gobuster, dirb and openvpn, each under its own licence. And the autonomous pipeline drives Claude Code and Codex, so your use is also governed by those providers' terms, not by PentestGPT's MIT grant.

On maintenance, the last push to the default branch was on 2026-07-14, and the most recent release is v1.0.0 from 2025-12-24. The repository is not archived. The gap between v0.14.0 in May 2024 and v1.0.0 in December 2025 is roughly nineteen months, which tells you the project moves in large steps rather than continuous small ones. Plan upgrades around releases, not around the branch.

The upgrade cost is concentrated in the internal layout. Because the build file names `unified_agent` as an obsolete compatibility copy and points to UnifedAgentWrapper as the maintained path, any local extension that imports from the compatibility directory carries migration work. The two reports at the repository root, one a migration report and one a production hardening report, are the place to look before you write code against those modules.

## Conclusion

Adopt PentestGPT if you already run Claude Code or Codex locally and want a staged pipeline for CTF boxes or internal asset discovery, and if you accept that the agent executes real tools on a real target. Do not adopt it if you need a packaged Windows or Android client, or if your engagement requires a fixed, auditable scan profile rather than model-directed tool selection. Verify first that Python 3.12 and uv are present, that `claude` or `codex` is authenticated, and that you can run `make docker-auth-status` with ROUNDTRIP=1 before pointing it at anything.

## FAQ

### What is PentestGPT used for?

It is an agentic penetration testing framework that uses large language models to work through a target. The README describes two modes: a CTF pipeline of recon, exploit and walkthrough, and a pentest pipeline of asset discovery, vulnerability identification and report.

### Is PentestGPT safe to use?

The repository does not document a dry-run mode, a scope file or a confirmation prompt, and the Dockerfile installs real tools such as nmap and gobuster. Whether a run is safe depends entirely on the target you point it at and on your authorization to test that target.

### Is PentestGPT free?

The source is MIT licensed and there is no pricing in the repository. The autonomous pipeline requires an authenticated Claude Code or Codex CLI, and legacy mode requires provider API keys, so any cost comes from those providers rather than from PentestGPT itself.

### How do I install PentestGPT?

The README requires Python 3.12 or newer, uv, and an authenticated Claude Code or Codex CLI. The documented steps are to clone the repository, change into it, and run `make install`, which the README says runs `uv sync`. A Docker path is also documented through `make docker-build` and `make docker-login`.

### How do I use PentestGPT?

Run `pentestgpt --target <address>` for the default CTF mode, or add `--mode pentest` for the asset discovery, vulnerability identification and report pipeline. The README shows `--instruction` for passing context such as a WordPress site, and `--list-sessions` for resuming saved sessions.

### What is PentestGPT?

It is an autonomous penetration testing agentic framework powered by large language models, maintained by GreyDGL and published at USENIX Security 2024. It ships an autonomous multi-stage pipeline plus a modernized legacy interactive mode.

## Sources

- [GreyDGL/PentestGPT on GitHub](https://github.com/GreyDGL/PentestGPT)
- [Issues](https://github.com/GreyDGL/PentestGPT/issues)
- [License: MIT](https://github.com/GreyDGL/PentestGPT/blob/main/LICENSE)
- [README](https://github.com/GreyDGL/PentestGPT/blob/main/README.md)
- [Releases](https://github.com/GreyDGL/PentestGPT/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/greydgl-pentestgpt
