Model or dataset
0xSteph/pentest-ai-agents avatar
0xSteph/pentest-ai-agents

pentest-ai-agents: 50 Claude Code subagents for authorized penetration testing

Turn Claude Code into your offensive security research assistant. Specialized AI subagents for authorized penetration testing plan engagements, analyze recon, research exploits, build detections, audit STIGs, and write reports.

2,298 stars433 forksShellMIT

At a glance

What is it?
A file-based pack of Claude Code subagents that route offensive security tasks to specialists. No runtime, no server, just Markdown definitions copied into your Claude Code setup.
Who is it for?
Adopt it if you already run Claude Code for authorized engagements and want structured routing from recon to reporting without standing up a service; skip it if you expected an autonomous scanner, since the repository states it packages Markdown definitions and ships no tooling.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 45 days ago.
What is it written in?
Mainly Shell, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The problem pentest-ai-agents addresses

A general-purpose coding assistant asked about Active Directory enumeration tends to answer at the level of a blog post. The knowledge is there, but it is not organized around the sequence a tester actually follows: scope, recon, exploitation, post-exploitation, detection, report. pentest-ai-agents is an attempt to fix that by splitting the domain into named specialists and letting Claude Code route to one.

The README describes the project as "a collection of 50 Claude Code subagents that turn Claude into an offensive security research assistant." Each agent carries domain knowledge for one area, and the listed set runs from `recon-advisor` and `web-hunter` through `ad-attacker`, `cloud-security`, `container-breakout`, `c2-operator`, `detection-engineer` and `report-generator`. The intended audience is a working tester or red teamer who already has Claude Code open and wants the assistant to stop being generic.

The framing matters. This is not a scanner and not a service. The README states plainly that there are "no servers, no Python deps, no setup beyond copying files." That single sentence defines both the appeal and the ceiling of the project.

How routing and the Tier 1 / Tier 2 split work

The mechanism is file-based. Agent definitions live under `agents/`, slash commands under `commands/`, and Claude Code reads them from the installed location. The README's flow is: install the agent files, open Claude Code, describe your task, and Claude routes to the right specialist automatically. The repository also ships an agent map as a Mermaid flowchart in the README, showing edges such as `osint-collector` to `recon-advisor` to `vuln-scanner`, and `vuln-scanner` fanning out to the exploitation agents.

The design decision worth noting is the tier system. Tier 1 agents are advisory and routable from any task. Tier 2 agents are execution-capable, require a declared scope, and sit in the offensive operations cluster. The README says the CI validator "requires the scope-guard block on every Bash-capable (Tier 2) agent." The v3.3 notes record that `cicd-redteam` was Bash-capable but missing that block, now fixed, and that the new CI check would have caught it.

That is a real architectural constraint rather than a marketing line: an agent that can run commands cannot be installed without the scope-enforcement text. The v3.2 notes add that `_scope-guard.md` carries an explicit hard-refusal list covering DoS, mass scanning, unattended worms, false-flag operations and safety-of-life systems. Whether those refusals hold in practice depends on the model, not on the file, and the README does not claim otherwise.

Installing as a plugin or with install.sh

Two install paths exist. The v3.3 notes describe the plugin path as two lines inside Claude Code, and the README says the `install.sh` curl path "still works unchanged." The plugin route is the shorter one:

bash
/plugin marketplace add 0xSteph/pentest-ai-agents
/plugin install pentest-ai-agents@pentest-ai-agents

The v3.3 notes also mention `./install.sh --global` as the global install form, and v3.1 added `install.sh --tools` as an opt-in installer for the underlying CLI tools. The v3.3 installer fixes include `curl | bash` no longer crashing under `set -u`, a corrected one-liner clone URL, slash commands installing alongside the agents, and `--uninstall` removing everything cleanly.

For an air-gapped setup, the repository ships a Dockerfile described in its own comments as "a minimal offline bundle of the pentest-ai-agents definitions" and explicitly "NOT a tool runner." It pins `debian:bookworm-slim` by digest, creates a non-root user with uid 10001, and copies `agents/`, `commands/`, `.claude-plugin/`, `db/`, `examples/` and the top-level docs. The default command prints usage rather than starting anything:

bash
docker build -t pentest-ai-agents .
docker run --rm pentest-ai-agents

Expect the container to print the bundle banner and exit. There is no server, no port, and no tooling baked in.

Slash commands and the findings database

Two features do more work than their billing suggests. The first is the routing command. `/recommend "freeform task"` returns the right agent plus concrete commands, and `/agents-for <tag>` filters the catalog by domain. This is the answer to the obvious problem with a 50-agent catalog: nobody memorizes 50 names, and a routing command is cheaper than a lookup table in a wiki.

The second is the findings database. v3.2 introduced `vulns.tool_used` for filtering findings by the tool that produced them, with new indexes on `cve` and `tool_used`. Existing engagements migrate forward through `db/migrate.sh`. v3.1 added `db/doctor.sh`, which audits which underlying CLI tools are installed on your machine, grouped by agent, showing a check or cross per tool with install hints.

That last script is the honest part of the project. The agents reference real tools across a wide surface, and `db/doctor.sh` is how you find out which of them your box actually has. The README lists 80+ tools tracked. Nothing in the repository installs them for you by default.

Where pentest-ai-agents is the wrong tool

The repository is candid about its own boundary. The Dockerfile comment says the image "runs nothing offensive" and packages files. The README says there are no servers and no Python dependencies. So anyone evaluating this as an autonomous pentest platform is looking at the wrong project: there is no scheduler, no scanning engine, and no execution layer beyond what Claude Code itself can invoke on your host.

A second limitation is scope enforcement. The CI validator checks that a scope-guard block is present in every Tier 2 agent, but presence of a Markdown block is not the same as enforcement. The guard is text handed to a model. If your rules of engagement depend on a hard technical control rather than a prompt-level instruction, this project does not supply that control, and the README does not claim it does.

A third is coverage honesty. Agents exist for `scada-attacker`, `iot-pentester`, `llm-redteam` and `container-breakout`, but the README does not document validation of those agents against real targets. The examples directory holds five files: an engagement plan, an nmap analysis, a detection rule, a STIG finding and a report excerpt. That is a sample set, not an evidence base.

The legal section of the README exists for a reason. The repository is a set of instructions for offensive work, and using it without written authorization is the failure mode that matters most.

How this differs from running a pentest framework

The closest comparison is a conventional framework such as Metasploit or a scanning platform like Nuclei. Those execute. They hold exploit modules, payloads and templates, and they run them against a target with deterministic results. pentest-ai-agents holds none of that. It holds prose that shapes how a model reasons about a task, plus a findings database and a set of slash commands.

The practical difference shows up in reproducibility. Run a Nuclei template twice and you get the same request. Ask an agent twice and you get two answers that may differ in ordering and emphasis. The README's own agent map shows a chain from `vuln-scanner` through `poc-validator` to `exploit-chainer` and `attack-planner`, which is a reasoning pipeline, not an execution pipeline.

That makes the two approaches complementary rather than competing. A framework produces raw output; these agents are positioned to interpret it, plan the next step, and turn the result into detection content and a report. The `detection-engineer`, `forensics-analyst` and `stig-analyst` agents sit on the defensive side of the same map, which is unusual for a tool in this category. The `swarm-orchestrator` and `report-generator` close the loop.

Maintenance, licence and upgrade cost

The last push to the default branch was on 2026-08-16, and the repository is not archived. Release cadence visible in the notes is roughly quarterly: v3.2.0 on 2026-05-03 and v3.4.0 on 2026-08-05, with v3.3 documented in the README between them. Each release has carried structural changes rather than content tweaks. v3.2 added four agents and tightened the scope guard. v3.3 added fifteen agents, moved to a plugin install path, and hardened the CI validator.

That cadence has a cost. A jump from 35 to 50 agents changes routing surface, and the plugin marketplace path is new enough that anyone on the older `install.sh` route should read the v3.3 notes before upgrading. The findings database has its own migration path through `db/migrate.sh`, so schema changes are handled rather than dropped.

The licence is MIT. In practical terms that permits commercial use, modification and redistribution provided the copyright notice and permission notice are included. It also means no warranty and no liability, which for a tool whose output feeds security decisions is worth reading rather than assuming. This is a description of the licence text, not legal advice; check it against your own policy.

Editorial conclusion

Adopt it if you already run Claude Code for authorized engagements and want structured routing from recon to reporting without standing up a service; skip it if you expected an autonomous scanner, since the repository states it packages Markdown definitions and ships no tooling. Before trusting it on a real engagement, run db/doctor.sh to see which underlying CLI tools are actually present on your box, and read _scope-guard.md to confirm the refusal list matches the rules of engagement you work under.

Frequently asked questions

What are the top 3 AI agents?

This repository does not rank agents, so it cannot answer that. What it documents is its own catalog: 50 Claude Code subagents split into Tier 1 advisory agents and Tier 2 execution-capable agents that require a declared scope.

What are the 5 types of AI agents?

The repository does not define a general taxonomy of agent types. It groups its own agents by domain instead, covering recon, web, Active Directory, cloud, mobile, wireless, payload crafting, reverse engineering, detection engineering and forensics.

Is pentesting illegal?

The repository ships a DISCLAIMER.md and a legal section in the README, and the agents are framed around authorized penetration testing. Whether a given engagement is lawful depends on your authorization and jurisdiction, not on the tool.

What is AI pentesting?

In this project it means using Claude Code subagents that carry offensive security domain knowledge to plan engagements, analyze recon, research exploits, build detections and write reports. The README states the agents are advisory or execution-capable, and that Tier 2 agents require a declared scope.

How do I pentest AI agents?

The repository includes an `llm-redteam` agent added in v3.2, described in the notes as covering OWASP LLM Top 10 testing, prompt injection, RAG poisoning, MCP server abuse and agent tool abuse. The README does not document validation of that agent against real targets.

Official sources

  1. 0xSteph/pentest-ai-agents on GitHub
  2. License: MIT
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/0xsteph-pentest-ai-agents.svg)](https://hysenlabs.com/projects/0xsteph-pentest-ai-agents)