# PentestCode: an autonomous pentest agent, and the authorization it assumes

> A hard fork of OpenCode rebuilt for offensive security, where thirteen agents share one engagement state that survives the session. Useful on engagements and CTFs you are authorized to run, and pointed at anything else it is an unattended attack tool.

**s0ld13rr/pentestcode** — PentestCode - Multi-agent AI penetration testing system with persistent engagement state, strategic coordination, and parallel autonomous operations.

- Repository: https://github.com/s0ld13rr/pentestcode
- Stars: 784 · Forks: 119
- Language: TypeScript
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/s0ld13rr-pentestcode

## An OpenCode fork with the editing half removed

PentestCode starts from OpenCode and cuts in a different direction. It is a hard fork of the MIT licensed OpenCode project, with the code editing focus stripped out and the rest rebuilt for offensive security work. The terminal stays the interface: you point it at a target you are authorized to test, and the agents behind it run the tools, read the output, and keep their own picture of what they have found.

The root `package.json` shows how much of the upstream skeleton survived. It is named `pentestcode`, marked private, and declares ES modules with `bun@1.3.14` as the package manager. Workspaces cover `packages/*` plus `packages/sdk/js`. The dev script runs the OpenCode package in the browser conditions, `lint` runs oxlint, `typecheck` goes through turbo, `postinstall` patches node-pty in `packages/core`, and `prepare` installs husky hooks. One detail tells you how the project treats its own safety net: the root `test` script prints that tests must not be run from root and exits 1.

The repository layout is wider than a single binary. Alongside `packages/`, `install.sh` and `patches/` sit `skills/`, `specs/`, `bench/`, `script/`, a `.pentestcode/` directory, and agent context files at `AGENTS.md`, `CLAUDE.md` and `CONTEXT.md`. There is also a `.gitleaksignore` next to `.gitignore`, which tells you the project expects secret material to pass through its working tree.

## Four ways to install, and what each one costs you

The published artifact is a single self-contained binary, so the runtime story is short: no Bun, no Node and nothing else to install once it lands. Linux and macOS are covered on x64 and arm64.

```bash
npm install -g pentestcode-ai

curl -fsSL https://raw.githubusercontent.com/s0ld13rr/pentestcode/main/install.sh | bash
```

The npm route and the piped shell script both end up with that binary. Arch Linux users have a third path, `paru -S pentestcode-bin`, through the AUR. Two environment variables change what the installer does: `PENTESTCODE_VERSION=0.1.7` pins a release, and `PENTESTCODE_INSTALL=/usr/local/bin` chooses the directory. Note that the pinned version in the documentation is 0.1.7, which sits behind the v0.2.6 release, so anyone following the pin example literally installs an older build than the current one.

Building from source is the expensive route and the only one that needs the toolchain:

```bash
bun install
bun run build --single --skip-embed-web-ui
# binary at packages/opencode/dist/pentestcode-<os>-<arch>/bin/pentestcode
```

The output path spells out the naming scheme, `pentestcode-<os>-<arch>`, and the `--skip-embed-web-ui` flag is there because the single binary does not carry the web interface.

## Provider login comes before the first target

PentestCode has no model of its own. It talks to whatever provider you connect through the ai-sdk layer, and there are more than twenty of them: Anthropic, OpenAI, Google, Azure, AWS Bedrock and Ollama are named, with the rest covered by the same interface. Connecting one is the first command in the quick start:

```bash
pentestcode auth login          # connect your LLM provider
pentestcode                     # interactive session
```

Omitting the second command's argument gives you an interactive session in the terminal; passing `--prompt` with a quoted instruction makes it one shot. Because Ollama is in the list, a local model is a legitimate option, which matters for engagements where target data must not leave the machine. That choice also sets the ceiling on everything else: every decision the coordinator makes, every parser call and every report line is bounded by what the connected model is willing and able to do.

The project marks itself beta and says so in the opening line rather than in a changelog. The stated expectation is rough edges on real engagements and CTFs, with an open issue tracker as the feedback path. The release history is quick and regular for a beta: v0.2.4 on 2026-08-02, v0.2.5 on 2026-08-10, v0.2.6 on 2026-09-02, which is also the date of the last push to the repository.

## Thirteen agents under one coordinator

The architecture claim is that a prompt pasted into a chat window is not the same thing as a team. PentestCode follows the strategist-coordinator model described in the HPTSA research paper the project links, citing a 4.3 times improvement over a single agent. The lead agent is named `pentest`. It plans, dispatches work and tracks state; the specialists below it each carry their own system prompt, their own tool permissions and their own domain knowledge, and they run in parallel rather than taking turns.

There are 13 agents in total. The named ones are recon, scanner, enumerator, exploiter, identity for Active Directory and Kerberos, infrastructure covering SNMP, IPMI and databases, webapp covering the OWASP Top 10, post-exploit, exploit-dev, a critic that checks for false positives, and a reporter. Two more are described only as hidden agents, one for context compression and one for session management. Reading that list as an inventory rather than a recipe: it tells you which specialties the tool claims to cover and, just as usefully, that the coordinator is expected to hand work sideways rather than push one linear script forward.

The permissions point matters for anyone running this in anger. Specialists do not share one tool surface, which is the mechanism behind the false-positive checker and the reporter existing as separate agents rather than as modes of a single prompt.

## One engagement state that every agent reads and writes

The second structural claim is a shared memory. When the scanner turns up something, the enumerator sees it immediately, not because they talk to each other but because every agent reads from and writes to one structured engagement state that survives the session. Close the terminal, come back the next day, and the run resumes where it stopped.

The state has eight named parts. Hosts and services hold IP, hostname, OS, ports, service versions and banners. Vulnerabilities carry severity, a status of suspected, confirmed or exploited, an evidence chain and a confidence score. Credentials record username, hash or password, type, domain and what they unlock. Access tracks who holds a shell, RDP or database session on which host and at what privilege level. Relationships form an entity graph with edge types such as EXPLOITED_VIA, CREDENTIAL_FROM, ADMIN_OF and PIVOT_TO. The AD domain model covers domain controllers, trusts, admins, password policy and GPOs. Network segments list VLANs, reachable networks and pivot hosts. Attack paths are computed through the relationship graph with cost-based Dijkstra plus Yen's K shortest routes.

Two things sit beside that structure. Mid-run queries are available through `/status`, `/vulns` and `/creds`. And a human-readable `findings.md` logs every vulnerability, credential and access gain with timestamps, so `tail -f findings.md` is enough to watch an engagement from another window. The evidence chain and confidence score on each finding are the parts a reviewer should look at first, since they are what separates a confirmed result from a guess the critic agent has not yet filtered.

## Parser tools are mandatory, so results cannot slip past the state

Beyond bash there are 18 built-in pentest tools, and a subset of them is compulsory. After a scan the agent is required to pipe the output through a parser rather than grep the XML by hand, so that every finding lands in the engagement state instead of dying in scrollback.

The parsers cover the common toolchain: `nmap_parse` turns nmap XML into hosts and services, `nuclei_parse` turns Nuclei JSON into vulns with severity, `cme_parse` turns NetExec output into creds, access and hosts, `gobuster_parse` classifies directory brute output, `bloodhound_parse` populates the AD model from SharpHound JSON, and `sqlmap_parse` extracts injection points. Two more analyze rather than import: `xss_detect` looks at responses for reflected or stored XSS, and `jwt_analyze` decodes a token and checks for alg none, a weak HMAC or an expired validity.

The remaining tools are planning and bookkeeping. `cred_spray` plans a credential spray across every discovered service, `scope_check` validates CIDR and wildcard scope, `attack_path_suggest` runs the cost-based path finding, `tunnel_manage` plans SSH, chisel and ligolo tunnels and tracks live sessions, `phase_control` manages phases with quality gates, `report_gen` produces markdown or JSON reports, and `state_update` records findings with more than 30 mutation types in batch mode. That is the shape of the tool surface: import results, analyze them, then record and report on them.

## The guardrails are a scope checker and a beta label

Nothing in the design prevents a run from leaving the network you meant to test. What exists is `scope_check` for CIDR and wildcard validation, `phase_control` for phases with quality gates, and the critic agent that filters false positives before the reporter writes anything up. Those are operator aids inside the engagement, not a permission system, and no gate in the repository substitutes for a signed scope.

The public shape of the project: MIT licensed, TypeScript, 726 stars, 116 forks and 6 open issues at the time of this snapshot, default branch main, last push 2026-09-02. Open issues on a beta offensive tool tend to be the interesting ones, because the failure modes of an autonomous pentest agent are not the ones a linter catches.

The honest framing is that this is a beta that its author says holds up on real engagements and CTFs while expecting rough edges. Used inside a scope you control, the shared state, the mandatory parsers and the report generator are the parts that repay the setup. Used outside one, the same persistence and credential reuse that make it competent is what turns it into an unattended intrusion tool.

## Conclusion

PentestCode belongs on a laptop with written authorization in hand: a scoped engagement, a CTF, or a lab you own, where scope_check and phase_control are wired to a real scope document. Do not point it at a network you have no permission to test, since credential reuse and session state are exactly what makes it effective and exactly what makes an unauthorized run a crime. Before the first run, confirm the provider you logged into, read the engagement state on disk, and check that the 0.1.7 pin in the install docs is not older than the binary you actually installed.

## FAQ

### Is it legal to run PentestCode against a network?

Only where you hold authorization, such as a scoped engagement, a CTF or a lab you own. The tool ships a scope_check tool for CIDR and wildcard validation, but that is an operator aid rather than a permission system.

### How many agents does PentestCode run at once?

Thirteen. A lead agent named pentest plans and dispatches, and specialists such as recon, scanner, enumerator, identity, webapp, a critic and a reporter each run with their own system prompt and tool permissions.

### Does PentestCode need Node or Bun installed to run?

No. The install produces a single self-contained binary for Linux and macOS on x64 and arm64, with no runtime required. Bun is only needed to build from source with bun install and bun run build --single.

### Which LLM providers can PentestCode use?

More than 20 through the ai-sdk layer, including Anthropic, OpenAI, Google, Azure, AWS Bedrock and Ollama. You connect one with pentestcode auth login before starting a session.

### Can I close a PentestCode session and come back to it later?

Yes. The engagement state survives the session, so the agents resume where they stopped, and a findings.md file logs vulnerabilities, credentials and access gains with timestamps you can follow.

## Sources

- [Issues](https://github.com/s0ld13rr/pentestcode/issues)
- [License: MIT](https://github.com/s0ld13rr/pentestcode/blob/main/LICENSE)
- [README](https://github.com/s0ld13rr/pentestcode/blob/main/README.md)
- [Releases](https://github.com/s0ld13rr/pentestcode/releases)
- [s0ld13rr/pentestcode on GitHub](https://github.com/s0ld13rr/pentestcode)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/s0ld13rr-pentestcode
