PentesterFlow: an approval-gated pentest agent for the terminal
Agentic offensive-security in your terminal
At a glance
- What is it?
- PentesterFlow wraps a local or hosted LLM in a terminal agent that plans against a scoped target, runs real shell and HTTP tools after approval, and writes evidence-backed findings to Markdown. It is for authorized pentests and bug bounty work, not for unattended scanning.
- Who is it for?
- Adopt PentesterFlow if you already run authorized engagements and want an agent that keeps shell and HTTP actions behind an approval prompt while writing findings you can hand to a client. Skip it if you need unattended scanning or a hosted service with a web console, since this is a terminal binary tied to a model endpoint you configure yourself.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 32 days ago.
- What is it written in?
- Mainly TypeScript, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 2, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What PentesterFlow is for, and who it is not for
PentesterFlow is a terminal assistant for authorized offensive-security work. The README frames it as human-in-the-loop, which is a design constraint rather than a slogan: the analyst sets scope, approves sensitive actions, and decides what counts as a finding. It is aimed at penetration testers and bug bounty hunters who already know how to test an API by hand and want the repetitive parts (recon, enumeration, evidence collection, write-up) handled by a model that can call real tools.
The scope is explicit. The README carries a warning that the agent can run shell commands, make HTTP requests, edit files, and process captured traffic after approval, and that it should only be used on systems where you have explicit authorization. That warning is the whole product boundary. If you want an unattended scanner that crawls a host and emails a report, this is the wrong shape of tool, because the approval step is not optional decoration, it is the interaction model.
The skills directory in the repository layout is a clue to the intended workflow. The README lists built-in pentest skills for recon, web vulnerabilities, SSRF, SSTI, JWT, GraphQL, race conditions, subdomain takeover, Supabase, and deserialization. Those are playbooks, not scanners. The agent reads a playbook and then drives tools, which means the quality of a session depends heavily on the model you point at it.
The agent loop, skills, and where findings actually go
The documented loop is plan, act, observe, verify, report, and learn, with auto-continue and context compaction between steps. Tools exposed to the model include shell, HTTP, file tools, search, browser capture, Burp ingest, MCP, jobs, and findings. The README describes execution as curl-first and reproducible, meaning the agent is expected to emit commands you could paste into a terminal yourself rather than opaque internal requests. That is a meaningful choice for auditability: a reviewer can read the command log and rerun a request.
Skills are Markdown playbooks, and the README notes an optional fork so that large playbooks stay out of the parent context. The quickstart transcript shows what that looks like in practice: a webvuln skill forks with four tools, the agent issues an HTTP GET against an orders endpoint, then falls back to a shell curl with a second user's bearer token, and on confirming cross-account access writes a high-severity IDOR finding to ./findings/idor-orders.md. The finding file is the unit of output. Reports are Markdown with evidence, impact, proof of concept, and remediation, and logs are JSON-lines.
The README is direct about hallucination: confirm_finding should be used only after reproduction with request and response evidence. That is a convention the model is instructed to follow, not a runtime guarantee. Nothing in the README describes a verifier that rejects a finding lacking evidence, so the discipline sits with the analyst reading the transcript.
Installing PentesterFlow and running a first session
The installers download the latest standalone binary for your OS and verify the published SHA-256 checksum when available. On macOS and Linux the documented command is a shell script piped to sh:
curl -fsSL https://raw.githubusercontent.com/PentesterFlow/agent/main/install.sh | shOn Windows, the README gives a PowerShell equivalent:
irm https://raw.githubusercontent.com/PentesterFlow/agent/main/install.ps1 | iexYou can pin a release or choose an install directory with environment variables. The README uses PENTESTERFLOW_VERSION and PENTESTERFLOW_INSTALL_DIR in this example:
PENTESTERFLOW_VERSION=v0.1.6 PENTESTERFLOW_INSTALL_DIR="$HOME/.local/bin" \
sh -c "$(curl -fsSL https://raw.githubusercontent.com/PentesterFlow/agent/main/install.sh)"Standalone binaries are also published per platform on GitHub Releases: pentesterflow-darwin-arm64 and pentesterflow-darwin-x64 for macOS, pentesterflow-linux-arm64 and pentesterflow-linux-x64 for Linux, and pentesterflow-windows-x64.exe for Windows. The README states the x64 standalone binaries are built with Bun's baseline runtime for older x86_64 CPUs and do not require AVX2. If you build from source instead, package.json sets engines.node to >=20.0.0 and packageManager to [email protected], with bin entries for pentesterflow and pentesterflow-browser-mcp.
For a first run, the quickstart pairs a local model with the CLI. Pull a model, then start the agent:
ollama pull qwen2.5-coder:32b
pentesterflowInside the CLI, the documented sequence is to pick a provider, set a target, and describe the task in plain language:
/provider
/target https://app.example.com
map the authenticated API surface and test for IDORTo continue an earlier assessment, the README gives a resume flag that also replays the previous session's persistent memory:
pentesterflow --resume <session-id>If you prefer flags over the interactive setup, backends can be selected at launch. The README shows an Ollama example and an OpenAI-compatible endpoint with a base URL and key:
pentesterflow --backend ollama --model qwen2.5-coder:32b
pentesterflow --backend openai-compat --base-url https://api.example.com/v1 --api-key sk-...Hosted providers follow the same pattern with an environment variable for the key, for example MOONSHOT_API_KEY with --backend kimi, GROQ_API_KEY with --backend groq, GEMINI_API_KEY with --backend gemini, and ANTHROPIC_API_KEY with --backend anthropic. What you should see after /target is a confirmation line echoing the target URL, and after a task, a running transcript of tool calls with their results.
Permission tiers are the real safety control
The README lists three permission tiers: ask, auto-safe, and yolo, alongside allow-once and allow-session grants and a plan mode. This is the mechanism that separates PentesterFlow from an autonomous scanner, and it is also the part most likely to be misconfigured. Ask means every sensitive action pauses for a modal. Auto-safe presumably lets a class of actions through without prompting, though the README does not enumerate which actions fall on which side of that line. Yolo is named and offered; the README does not describe its bounds either.
That gap matters. If you cannot state from the documentation exactly which commands auto-safe will execute without asking, treating it as a convenience setting on a client engagement is a risk you are taking on faith. The safer default for anything with a signed scope is ask, with allow-once grants for the specific commands you have already reviewed in the transcript.
Plan mode is the other half of the control story. It lets the agent propose a sequence before executing, which is useful when you want to sanity-check whether the model understood the target before it starts issuing requests. Neither plan mode nor the permission tiers replace scoping. The agent works against whatever /target you set, and the README's authorization warning is the only boundary it draws.
Memory, sessions, and what persistence costs you
Long engagements are the stated motivation for the memory system. The README describes session memory, curated facts added with a # prefix, context snapshots, a resume recap, and a local intelligence layer built from project and personal knowledge bases. On resume, the CLI shows a recap of persistent memory so you do not rebuild context by hand, and the repository layout includes a skills directory alongside src, which is where the Markdown playbooks live.
The trade-off is data retention. A personal knowledge base that improves future sessions is, by definition, storing details from past engagements on disk. The README does not describe an encryption scheme, a retention policy, or a purge command for that store. If you work under contracts that specify how client data is handled, that is a question to answer before you let the agent learn from a live target rather than after. The same applies to captured traffic: browser capture and Burp ingest both pull real request and response data into the session.
Compaction and context snapshots exist because long sessions exceed a model's window. Compaction is a lossy operation by nature, and the README does not document what gets dropped or whether a snapshot can be replayed to recover detail. For a short bug bounty session this is a non-issue. For a multi-day engagement, verify that the resume recap actually carries the facts you care about before relying on it.
How PentesterFlow differs from a general coding agent
The obvious comparison is a general-purpose terminal agent such as Claude Code or an OpenAI-compatible coding assistant pointed at a security task. Those tools are good at reading a repository, editing files, and running commands, and they can be told to test an API. The difference is what ships in the box. A coding agent has no built-in playbooks for SSRF, SSTI, JWT, GraphQL, race conditions, or subdomain takeover, no finding schema, and no Burp ingest. With PentesterFlow those are part of the repository rather than prompts you write yourself.
The second difference is the approval model. Coding agents typically ask before file writes or destructive shell commands, but the category of action they guard is different. PentesterFlow guards HTTP requests against a live target, which is the action that can get you in trouble on an engagement. The permission tiers, allow-once grants, and plan mode are built around that specific risk.
The third difference runs the other way. A general coding agent backed by a frontier hosted model will often reason better about an unfamiliar API than a local 14B or 32B model will. PentesterFlow supports both, with backends listed for Ollama, LM Studio, Kimi, Groq, Gemini, Anthropic, OpenAI-compatible endpoints, and OpenRouter, but the skills and the finding format do not compensate for a weak model. If your only available model is small, expect more false leads and more time spent reading transcripts.
Maintenance, licence, and upgrade cost
The repository is not archived, and the last push was on 2026-08-31. That is recent enough that the project is being worked on, but the release history is worth reading carefully: the most recent listed release is v0.1.20 from 2026-06-14, while package.json carries version 0.3.0 and the README transcript shows PF v0.3.0. The 0.3.x line does not appear in the published release list, so if you install via the shell script you are pulling whatever the latest release is, and if you build from source you get 0.3.0. Those are not the same artifact.
Version drift between the README and the releases is a real upgrade cost. The README documents PENTESTERFLOW_VERSION for pinning, which is the right lever: pin to a known release rather than tracking latest, and read CHANGELOG.md before moving. The repository also ships AUDIT.md and PROJECT.md, and the CI script in package.json runs typecheck, lint, unit tests, OpenTUI tests, and a build, so a source build has a defined verification path.
On licensing, the project is Apache-2.0 and the README badge and package.json agree. Apache-2.0 includes an explicit patent grant and requires that you preserve notices and state changes when you redistribute. It does not tell you whether your use of the tool on a given target is lawful; that question sits with your authorization and your contract, not with the licence file. Bundling the binary into a commercial service is a different analysis from running it internally, and the licence text is the thing to read rather than a summary of it.
Editorial conclusion
Adopt PentesterFlow if you already run authorized engagements and want an agent that keeps shell and HTTP actions behind an approval prompt while writing findings you can hand to a client. Skip it if you need unattended scanning or a hosted service with a web console, since this is a terminal binary tied to a model endpoint you configure yourself. Before trusting it on a live scope, verify three things: that your chosen backend answers through /provider, that your permission tier is set to ask rather than auto-safe or yolo, and that a confirmed finding actually lands in ./findings/<slug>.md with the request and response you reproduced.
Frequently asked questions
How do I install PentesterFlow?
On macOS and Linux the README gives a shell installer, curl -fsSL https://raw.githubusercontent.com/PentesterFlow/agent/main/install.sh | sh, and on Windows a PowerShell equivalent, irm https://raw.githubusercontent.com/PentesterFlow/agent/main/install.ps1 | iex. The installers download the latest standalone binary and verify the published SHA-256 checksum when available. You can pin a release with PENTESTERFLOW_VERSION and choose a directory with PENTESTERFLOW_INSTALL_DIR.
Which model backends can PentesterFlow use?
The README lists Ollama, LM Studio, Kimi, Groq, Gemini, Anthropic, OpenRouter, DeepSeek, Naraya, and generic OpenAI-compatible endpoints. You select one interactively with /provider or at launch with flags such as --backend ollama --model qwen2.5-coder:32b, and hosted providers read their key from an environment variable like ANTHROPIC_API_KEY or GROQ_API_KEY.
Does PentesterFlow run commands without asking?
It depends on the permission tier you choose. The README documents ask, auto-safe, and yolo, plus allow-once and allow-session grants and a plan mode. The README does not enumerate which actions auto-safe or yolo permit without a prompt, so the documented behaviour for a reviewed engagement is the ask tier with allow-once grants for commands you have already seen in the transcript.
Where does PentesterFlow write its findings?
Confirmed findings are written to ./findings/<slug>.md as Markdown containing evidence, impact, a proof of concept, and remediation, and logs are JSON-lines. The README states that confirm_finding should only be used after reproduction with request and response evidence.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/pentesterflow-agent)