Open-source project
bugbasesecurity/pentest-copilot avatar
bugbasesecurity/pentest-copilot

Pentest Copilot: an agent that drives a Kali box from a browser UI

Pentest Copilot is an AI-powered browser based ethical hacking assistant tool designed to streamline pentesting workflows.

1,530 stars284 forksTypeScriptMIT

At a glance

What is it?
A read of Bugbase's Pentest Copilot: what the agent actually executes, how it wraps Kali, Burp and a browser into one session, and what the README says about running it outside Linux.
Who is it for?
Pentest Copilot is built for a Linux operator with Docker who wants an LLM driving an existing Kali toolchain rather than reasoning about output alone. The bundle it ships is substantial: 16 agent tools, a registry of over a hundred security packages, Burp Suite integration and real browser automation with a VNC view.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 54 days ago.
What is it written in?
Mainly TypeScript, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 7, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The agent runs commands, it does not suggest them

The README's one-line summary is that Pentest Copilot connects to a Kali attack box, runs tools autonomously, analyses results and iterates. That word autonomous is doing the work. The documented behaviour is an agentic loop: the model runs commands directly on the attack box, reads the output, decides the next step and repeats, with up to 25 iterations per turn and no manual nudging required.

That is the whole architectural claim, and it is a different thing from a chat assistant that suggests commands for you to paste. The tool list makes it concrete. There are 16 agent tools covering bash, Python script execution, tool installation, shell management, Google search, subagent spawning, Burp Suite functions (proxy history, Repeater, Intruder, Collaborator) and browser automation. Each of those is an execution path, not a suggestion surface.

Tool supply is handled by a registry described as over 100 curated capabilities across seven categories: network, rev, pwn, crypto, forensics, stego and core. You select what a session needs and the agent installs the rest, which is the difference between a container that breaks on a missing dependency and one that can fetch it mid-run.

The repository is TypeScript under the MIT licence, with the homepage at pentest.bugbase.ai. The tree is a conventional split: backend/, frontend/, a kali/ directory for the attack environment, deploy/, docs/, three compose files for different modes and run.sh at the root. There are no GitHub releases, so the version a user runs is whatever the repository's default branch holds. The last push was on 2026-08-14.

Three commands and four services to wait for

The quick start is short enough to state in full:

bash
git clone https://github.com/bugbasesecurity/pentest-copilot.git
cd pentest-copilot
./run.sh start

Then open http://localhost:3000, register and start a session. Configuration of the model happens after the first start, under Settings -> Models.

What `run.sh start` actually brings up is worth naming, because it is four services rather than one. The compose file defines MongoDB 7 with a persistent volume and a healthcheck running a ping command through mongosh, Redis 7 on alpine with a `redis-cli ping` healthcheck, and a backend service that depends on both being healthy before it starts. The backend gets the Mongo connection string and the Redis URL as environment variables, plus `CODEX_HOME` pointing at /root/.codex and `SSH_CONFIG_FILE` pointing at /root/.ssh/config inside the container.

The script is explicit about waiting. It waits for the frontend, backend, MongoDB and Redis to be ready before reporting success, and if startup fails it prints the affected container status and recent logs rather than a bare failure. That is a small design decision worth copying.

Port mapping is tight. MongoDB, Redis and the backend ports are all bound to 127.0.0.1 rather than 0.0.0.0, so a mistake in the configuration does not expose a database to the network. The backend also publishes 6080, which is the noVNC port for watching the agent's browser, and 9020.

Bringing your own model, including your existing subscriptions

The model side is the most interesting design decision. Rather than insisting on one provider, the tool accepts OpenAI, Anthropic by API key or OAuth, Google, Mistral, or any OpenAI-compatible endpoint. First-class model families named in the README include GPT-5.6 Sol, Terra and Luna, Claude Fable, Opus and Sonnet 5, and Kimi K3 either through the direct Moonshot API or OpenRouter.

The second option is more unusual. If you already pay for a coding agent subscription, the CLI can be used as the inference provider:

bash
codex login
claude auth login

After authenticating on the machine running the CLI, you select Use Codex or Use Claude Code in the model settings. The README is careful about the boundary here: the official CLI owns login, refresh and subscription entitlement handling, and Pentest Copilot neither copies nor replays OAuth tokens. Subscription transports receive the same conversation history and function schemas as the API providers and return the same assistant and tool call contract, so the tool loop and consent checks keep working.

Mechanically, the Docker backend includes the Linux Codex CLI and mounts only the host's file-based ~/.codex/auth.json. Set `CODEX_AUTH_FILE` before `docker compose up` if that file lives somewhere else. The README warns that the CLI may refresh the file during normal use and that it should never be committed or shared. Host Keychain-only credentials and Claude Code remain available only in developer or host mode until a host inference bridge is configured.

The terms caveat is stated too: Claude subscription use is local CLI control and must comply with Anthropic's current third-party product and subscription terms. Worth reading before pointing it at an engagement, and not a line a project would add casually.

Burp Suite, browser automation and VPN handling

The integrations are what separate this from a shell wrapper around an LLM.

Burp Suite support is real rather than nominal. There is a proxy history viewer, the ability to send requests to Repeater and Intruder, and Collaborator for out-of-band testing, all reachable by the agent and through the UI. That matters for workflow, because the agent can take a request it found during enumeration, push it into Repeater and mutate it, without a human copying payloads between tools.

Browser automation uses the Magnitude library, described as real browser automation rather than a headless request generator. The stated use is testing login flows, filling forms and interacting with JavaScript heavy applications, with traffic optionally proxied through Burp. In Docker mode you watch the browser through the built-in VNC stream; in developer mode it opens on your local desktop. That split is the whole reason for the noVNC port in the compose file.

VPN management handles `.ovpn` and `.conf` bundles with their referenced certificates, keys or credentials, connect and disconnect from the browser, and multiple simultaneous connections. The backend service requests the NET_ADMIN capability, and the compose file comments that it is required only when a workspace selects a Local work host and starts OpenVPN there, which is a good example of a capability grant scoped and explained rather than blanket.

Subagent parallelism lets background agents run tasks concurrently, with a directory brute force and a subdomain enumeration at the same time given as the example.

Consent, safety checks and what runs where

An agent that executes commands unattended needs a brake, and the README describes one: dangerous commands such as recursive deletes, device writes and fork bombs require explicit approval even in auto-run mode. So the consent loop is not bypassed by turning on automation, which is a design choice that differs from most agent tooling.

The work host model is the other half. In Docker mode the tool mounts the host's ~/.ssh and ~/keys directories read-only, and each workspace selects a concrete `Host` alias from ~/.ssh/config. Every session in that workspace uses the same host and work folder, and the README points out that this avoids copying private keys into MongoDB. Set `HOST_SSH_DIR` or `HOST_SSH_KEYS_DIR` before starting Docker if those directories live elsewhere.

Named aliases rather than wildcard entries are recommended, and the shape is ordinary SSH config:

ssh-config
Host lab-box
  HostName 10.10.10.10
  User root
  IdentityFile ~/.ssh/lab-box.pem

The compose file also passes HTTP_PROXY, HTTPS_PROXY and NO_PROXY through to the backend, with the default NO_PROXY list naming localhost, 127.0.0.1, mongodb, redis and kali. That default is a small correctness detail: without it, the internal service hostnames would try to resolve through whatever proxy the operator has set, and the container would fail in a way that looks like a network bug.

Windows users need WSL2, and the README says so

The platform guidance is short and negative, which is refreshing. On Windows, run Pentest Copilot inside WSL2 with Docker Desktop's WSL integration enabled. Native PowerShell and Windows SSH work hosts are not supported, and the stated reason is that workspace commands require a POSIX shell.

The advice that follows is a performance and permissions point rather than a preference: clone the repository into the WSL filesystem, not under /mnt/c. A repository on a mounted Windows drive loses file permission semantics and runs noticeably slower, and a pentesting tool that shells out constantly feels that difference immediately.

Other structural notes. The project has no GitHub releases, so tracking a specific version means tracking a commit. `config.toml.template` at the repository root is mounted read-only into the backend container, so the configuration path is explicit rather than implicit. There are separate compose files for the default, development and Kali configurations, which is what you would expect from a project where the attack environment itself is versioned differently from the UI.

On authorisation, the README names its intended use: real-world engagements, boot2root boxes and CTFs. An agent that will run up to 25 iterations against a target on its own makes that scope note worth taking seriously, since the tool has no notion of whether a host is in front of it.

Editorial conclusion

Pentest Copilot is built for a Linux operator with Docker who wants an LLM driving an existing Kali toolchain rather than reasoning about output alone. The bundle it ships is substantial: 16 agent tools, a registry of over a hundred security packages, Burp Suite integration and real browser automation with a VNC view. The constraints are equally clear, with WSL2 required on Windows, a POSIX shell required for workspace commands, and read-only mounts of the host SSH directory in Docker mode. Start with the three-command quick start, assign a model under Settings -> Models, and scope every session to hosts you are authorised to test, since the agent loop will run up to 25 iterations on its own.

Frequently asked questions

What does Pentest Copilot need to run?

It needs Docker with a Kali attack box, and on Windows that means WSL2 with Docker Desktop's WSL integration enabled, because workspace commands require a POSIX shell. Native PowerShell and Windows SSH work hosts are not supported. Bring a model through Settings -> Models after the first start.

Can Pentest Copilot use my existing ChatGPT or Claude subscription?

It can use an authenticated Codex CLI in Docker or host mode, and Claude Code in host or developer mode, as inference providers. The official CLI keeps ownership of login, refresh and entitlement, and the tool does not copy or replay OAuth tokens, so the tool loop and consent checks behave the same as with API providers.

Does Pentest Copilot ask before running dangerous commands?

It does. The README states that dangerous commands such as recursive deletes, device writes and fork bombs require explicit approval even in auto-run mode, so enabling automatic iteration does not remove the consent check.

Official sources

  1. bugbasesecurity/pentest-copilot on GitHub
  2. Issues
  3. License: MIT
  4. Project website
  5. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/bugbasesecurity-pentest-copilot.svg)](https://hysenlabs.com/projects/bugbasesecurity-pentest-copilot)