Model or dataset
S1N6H/pentest-harness avatar
S1N6H/pentest-harness

Pentest Harness makes the model, the tools, and the session all replaceable

Pentest Harness — Heaven for Hackers. A self-hosted AI agent harness for authorized pentests, bug bounty, security labs, and CTFs. Bring your own AI model API; sessions stay local.

418 stars66 forksTypeScriptMIT

At a glance

What is it?
Pentest Harness is a self-hosted agent workspace for authorized security engagements, built on Cordis so that every layer, including the model adapter, the toolset, session storage, and credential handling, is a plugin rather than a core you cannot replace. It ships a Pentest Mode operating standard, a dark-only interface, and a credential store that keeps API keys out of your settings file.
Who is it for?
Pentest Harness fits a security practitioner who wants a local agent workspace, brings their own model key, and would rather replace a component than patch around one. It does not fit a team that needs shared multi-user infrastructure or an audit trail, since the documented defaults are a loopback URL and a single-user credential file.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 3 days ago.
What is it written in?
Mainly TypeScript, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 10, 2026, and from our analysis. They are not legal advice.

Editorial analysis

Four commands and a loopback URL, and the only hard requirement is a container runtime

The quick start is four commands:

sh
git clone https://github.com/S1N6H/pentest-harness.git
cd pentest-harness
pnpm install
pnpm build
pnpm dsh web

The web interface comes up at `http://127.0.0.1:2323`. Requirements are Node.js 20 or newer with Node 22 recommended, and pnpm. If the `pnpm dsh` binary is not on your PATH, the documented alternatives are `pnpm exec dsh web` or invoking `node apps/cli/lib/index.js web` directly, which is a useful thing to know when a global install and a workspace install disagree about where the CLI lives.

Changing the port is a flag on the same command:

sh
pnpm dsh web --port 3000

Note what is not in the requirements list. There is no database server, no message broker, and no cloud account. The defaults table gives three paths and that is the whole configuration surface at startup: the web UI address, a settings file at `$DSH_HOME/settings.yaml` which defaults to `~/.dsh/settings.yaml`, and a credentials file at `$DSH_HOME/.credentials.yaml` with owner-only permissions.

Two details there are worth more than they look. The settings and credentials files are separate files, not one file with a section for secrets. And the loopback address means the UI is not reachable from another machine by default, which is the right default for a tool holding shell access to the machine it runs on.

Cordis is the substrate, and the claim is that there is no closed core to fight

The architectural claim is stated in one line and it is the line to evaluate the project on: everything is a plugin, built on Cordis, with no closed core to fight.

Taken literally, that covers every layer the README names as replaceable from configuration: model adapters, tools, sessions, settings, and credentials. The Cordis framework is what makes that coherent, because it is a plugin container with its own lifecycle rather than a set of hooks bolted onto a fixed application.

The repository layout backs the claim up. The workspace list in the root manifest covers `vendor/*`, `packages/*/*`, `native/landlock-run` and its own packages, `apps/*`, and `website`. Vendored dependencies sitting in the workspace means they are built from source alongside the project rather than consumed as published tarballs, which is a choice you inherit whether you wanted it or not.

The examples directory is the clearest evidence that the plugin architecture is meant to be used rather than admired. Named examples include an ACP agent, a headless agent, a JSON-RPC agent, an MCP memory integration, a Cordis web example, and a scheduling example. Those six cover the integration surfaces that matter for an agent harness: another protocol, no interface at all, a structured transport, a memory backend, the plugin host itself, and timed execution.

A native directory named for Linux Landlock is also in the workspace, which suggests the toolset is meant to be confined rather than merely sandboxed by convention. Whether that confinement is enforced on your platform is worth checking before you rely on it.

Any OpenAI-compatible endpoint, and the provider ID fills itself in

Model connection is the first-run flow and it is built around the assumption that you already have a key and a preference about vendors. The named providers are OpenAI, Anthropic, DeepSeek, Google, Mistral, Groq, OpenRouter, and Azure OpenAI, plus any OpenAI-compatible gateway.

Underneath, the engine speaks three protocol shapes: OpenAI Chat Completions, the OpenAI Responses API, and Anthropic Messages. DeepSeek is called out alongside them. Anything that speaks one of those, or a gateway that speaks OpenAI, gets model auto-discovery.

The six-step flow is worth following literally because two of the steps are less obvious than they look. You open Settings and go to Models, click Add a custom provider, and paste your API base URL. The provider ID and display name auto-fill from what you pasted, which means a typo in the URL produces a confidently named wrong provider rather than an error. Then you enter the key, models auto-discover from the endpoint, and you click one to add it or type one by hand.

The provider cards in the main interface are the part that turns this from a config form into a tool you can trust at a glance: live connection testing, an enabled and disabled toggle per provider, and a context badge per model. The badge matters more than it looks, because the difference between a model that fits your context budget and one that silently truncates is invisible until something fails oddly.

Pentest Mode is a single sentence in the README and the whole weight of the project

One feature is listed as Pentest Mode, described as a professional offensive-security operating standard for authorized engagements. That is the entire description in the README.

Everything else in the toolset is general-purpose, and the features list is explicit that it is: a shell, a filesystem, web research, skills, goals, subagents, background jobs, and workflow control. Read on its own, that is a capable general agent harness with a pentest label. What makes it something else is the operating standard applied over the top.

That is a documentation artefact rather than a technical one, which has two consequences. The good one is that an operating standard can encode judgement that no tool call expresses: when to stop, what to write down, what requires a human decision before it happens. The cost is that the README does not tell you what is in it, so you would be adopting an unseen policy.

The project targets authorized work, and the README says so in the same breath as the offensive-security framing: authorized penetration tests, bug bounty research, security labs, and CTF engagements. The scope statement is part of the tool's design, not a disclaimer bolted on afterwards. If your engagement is not one of those four, the mode is the wrong default regardless of how good the shell integration is.

The last line of that feature list is the one that makes the tool usable during a long session, and it is unrelated to offense: subagents, goals, and background jobs let work continue while you are looking at something else.

Sessions persist to JSONL or SQLite and can be replayed

Durable sessions are described as JSONL and SQLite persistence with replay, and the promise is that you resume exactly where you left off.

That promise is doing more work than it first appears to. A pentest engagement is not a conversation; it is a sequence of observations over hours or days, often across a context window that the findings themselves will exceed. Losing the thread means re-deriving what you already tried, and re-deriving is where an agent wastes tokens and misses a thread it had already pulled.

Replay is the more interesting half. A JSONL session file is a transcript, and a transcript that can be replayed means you can go back and look at what the agent actually did rather than what its summary says it did. For work where you may need to justify a step later, that difference is the whole value.

The companion feature is context management under the heading of context that never dies: token metering, automatic compaction, and tool-result pruning. Metering tells you where the budget went, compaction is what keeps a long session alive, and pruning is the one to think about, because a pruned tool result is a detail the agent will not have. Whether that is acceptable depends entirely on the task. A reconnaissance sweep tolerates losing a verbose curl output; a step where the exact bytes of a response mattered does not.

Note the persistence format is plural. JSONL and SQLite are both offered, and the choice is a plugin decision, which is consistent with the rest of the architecture.

Keys are references, not values, and the split is enforced by the file layout

The security claim is specific enough to check: API keys live in an owner-only credential store as references, and never in settings files or logs.

The mechanism behind that word references is worth understanding. If the agent configuration contains a reference rather than the key itself, then anything that reads the configuration, serialises it, or writes it into a log gets a pointer instead of a secret. The defaults table shows the physical arrangement: `settings.yaml` for configuration, `.credentials.yaml` next to it with owner-only permissions for the values.

Two files in one directory is a small design choice with a real consequence. You can commit your settings, or a copy of them, without leaking keys, as long as the second file stays out of version control. The risk is the inverse: someone who sees a settings file may assume it is the whole configuration, and paste a key into it. The README's own claim is that the credential store holds references, so a key pasted into a settings field is not stored the way the feature intends.

The dark-only theme is a smaller item in the same list. It is described as a focused, low-glare interface built for long engagements, which is a usability claim rather than a security one, and a defensible one for a tool someone stares at for hours.

There is a third-party notices file and a security policy in the repository root, plus a set of brand guidelines with translated variants, which is more governance surface than most projects of this size carry.

A 0.1.1 release candidate, seven Vitest configurations, and no GitHub releases

The version in the root manifest is `0.1.1-rc.2`, the package is marked private, and the license is MIT. There are no GitHub releases, and the last push to the repository was on 2026-09-11.

A release candidate is worth pausing on. The plugin architecture, the model adapters, and the Pentest Mode standard are all things you would want to read before trusting with shell access, and none of that implies a stable tag. There is no changelog to diff between versions, which means tracking upstream means watching commits.

The test configuration tells you what the maintainers consider distinct enough to test separately. There are separate Vitest setups for the base run, end-to-end, snapshots, the web client, web stress, and web performance, plus a shared config and a coverage-partition script. A performance and a stress configuration for a web UI is a project that expects people to leave sessions running, which matches the long-engagement framing.

The rest of the toolchain is unglamorous and thorough: tsdown for bundling with separate host and client build faces, a custom build script with an official profile, oxlint for linting, jscpd for duplication detection with its own config, knip for unused exports, lefthook for git hooks, and a Python component alongside the TypeScript with its own pytest configuration.

The last of those is the one to notice. A JavaScript project with a pytest configuration and a `python/` directory at the root is doing something in Python on purpose, and the README does not say what.

Editorial conclusion

Pentest Harness fits a security practitioner who wants a local agent workspace, brings their own model key, and would rather replace a component than patch around one. It does not fit a team that needs shared multi-user infrastructure or an audit trail, since the documented defaults are a loopback URL and a single-user credential file. Before you point an agent at a target, read the Pentest Mode definition in the documentation and confirm it matches the engagement you are actually authorized to run, since that operating standard is the part of this project that carries the most weight and the README describes it in a single line.

Frequently asked questions

What does pentest stand for?

It stands for penetration testing. Pentest Harness describes itself as a dark-first AI agent harness for authorized penetration tests, bug bounty research, security labs, and CTF engagements, and it includes a Pentest Mode described as an offensive-security operating standard for authorized work.

Is pentesting illegal?

The scope Pentest Harness is built for is authorized work: penetration tests, bug bounty research, security labs, and CTF engagements. The Pentest Mode feature is described as an operating standard for authorized engagements, and the tool gives an agent a shell, so what it is pointed at is the operator's responsibility.

What does harness mean in coding?

In this project it means the surrounding runtime that drives an agent rather than the agent itself. Pentest Harness is built on Cordis so that model adapters, tools, sessions, settings, and credentials are all replaceable plugins, and the harness is the thing that hosts them.

Which AI models can I use with Pentest Harness?

You bring your own key from OpenAI, Anthropic, DeepSeek, Google, Mistral, Groq, OpenRouter, Azure OpenAI, or any OpenAI-compatible gateway. The engine speaks OpenAI Chat Completions, the OpenAI Responses API, and Anthropic Messages, and models auto-discover from the endpoint you paste in.

How are API keys stored in Pentest Harness?

Keys live in an owner-only credential store as references, never in settings files or logs. The defaults are a settings file at `$DSH_HOME/settings.yaml` and a credentials file at `$DSH_HOME/.credentials.yaml` with owner-only permissions.

Official sources

  1. Issues
  2. License: MIT
  3. Project website
  4. README
  5. S1N6H/pentest-harness on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/s1n6h-pentest-harness.svg)](https://hysenlabs.com/projects/s1n6h-pentest-harness)