# hackingBuddyGPT verifies success against real command output, and its last release predates a year of commits

> A framework for LLM driven security testing agents that ships privilege escalation, web, API and Active Directory experiments as sub-commands of one CLI. Success is checked against ground truth rather than the model's claim, runs are capped by four independent limits, and every provider is reached through a single upstream library.

**ipa-lab/hackingBuddyGPT** — Helping Ethical Hackers use LLMs in 50 Lines of Code or less..

- Repository: https://github.com/ipa-lab/hackingBuddyGPT
- Website: https://hackingbuddy.ai/
- Stars: 1,249 · Forks: 218
- Language: Python
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/ipa-lab-hackingbuddygpt

## Success is checked against command output, not against the model's word

The single most important design decision in this framework is that a claimed compromise is verified rather than believed.

The minimal tool-calling use-case makes that explicit: its success tool is checked against ground truth, so a hallucinated or prematurely conceding announcement that root was obtained cannot score a false success. The same mechanism is listed as a framework feature, ground truth verified success detection for privilege escalation.

The connection settings carry the other half of it. Besides the target address and credentials there is a separate hostname value, described as the hostname the target reports, and it exists for root and success detection. So the check is not only local: the framework also asks the target what it thinks it is called.

The warning at the top of the readme is the reason this matters more here than in most agent frameworks. This software executes real commands on live systems. In local shell mode those commands run on your machine, and in SSH or psexec mode they run on the target you point it at. The stated advice is to run it only against systems you own or are explicitly authorised to test, and to prefer isolated VMs or containers.

## Two execution styles share one loop, and the old one is a template plus a parser

There are two ways to drive a target here, and both sit on the same shared loop.

The older style is a simple text command strategy. A Mako template renders the whole history into a single prompt each round, the model replies, and the framework parses one bare command out of the reply and runs it. That is described as the classic hackingBuddyGPT loop, and the minimal privilege escalation use-case built on it is around twenty lines long, which is where the fifty lines or less claim in the project description comes from.

The newer style is a native tool calling agent: a real chat history is kept and the model drives the target through function calls. The two styles are not separate frameworks. They share the loop, which is why a use-case can be small in one style and elaborate in the other without the surrounding machinery changing.

The practical difference is where the state lives. The template style resends history as text and parses a command out of prose, while the tool calling style keeps conversation state and lets the model emit structured calls. The first is easy to read and hard to constrain; the second is the opposite, and the framework offers both because a twenty line experiment and a research run need different things.

## Ten experiments are sub-commands, and one of them calls another

Every experiment is registered as a sub-command of a single CLI, and the catalogue is grouped by target type rather than by difficulty.

Five work on privilege escalation. The two minimal ones differ only in execution style, with the tool calling twin adding verified success detection. A full featured Linux case adds optional retrieval augmented generation through a rag path option, chain of thought, state tracking and structured guidance, and its Windows counterpart drives the target through psexec instead of SSH. The fifth is the interesting one: it runs a Linux smart enumeration script on the target first, turns that output into hints, and then orchestrates the full Linux case once per hint. It is an example of a use-case calling another use-case.

Three more test the web. One works over HTTP and lets the model narrate its reasoning with an OWASP style playbook capability, one adds shell access to a Kali style attacker box, and one is a top level agent with no direct target access that delegates to bounded sub-agents.

The last two cover a REST API, which is pentested after detecting either an OpenAPI specification or a website sitemap, and Active Directory, where a persistent strategic planner maintains a task tree and hands each task to a fresh, memoryless executor. That last design was ported from an existing attack tool rather than written here.

## Four independent ceilings, and one of them is in dollars

Run limits are the part of this framework most likely to save you money, and they are four separate caps rather than one.

A run can be bounded by rounds, by tokens, by cost in dollars and by wall-clock duration, exposed as limit flags for maximum rounds, tokens, cost and duration. They are described as unified, which means one run inherits all four rather than choosing between them, so a loop that is cheap per iteration but never terminating still hits the round or time ceiling.

The fourth kind of limit is on the model side. A context size setting in the environment file is described as the maximum context size of the model, used for prompt trimming, and the example value shipped in the template is 16385. That number is the one to get right, because the template style resends the entire history every round and the tool calling style keeps it in the conversation. If the configured size is smaller than what the model actually supports, you pay for trimming on every iteration.

Because cost is one of the four ceilings, the framework is usable against a paid API without a separate budget mechanism, which is the main reason the limit flags exist at all.

## One upstream library, and every provider sits behind a single model string

Model access is deliberately not abstracted into the framework itself. A single upstream library handles every provider, and the project names it as the only LLM dependency.

The dependency file makes the reasoning explicit: that library pulls one vendor's SDK in transitively, but the project no longer depends on or imports it directly. So the framework holds no provider code of its own, and a model is chosen by a string. The shipped examples cover a router hosted model, a direct vendor model where a base URL must also be set, and a local model served by Ollama in its chat form.

Two of those settings exist for testing your own agent rather than the model. A proxy setting routes requests through an intercepting proxy such as Burp or mitmproxy, and a companion flag allows an insecure certificate for exactly that case. If you are evaluating agent behaviour or prompt handling, being able to read the wire traffic is the difference between guessing and inspecting.

The default endpoint is a router rather than a vendor, which means an unset base URL sends your prompts somewhere other than where a reader might assume.

## The documented install uses uv and the container image uses pip

Two installation paths are given, and they do not use the same tooling.

The requirement is Python 3.13 or newer. The recommended path clones the repository and lets the recommended environment manager create a virtual environment and install the project, with a comment in the readme noting that a plain virtual environment plus pip works too. Optional dependency groups exist for testing, for linting, and for the retrieval stack used by the retrieval augmented use-case.

The container takes the other path. The image is built on a slim Python 3.13 base, copies the whole repository in, installs the project with an editable pip install, and sets the entrypoint to the CLI itself. So a container run needs no command arguments at all to start, and the first thing it can do is list the registered experiments.

Running one looks like this, with the connection settings passed as flags:

```bash
wintermute MinimalPrivEscLinux \
    --conn=ssh --conn.host=192.168.122.151 \
    --conn.username=lowpriv --conn.password=trustno1
```

Note where that points: a lab address, a low privilege user and a throwaway password, which is the shape of a deliberately vulnerable practice box rather than anything you should aim at.

## The last release is a year old, and the version number still says 0.5.0

The repository is actively changed and formally almost static, and both facts are visible in the metadata.

The newest published release is 0.5.0, from August 2025, preceded by 0.4.0 in April 2025 and 0.3.0 in August 2024. The last push to the main branch came at the start of October 2026. So the branch has more than a year of work that no tag describes, and the version in the project file is still 0.5.0, matching the oldest of those three recent releases rather than describing current code.

The classifiers call it beta, the licence is MIT, and the maintainer list is two named people with an institutional address, which matches the project's origin as reproducible security research rather than a commercial product. That origin shows elsewhere too: a citation file, a publish notes file, a benchmark script at the repository root, a separate privilege escalation benchmark maintained alongside it, and open access reports built from that benchmark.

One packaging detail is worth knowing if you install from source. The distribution name normalises to a lowercase form while the source lives under a mixed case directory, so the module name is pinned explicitly in the build backend configuration to keep the import working.

## Conclusion

This is a research framework for authorised testing, and its own readme is blunt about the boundary: it executes real commands on live systems, on your machine in local shell mode and on the target in SSH or psexec mode, so use it only on systems you own or are explicitly authorised to test, inside a throwaway VM or container. Within that boundary it is unusually careful about the two things that make agent security tools untrustworthy, which are claiming success it did not achieve and spending without a ceiling. Verify the licence and requirements before you start, since the package requires Python 3.13 or newer, and note that the newest published release is version 0.5.0 from August 2025 while the branch has moved on for more than a year.

## FAQ

### Which LLM is best for pentesting with hackingBuddyGPT?

The project does not name a winner; it publishes a paper comparing multiple LLMs for this purpose and treats model choice as configuration. Every provider is reachable through one upstream library via a model string, including a router hosted model, a direct vendor model with a base URL, and a local Ollama model. The context size setting is what matters operationally, because it drives prompt trimming.

### What is hackingBuddyGPT used for?

It is a framework for building LLM driven security testing agents. It supplies the groundwork: LLM connectivity, target connectors over SSH, a local shell or psexec style remote execution, capability and tool wiring, run limits and structured logging, so a new experiment can be written as a use-case in a few dozen lines. Each use-case becomes a sub-command of the wintermute CLI.

### How do I run hackingBuddyGPT against a target?

Install it with Python 3.13 or newer, copy the example environment file to .env and fill in the key and connection settings, then run wintermute with no arguments to list every registered use-case and wintermute with a use-case name and --help for its options. Connection settings can be passed as flags on the command line, which override the .env file.

### How does hackingBuddyGPT stop an agent from claiming false success?

By checking claims against ground truth. In the tool calling privilege escalation use-case the success tool is verified against real command output, so a hallucinated or conceding announcement that root was obtained cannot register as a success, and the framework lists ground truth verified success detection as a core feature. A separate target hostname setting is used for root and success detection.

### Can I cap what hackingBuddyGPT spends on a run?

Yes, with four unified limits applied to every run: maximum rounds, maximum tokens, maximum cost in dollars and maximum wall-clock duration, all set through limit flags. Separately, a context size setting caps the model's context and is used for prompt trimming, which matters most in the template style that resends the whole history each round.

### Is it safe to run hackingBuddyGPT on my own machine?

Not without isolation. The readme states that the software executes real commands on live systems: in local shell mode they run on your machine, and in SSH or psexec mode on the target you point it at. The stated guidance is to run it only against systems you own or are explicitly authorised to test, and to prefer isolated VMs or containers.

## Sources

- [ipa-lab/hackingBuddyGPT on GitHub](https://github.com/ipa-lab/hackingBuddyGPT)
- [License: MIT](https://github.com/ipa-lab/hackingBuddyGPT/blob/main/LICENSE)
- [Project website](https://hackingbuddy.ai/)
- [README](https://github.com/ipa-lab/hackingBuddyGPT/blob/main/README.md)
- [Releases](https://github.com/ipa-lab/hackingBuddyGPT/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/ipa-lab-hackingbuddygpt
