# BoxPwnr's benchmark table has an empty Completion column and traces-per-solve ranging from 1.8 to 54

> A harness that points language models and CLI coding agents at capture-the-flag challenges and vendor security labs, then publishes every solving trace for replay. The measurement discipline is the interesting part and also the weakest: sixteen result rows with no model, solver or date attached, a column that is blank throughout, and a trace count that is not comparable across platforms.

**0ca/BoxPwnr** — A modular framework for benchmarking LLMs and agentic strategies on security challenges across HackTheBox, TryHackMe, PortSwigger Labs, Cybench, picoCTF and more.

- Repository: https://github.com/0ca/BoxPwnr
- Stars: 457 · Forks: 59
- Language: Python
- License: AGPL-3.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/0ca-boxpwnr

## The Completion column is empty in all sixteen rows

The results table is the centre of the project and it has a defect that is visible immediately.

The header is Platform, Solved, Completion, Traces. Every one of the sixteen data rows fills Platform, Solved and Traces. The Completion cell is empty in all sixteen. Not zero, not a dash: blank.

A completion percentage is derivable from the two columns beside it, so nothing is lost analytically. But the column exists, it is generated rather than typed, and it is broken in every row, which points at the generation script rather than at the data. The table is fenced by BEGIN_BENCHMARK_STATS and END_BENCHMARK_STATS comments, so something under scripts/ writes it into the README on every run. A script that emits a header it never fills is a script that has drifted, and the same script is presumably what decides whether Solved and Traces mean the same thing this month as last month.

That is the practical reason to read this table sceptically rather than as a result set. The numbers may be right. The presentation says the pipeline that produces them was not checked recently.

Rows are individually linkable to a per-platform page on a separate site, boxpwnr.info, which is where the definition of a trace, the solver used and the attempt budget would presumably live. None of that is in the repository.

## Traces per solve ranges from 1.8 to 54, so the column is not an effort measure

Dividing the Traces column by the Solved column gives a number that is not printed anywhere and should not be.

At the top: Cybench solved 40 of 40 with 2165 traces, which is 54 traces for every solved target. HTB Starting Point solved 25 of 25 with 770 traces, 30.8 each. ExploitBench solved 2 of 42 with 58 traces, 29 each. Argus solved 47 of 60 with 1026 traces, 21.8 each.

At the bottom: BSidesSF CTF 2026 solved 43 of 51 with 76 traces, 1.8 each. CyberGym solved 476 of 1507 with 977 traces, 2.1 each. HTB Challenges solved 324 of 818 with 732 traces, 2.3 each. PortSwigger Labs solved 163 of 270 with 377 traces, 2.3 each.

That is a thirtyfold spread, and it does not track difficulty. If traces measured effort, the two hardest rows should have the highest ratios, and instead Cybench, a complete solve, has the highest ratio in the table at 54.1 while BSidesSF, at 84 percent solved, has the lowest at 1.8.

What actually explains the split is target count against attempt budget. A platform with 500 targets and a fixed budget yields a small trace count per solve. A platform with 16 or 40 targets and the same fixed budget yields a large one. Cybench at 40 targets and ExploitBench at 42 targets both sit near 30 to 54, while every platform above 250 targets sits between 1.8 and 5.1. Neurogrid, at 36 targets, sits at 11.6, between the two clusters.

So Traces counts attempts, Solved counts challenges cleared at least once, and the two columns measure different units. Any comparison that puts them in the same row as if they were commensurable is misleading.

## No row names a model, a solver or a date

Look for an attribution column in that table and there is none.

Sixteen rows, four columns, and not one of them says which language model produced the numbers, which of the nine solvers was selected, how many attempts each target got, or when the run happened. The platform links go to a separate site where that context may live, but the repository does not carry it.

The consequences are concrete. The solve rates cannot be attributed to a model, so a reader cannot conclude that one model is better at security reasoning than another from this table alone. They also cannot tell how much of the variation is the model and how much is the target set. ExploitBench at 2 of 42 and CyberGym at 476 of 1507 might be two hard benchmarks, or one hard benchmark and one where a lucky heuristic clears the easy boxes, and this table cannot distinguish those.

The spread itself does track apparent difficulty in a way that suggests the targets are doing the work. The top of the table is the easiest material: HTB Starting Point at 25 of 25, Cybench at 40 of 40, picoCTF at 502 of 503, XBOW at 102 of 104. The bottom is the hardest: ExploitBench at 2 of 42, LevelUpCTF at 50 of 254, CyberGym at 476 of 1507.

That is a legitimate result for a harness that aggregates across many runs, and the project is upfront that this is an experiment rather than a claim about any vendor. The README describes it as a fun experiment to see how far models can go on CTF challenges and security labs. The trace site makes the underlying data inspectable, which is the right mitigation, since you can go and read what actually happened on a specific target rather than trusting the aggregate.

## Fifteen platform flags against sixteen result rows, and three platforms you cannot select

The command surface and the published results disagree about what exists.

The --platform flag accepts fifteen values: htb, htb_ctf, htb_challenges, portswigger, ctfd, local, xbow, hackbench, cybench, cybergym, exploitbench, picoctf, tryhackme, levelupctf, argus.

The results table has sixteen rows, and three of them have no flag: BSidesSF CTF 2026, Cloud Village CTF 2026, and Neurogrid CTF, described in its own link text as the ultimate AI security showdown. Two more rows, HTB Starting Point and HTB Labs, share the single `htb` value rather than having flags of their own.

So the mismatch runs both ways. Three event CTFs appear in the results with no way to request them from the CLI, and two rows collapse into one flag. Meanwhile ctfd and local are selectable but produce no row in the table at all. ctfd is the platform family that most self-hosted challenge servers implement, and local is the one you would use to point the harness at something on your own machine.

The likely explanation is that the results are an aggregate across everything ever run, including event targets reached through a path the flag list has since been narrowed to describe. Either way, someone reading the table to decide what they can run locally has to cross-reference two lists that do not line up.

Per-platform documentation exists at src/boxpwnr/platforms/README.md, which is where the ctfd and local adapters would be described.

## For five of the nine solvers, BoxPwnr wraps somebody else's agent loop

The solver list is nine values and they are not nine comparable things.

Five are CLI agents that ship their own loop: claude_code, codex, grok, cursor-cli and kiro_cli. The documentation is explicit that these run their own agent loop and stream results back to BoxPwnr. BoxPwnr supplies the environment, the target, the timeouts and the trace. The iteration strategy, the tool selection and the context management all belong to the vendor's binary.

Four are loops BoxPwnr drives itself: single_loop, single_loop_xmltag and single_loop_compactation, plus hacksynth, which has no description anywhere in the visible text. The three single_loop variants differ in output protocol and, judging by the name, in whether context is compacted. That is the actual experimental design: holding the environment fixed and varying the protocol the model must emit commands through.

That split has a measurement consequence worth stating plainly. Comparing two CLI solvers against each other compares two vendor products, which is a legitimate thing to want to know. Comparing a CLI solver against a single_loop variant compares a vendor's agent against a bare prompt loop, which measures the difference between a full agent and a loop rather than between two models. Anyone building a comparison table from this needs to say which of those two questions they are asking.

One detail worth knowing before typing: the compaction variant is spelled single_loop_compactation in the flag list. That is the literal value, misspelled, and it has to be typed exactly.

## playwright-stealth and live web search are in the base dependency list

Two entries in the dependency list say more about how the harness operates than the README does.

The first is playwright-stealth, alongside playwright itself. That is a bot-detection evasion plugin for browser automation. In a CTF and lab harness it is aimed at the platforms that sit behind bot protection, and it means the tool presents itself to those sites as a browser that does not look automated. That is a reasonable requirement for the job and also a methodological fact: runs against protected platforms are not observing the target the same way an unthrottled client would.

The second is ddgs, listed with a comment saying it is required by DuckDuckGoSearchRun for the live web_search tool. Live web search is therefore inside the default loop, not an opt-in. Agents solving a lab can search the public web while doing it, and published traces will contain those searches.

Also in the list: mcp, so the harness speaks the Model Context Protocol; tokencost, for the per-run cost accounting behind --max-cost; pdfminer.six, which fits the document-heavy PortSwigger labs; playwright for the browser; and ten LangChain integration packages, one per provider.

Two of those LangChain packages carry comments explaining why they are there, and both comments are about workarounds rather than features. langchain-google-genai is required for gemini-3-flash-preview and improved generation_config support. langchain-model-profiles is required for model.profile context window detection.

That second one matters for anyone running long-horizon attempts. Context window limits come from a profile lookup in a package versioned at 0.0.5, so the bound on how long a run can go before truncation is whatever a pre-1.0 profile package reports for a given model name.

## Ten LangChain packages with lower bounds, one exact pin, and an example.com author email

The dependency list is twenty-seven entries and the pinning policy is inconsistent in a way that shows the project's history.

Twenty-six use a lower bound with no ceiling: langchain at 1.2.0 and up, langchain-core at 1.2.0, langchain-openai at 1.1.0, langchain-anthropic at 1.3.0, langchain-community at 0.4.0, langchain-xai at 1.1.0, langchain-google-genai at 4.1.0, langchain-google-vertexai at 3.2.0, openai at 1.101.0, anthropic at 0.75.0, playwright at 1.57.0, pydantic at 2.10.6. None of them has an upper bound, so a sync can move any of them forward.

One uses an exact pin: deepseek at 1.0.0. It is the only `==` in the list, and it sits among packages whose floors are several major versions ahead of it.

The model-profiles package at 0.0.5 is a pre-1.0 dependency doing load-bearing work for context window detection, which is the kind of thing that breaks silently on a minor bump rather than failing a build.

The metadata has its own oddity. The project description in the manifest is Automated HTB machine solver. That is the original scope of the project, from when it solved HackTheBox machines. The repository description, by contrast, describes a modular framework for benchmarking across fifteen platforms. So the published package on PyPI still presents itself as single-platform while the project covers fifteen, and the manifest has not been updated since the expansion.

The author field is worse for contact purposes: the name 0ca with the email address oca@example.com, the reserved example domain. Anyone reaching out through package metadata is writing to a placeholder.

## boxpwnr/ sits beside src/boxpwnr/, and make ci-test wants your real keys

Two things in the build setup deserve checking before you rely on either.

The first is the layout. The top-level listing contains both a `boxpwnr` directory and a `src/` directory, and the version file is written to src/boxpwnr/_version.py. With hatch-vcs and a src layout, the package root is src/boxpwnr and there should be no second copy at the top. A top-level `boxpwnr` alongside `src/boxpwnr` is either a shim, a stale leftover, or a real duplicate, and the console script `boxpwnr = boxpwnr.cli:main` does not say which one gets imported. On a src-layout install the resolution is unambiguous; on a checkout run in place it may not be.

The second is the Makefile. There are three targets that simulate GitHub Actions locally with act, and two of them pass your secrets file straight into the run:

```bash
act push -W .github/workflows/ci-free-models-tests.yml --secret-file .env -P ubuntu-latest=catthehacker/ubuntu:act-latest
```

So make ci-test and make ci-integration cannot run without a populated .env holding real provider credentials. That is fine locally and wrong as a default, because a target named ci-test that needs live keys is a target a contributor runs with production keys in a file that a careless git add would publish. The README does note that keys are saved to .env on first run.

The runner image is also worth a look. Both act targets pin the ubuntu runner to catthehacker/ubuntu:act-latest, a personal Docker Hub namespace rather than a first-party image, so the environment your local CI simulation runs in is not the environment the real runner uses.

The ordinary targets are fine: make test runs pytest through uv run, and make test-changed calls scripts/pytest_changed.py. Linting is flake8 from a .flake8 config at the root, with black in the dev extra alongside it.

One last consistency check. v0.4.0 shipped 2026-07-18 and the last push was 2026-07-24, so the branch is a few days ahead of the tag. The version comes from git tags rather than a literal, and the licence is AGPL-3.0, which is worth reading alongside the fact that every solving trace is published.

## Conclusion

Use BoxPwnr to compare agent harnesses against each other on the same lab targets, and treat its published table as a leaderboard rather than a model ranking, because no row identifies a model or a solver. Do not cite the aggregate solve rates without recomputing them, since the table ships an empty Completion column and a trace count whose meaning varies thirtyfold across platforms. Verify first that every trace is published by default before you point it at anything you would rather not have transcribed, and check the dual boxpwnr directories in the tree before building on the import paths.

## FAQ

### What does BoxPwnr do and what can it be pointed at?

It runs language models and CLI coding agents against capture-the-flag challenges and security labs, resolving platforms with the --platform flag: htb, htb_ctf, htb_challenges, portswigger, ctfd, local, xbow, hackbench, cybench, cybergym, exploitbench, picoctf, tryhackme, levelupctf and argus. Commands run by default in a Docker container with Kali Linux, built automatically on first run, and a VPN is established when the platform needs one.

### Which solvers does BoxPwnr support?

Nine values: claude_code, codex, cursor-cli, grok, kiro_cli, external, single_loop_xmltag, single_loop, single_loop_compactation and hacksynth. The five CLI agents run their own agent loop and stream results back, while the single_loop variants are loops BoxPwnr drives itself, differing in output protocol and context compaction.

### What results does BoxPwnr publish?

A sixteen-row table of solve rates and trace counts across platforms, from HTB Starting Point at 25 of 25 down to ExploitBench at 2 of 42, with full conversation logs and an interactive replay viewer on a separate site. No row names a model, a solver or a date, and the Completion column is empty in every row.

### How do I install and run BoxPwnr?

Clone with submodules using git clone --recurse-submodules, install uv, then uv sync to create the virtual environment. Docker must be installed and running. Run a target with uv run boxpwnr --platform htb --target meow, and on first run you are prompted for API keys, which are saved to .env.

### Does BoxPwnr publish everything an agent does on a target?

The project states that all solving traces are available on its traces and benchmarks site, and that each trace includes full conversation logs showing model reasoning, commands executed and outputs received, replayable in an interactive web viewer. Live web search is available inside the default loop, so those searches appear in traces too.

## Sources

- [0ca/BoxPwnr on GitHub](https://github.com/0ca/BoxPwnr)
- [Issues](https://github.com/0ca/BoxPwnr/issues)
- [License: AGPL-3.0](https://github.com/0ca/BoxPwnr/blob/main/LICENSE)
- [README](https://github.com/0ca/BoxPwnr/blob/main/README.md)
- [Releases](https://github.com/0ca/BoxPwnr/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/0ca-boxpwnr
