# aaabench hands a coding agent a brief and takes hints off the table

> SelfStarter, the harness in the aaabench repository, boots a real Unreal Engine editor over MCP, gives an agent a written brief, and then stays out of the way across sessions. The design bet is that whether the agent notices its own mistakes is the result worth measuring, so the operator is barred from offering diagnosis, fixes, or answers.

**ukanwat/aaabench** — A long-horizon benchmark harness: give a coding agent a real game engine, professional conditions and time, and ask it to build an open-world game. Harness only, no results.

- Repository: https://github.com/ukanwat/aaabench
- Website: https://utkarshkanwat.com
- Stars: 388 · Forks: 71
- Language: Shell
- License: MIT
- Published: 2026-09-17 · Updated: 2026-09-17 · Language: en
- Canonical page: https://hysenlabs.com/projects/ukanwat-aaabench

## The one rule is that you may not help, and breaking it invalidates the run

The harness states a single constraint and treats it as the point of the exercise. An operator provides conditions, resources, and the brief. Diagnosis is not allowed, the fix is not allowed, an answer is not allowed. The reasoning is stated plainly in the repository: whether the model notices its own mistakes is the whole point of the run, so every hint is a result you can no longer claim.

That framing explains a lot of the design choices downstream. If the agent cannot be rescued, the harness has to give it everything else: a real editor rather than a mock, a handbook to consult, sensors to look at its own output, and a rule about when a human decision is actually requested.

The boundary between operating the harness and doing the agent's job is written down in a file called HARNESS-RULES.md, and the path table marks it as the document to read first. Keeping that line in a separate file rather than in the prompt itself is a deliberate choice. An operator running an unattended session will be tempted at the exact moment they should not act, so the rule needs to be somewhere they can re-read between sessions rather than something buried in the brief.

## bin/run-agent.sh keeps one session alive while the editor dies underneath it

The starting sequence is four commands, and each one has a distinct job:

```bash
git clone https://github.com/ukanwat/selfstarter && cd selfstarter
cp -R project AgentCity        # the project skeleton
./bin/setup-capabilities.sh    # plugins, renderer features, python libraries
./bin/run-agent.sh             # boots the editor, hands over the brief, keeps it going
```

`bin/run-agent.sh` is described as one session. It boots the editor, waits for MCP, hands over the demand, resumes the session if it stops early, and relaunches the editor if it dies. That last pair of behaviours is the part that distinguishes this from a wrapper around a chat prompt. A long run is expected to hit an editor crash, and the harness treats the crash as an event to recover from rather than a run to abandon.

For repeated runs there is `bin/run-many.sh`, which does sequential unattended sessions under a single-instance lock. The lock matters because two agents writing the same Unreal project would produce a result nobody could interpret. And `bin/prep-project.sh` exists for the case where you are starting from nothing rather than copying the `project/` skeleton.

## setup-capabilities.sh widens what the agent can reach, and is safe to rerun

An agent driving a game editor through a scripting interface hits a wall quickly if the editor is in its default state. `bin/setup-capabilities.sh` exists to remove that wall, and it is specific about the four things it changes: it enables engine plugins in the `.uproject`, turns on renderer features that ship switched off, installs the Python libraries a world generator wants, and adds the optional local image, mesh, and audio tools.

That last item is the one worth pausing on. The agent is expected to produce assets, and doing that without a network service means local generation tooling has to exist on the machine first. So the harness's image and mesh tools are a precondition for the work rather than an add-on, and whether they are installed changes what a run can attempt.

The script is described as idempotent, which is a small but load-bearing word in a harness like this. Sessions get resumed and editors get relaunched, so setup will be run more than once against the same project directory. A setup step that is not idempotent becomes a source of divergence between run one and run four, and divergence between runs is exactly the variable a long-horizon benchmark cannot afford.

The `project/` directory holds the `.uproject` with the required plugins and the config that auto-starts the MCP server. Copying it to `AgentCity/` is the documented way to begin, which means the skeleton is versioned with the harness rather than generated per run.

## tools/ue_qa.py is a viewport camera, and self-verification is a scored capability

The harness gives the agent eyes. `tools/ue_qa.py` is a sensor for viewport capture and inspection, so the agent can screenshot its own work and look at it. Alongside that, the run includes play-in-editor and reference photographs of the real world.

Why this matters is spelled out in the list of what a run tests. Self-verification is its own item: does the agent look at its own work, notice what a stranger would call fake, and fix it without being told. The same item repeats the constraint that nobody points at anything for it. Whether it notices is the point.

That makes `tools/ue_qa.py` a measurement instrument rather than a convenience. Without it you would be grading a world from screenshots the harness took on your behalf, and the difference between an agent that iterated and one that got lucky on the first render would be invisible. With it, the iteration is something the agent chose to do.

The tools directory is described as what the agent uses for eyes, widgets, and image generation, so this is a small toolkit rather than a single script. The widgets are what let the agent drive editor UI it cannot script directly, which is how a build ends up with the game's own screens, a map with named districts, rather than only geometry.

## docs/ is a handbook the agent may consult, and .claude/skills/ holds 21 packs

Two directories carry the knowledge. `docs/` is the handbook, and its contents describe what kind of reference material it is: production workflow, level pipeline, systems budgets, detail and density, parallelism, a world inventory of hundreds of kinds of real-world object mined from OpenStreetMap, sources for assets, mocap, map data, and rendering, and the engine's version traps.

That last category is the clearest signal about what kind of failure this harness is built around. Version traps are the places where a model recalls an API that has moved, and the run description names this directly as part of engineering under a hostile surface: a real editor that crashes, an API it has to discover rather than recall, tools that fail silently, and a generator whose output has to be validated because plausible numbers describe impossible places.

The second directory is `.claude/skills/`, holding 21 skill packs. Named areas include game AI, level design, game feel, shaders, Niagara, Blueprints, Enhanced Input, behaviour trees, physics tuning, camera, dialogue, audio, save systems, performance, and reference-image search. The count is what makes it worth noting: this is not a single prompt file but a library of specialisations the agent can draw on, and it is the closest thing in the repository to a conventional plugin architecture.

The boundary between the two matters. `docs/` is reference the agent may consult. `.claude/skills/` is procedure. Keeping them apart means you can extend the knowledge base without rewriting the skills that read it.

## The bar is a street that behaves, not a level that loads

The pass criteria are stated in the README as four artefacts, and all four come from a single run: a street that behaves, with traffic, pedestrians, and a player in it; the game's own screens, built by the agent, including a map with named districts; a city that reads as one from any height; and a place rather than a level, meaning geography the city had to be built around.

The framing sentence is worth repeating because it sets the whole difficulty. The bar is not that a level loads, it is that a place survives a stranger looking at it, and a game that opens like a game.

What makes this a hard test rather than a demo is the reasoning requirement underneath it. The brief demands that nothing be placed because it looked good somewhere else. Deep water decides where the port goes, industry follows the rail, money builds uphill and upwind, sunlight limits how tall a street can be, and a courthouse grows bail bonds around it. A world built without that knowledge looks wrong instantly, to anyone, with no expertise required to say so.

That last clause is the design decision that makes the harness usable by someone who is not a game developer. The evaluation does not require a domain expert, because a world that ignores where ports go is legible as broken to a person who has never opened Unreal.

## Early development, one example run, and no published results

Status first, because it changes how much weight the rest of this carries. The README says the project is in early development and that what exists today is an example run: one agent, one brief, a real game engine, and no human help. The repository description is blunter and describes it as a harness only, with no results.

The two statements are consistent, and the README explains how. Everything an agent produces belongs to the run that produced it, not to the repository, so the example run's artefacts are the only run output present. The screenshots and the four pass criteria are there to show what the demand asks for, not to prove a score.

There are no GitHub releases. The last push to the repository was on 2026-09-23, and the primary language is Shell, which fits a repository that is mostly scripts, Markdown, and an Unreal project skeleton rather than an application.

If you are evaluating whether an agent can hold a brief for days, this repository gives you the means to find out and not the answer. The README is explicit that it is in early development, so expect the harness itself to need adjusting, and read HARNESS-RULES.md before your first run rather than after your first violation of it.

## Conclusion

aaabench fits a research team that wants to measure what an agent does across days of unattended work rather than in a single chat, and that already has an Unreal editor and the patience to supervise a run without intervening. It does not fit a production team looking for a way to ship a game, because the repository contains a harness and one example run's worth of description, not a game and not published results. Before you start, read HARNESS-RULES.md in full and decide in advance what you will do when the agent is stuck, because the whole design depends on your not answering, and that decision is much harder to hold once you are three hours into a run.

## FAQ

### What is aaabench?

aaabench is a long-horizon benchmark harness, and the README names the project SelfStarter. It gives a coding agent a real game engine, professional conditions, and time, then asks it to build an open-world game. The repository contains the harness, with no results.

### How do I start an aaabench run?

Clone the repository, copy the `project` skeleton to `AgentCity`, run `./bin/setup-capabilities.sh` to enable plugins, renderer features, and Python libraries, then run `./bin/run-agent.sh`, which boots the editor, hands over the brief, and keeps the session going.

### Why is the operator not allowed to help the agent in aaabench?

Because whether the model notices its own mistakes is the whole point of the run, so every hint is a result you can no longer claim. HARNESS-RULES.md sets the line between operating the harness and doing the agent's job.

### Does aaabench have published results?

No. The project is in early development and the repository contains the harness only. The example run in the README is there to show what the demand asks for, and everything an agent produces belongs to the run that produced it.

### How does the agent in aaabench see its own work?

Through `tools/ue_qa.py`, a sensor for viewport capture and inspection, alongside play-in-editor and reference photographs of the real world. The tools directory also holds widgets and image generation used by the agent.

## Sources

- [Issues](https://github.com/ukanwat/aaabench/issues)
- [License: MIT](https://github.com/ukanwat/aaabench/blob/main/LICENSE)
- [Project website](https://utkarshkanwat.com)
- [README](https://github.com/ukanwat/aaabench/blob/main/README.md)
- [ukanwat/aaabench on GitHub](https://github.com/ukanwat/aaabench)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/ukanwat-aaabench
