# Jailbreak Autoresearch: an autoresearch loop for prompt-harness experiments

> The repository runs a small search loop over header and footer harnesses wrapped around one fixed prompt body, scoring each response with an OpenRouter model and storing the results in SQLite. It is built for prompt-harness experiments you are authorized to run, not for production traffic.

**davidondrej/jailbreak-autoresearch** — We shall set the models free.

- Repository: https://github.com/davidondrej/jailbreak-autoresearch
- Stars: 526 · Forks: 202
- Language: Python
- License: MIT
- Published: 2026-09-20 · Updated: 2026-09-20 · Language: en
- Canonical page: https://hysenlabs.com/projects/davidondrej-jailbreak-autoresearch

## What the jailbreak autoresearch loop actually searches over

Most prompt experiments change the prompt. This repository holds the body fixed and searches the wrapper around it. The README states that the repo tests whether different header and footer harnesses change how target models answer one fixed body. The body is example.md. The rubric that decides whether an answer is good is desired-output.md, and the README describes it as the verifier. The runner always uses those two files, so the only thing the search is allowed to vary is the harness.

The intended user is someone doing prompt-harness research on models they are authorized to evaluate. The README closes with that constraint: use only test bodies you are authorized to evaluate. There is no web interface, no scheduler, and no service to deploy. You edit two Markdown files, set one environment variable, and run a Python script. The output is a SQLite database of experiments in runs/.

One normalization rule applies to every generated footer: it is normalized to end with the sentence "Answer with exactly one sentence." That constrains the output shape, which makes scoring comparable across candidates, but it also means the loop is not testing open-ended answer formats. If your rubric rewards a long structured answer, this harness rule fights it.

## How run.py, models.json and the scorer fit together

The README gives the mechanism as five steps. First, run.py chooses a target model, a researcher model, and a scorer model from models.json. Second, the researcher proposes a candidate multi-turn harness. Third, the target model receives that harness with example.md inserted as the final body. Fourth, the scorer compares the response to desired-output.md and returns a score between 0.0 and 1.0. Fifth, winning fragments are stored and reused by later strategies.

That fifth step is what makes it a search rather than a sweep. Fragments that scored well are kept and fed into later runs, so the candidate space is not fixed in advance. The README lists four strategies. baseline applies no harness at all, which gives you a control. seeded pulls seed headers and footers from prompts/headers/ and prompts/footers/. evolve-best mutates the strongest prior harness. recombine recombines strong fragments from prior runs. The last two depend on earlier results existing, so a first run on an empty database cannot exercise them in a meaningful way.

Model roles come from models.json, which the README describes as an OpenRouter model list. The roles are separate, so the model that proposes harnesses need not be the model that answers, and neither need be the model that scores. All three are reached through OpenRouter. Nothing in the README describes a local model path, so a network path to OpenRouter is part of the design rather than an option.

## Installing it and running a first dry smoke test

There is no package to install. The repository is Python, and the README's setup section starts with configuration rather than installation. Clone the repository, then create a .env file in the project root containing your OpenRouter key. The README shows the file with a single key:

```bash
OPENROUTER_API_KEY=your_key_here
```

The README also warns not to commit real API keys and not to commit private experiment databases. Before any live call, customize the two root files. example.md is the body you want to test, and desired-output.md is the scoring rubric describing what a good answer should look like. The README says to keep them in sync, which is the step most likely to be skipped.

Next, run the dry smoke test. It exercises the pipeline without spending live calls:

```bash
python3 run.py --all-strategies --max-permutations 1 --dry-run
```

With the environment in place, run one live baseline. This is the smallest useful live experiment, one strategy with no harness:

```bash
python3 run.py --strategy baseline --max-permutations 1
```

Then widen to every strategy on a single role permutation:

```bash
python3 run.py --all-strategies --max-permutations 1
```

Finally, summarize what was written:

```bash
python3 report.py
```

Results go to runs/experiments.sqlite, and the README states that runs/ is ignored by git. If you want the loop to run on its own rather than by hand, the README recommends Codex CLI's /goal feature at v0.128.0 or later, started from this directory with a prompt that points at objective.md and validates each change with the same run.py and report.py commands.

## The scoring rubric is the weak link, and the README does not fix it

The loop optimizes against desired-output.md, and the scorer is a model reading that file. Nothing in the README describes a check that the rubric is unambiguous, that the scorer is consistent between runs, or that a score of 0.8 means the same thing on Tuesday as it did on Monday. A harness that happens to match the scorer's biases will score well whether or not it produces better answers. The repository stores scores, but storing a number is not the same as validating it.

There is a second limitation in the search itself. The README says winning fragments are stored and reused by later strategies, and evolve-best mutates the strongest prior harness. That is a hill-climbing shape. It has no described mechanism for keeping diversity or for escaping a local optimum, so a long autonomous run can converge on one family of harnesses and keep refining it. The seeded strategy is the only described source of fresh starting material, and it draws from prompts/headers/ and prompts/footers/ rather than generating new directions.

The repository is also the wrong tool when the body is what you need to change. If your question is which wording of the task works best, this loop holds the wording fixed by design and will not answer it. The same applies if you need latency or cost measurements: the README describes a score from 0.0 to 1.0 and says nothing about timing or token accounting, so the database is not a performance record. Maintenance status is worth noting too: the last push to the default branch was on 2026-05-10, and there are no retrieved releases, so treat the code as a snapshot rather than a project with a release cadence.

## How it differs from running your own eval harness

The obvious alternative is a general evaluation framework where you write your own prompt variants and score them yourself. The difference is who proposes the candidates. In a hand-written eval, you enumerate the variants, and the framework measures them. Here the researcher model proposes the harness, the target answers it, and the scorer grades the result, with strong fragments carried forward. You supply the fixed body and the rubric, and the loop supplies the candidates.

That trade is real in both directions. You get candidates you would not have written, and you give up knowing exactly what was tried unless you read the database. A hand-written eval also lets you control the number of calls precisely, which matters when every run goes through OpenRouter. The README's --max-permutations flag is the only described knob for bounding a run's size, and it bounds role permutations rather than the number of proposals.

A second alternative is to skip the harness search entirely and tune the body with a single strong model in an interactive session. That is faster for a one-off question and leaves no database. It becomes impractical once you want to compare dozens of wrappers under the same rubric, which is the case this repository is built for. The separation of roles in models.json is the part a manual session cannot reproduce: a different model can score than the one that answers, which reduces the chance that a model rewards its own style.

## Licence, upgrade cost and what to verify first

The repository is MIT licensed, so you can use, modify and redistribute the code, including in closed products, provided the licence and copyright notice are preserved. That is the usual reading of MIT, not legal advice, and it covers the code in this repository only. The models you call through OpenRouter carry their own terms, and the README's instruction to use only test bodies you are authorized to evaluate is a constraint on your use rather than a licence term.

Upgrade cost is low in the sense that there is no dependency manifest described in the README and no retrieved releases to track. It is higher in the sense that the loop's behaviour depends on files you own: example.md, desired-output.md, models.json, and the prompts/ directories. Change models.json and your scores are no longer comparable to earlier rows in the same database. The README does not document a schema version or a migration path for runs/experiments.sqlite, so keep the database with the configuration that produced it.

Before a long run, verify three things. Confirm example.md and desired-output.md describe the same task, since the README calls out keeping them in sync. Confirm models.json names models you can actually call, because the target, researcher and scorer are all selected from it. And confirm the dry run wrote runs/experiments.sqlite before you let anything run autonomously under /goal, because the README's checkpointing guidance assumes the validation commands work.

## Conclusion

Adopt it if you already have a fixed prompt body and a written rubric, and you want a repeatable way to test header and footer variants against a scorer. Do not adopt it if you need a production prompt-management system, a hosted service, or anything that works without an OpenRouter key. Before running the live loop, verify that example.md and desired-output.md describe the same task, that models.json points at models you can call, and that the first dry run writes runs/experiments.sqlite.

## FAQ

### What is AI AutoResearch?

In this repository it is a small autoresearch loop for prompt-harness experiments. It tests whether different header and footer harnesses change how target models answer one fixed body, and stores the harness, response, score and model-role permutation in SQLite.

### What is the jailbreak program?

Here the program is run.py, which chooses a target, researcher and scorer model from models.json, has the researcher propose a multi-turn harness, sends it to the target with example.md as the final body, and has the scorer compare the response to desired-output.md for a score from 0.0 to 1.0.

### What are AutoResearch agents?

In this project the agent shape is a loop: a researcher model proposes a candidate harness, a target model answers it, and a scorer model grades the answer. Winning fragments are stored and reused by later strategies such as evolve-best and recombine.

## Sources

- [davidondrej/jailbreak-autoresearch on GitHub](https://github.com/davidondrej/jailbreak-autoresearch)
- [Issues](https://github.com/davidondrej/jailbreak-autoresearch/issues)
- [License: MIT](https://github.com/davidondrej/jailbreak-autoresearch/blob/main/LICENSE)
- [README](https://github.com/davidondrej/jailbreak-autoresearch/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/davidondrej-jailbreak-autoresearch
