Model or dataset
davebcn87/pi-autoresearch avatar
davebcn87/pi-autoresearch

pi-autoresearch: an autonomous experiment loop inside the pi coding agent

Autonomous experiment loop extension for pi

8,135 stars464 forksTypeScriptMIT

At a glance

What is it?
pi-autoresearch is a pi extension that turns an optimization goal into a repeatable loop of run, measure, keep or revert. It suits engineers with a command that prints a number, and it depends entirely on that command being honest.
Who is it for?
Adopt pi-autoresearch if you already run pi and can express your target as a command that prints METRIC name=number lines, because the loop is only as trustworthy as that measurement. Skip it if your goal has no automatable metric, or if you cannot review the commits the loop writes.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 20 days ago.
What is it written in?
Mainly TypeScript, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The problem pi-autoresearch solves, and who it is for

Optimization work has a boring shape. You change one thing, run the benchmark, look at the number, and decide whether to keep the change. Doing that by hand is slow, and doing it in a chat session is worse: the agent forgets which variants already failed, and you end up re-testing dead ends. pi-autoresearch packages that loop as an extension for pi, the terminal coding agent, so the agent can run the cycle itself and keep a record.

The README frames the loop in one line: try an idea, measure it, keep what works, discard what does not, repeat. The stated targets are test speed, bundle size, LLM training and build times, with Lighthouse scores named as another example. The common property is that each target resolves to a command whose output includes a number.

That constrains the audience. This is not a tool for refactoring for readability or for fixing a bug with a fuzzy definition of done. It is for someone who can write a script that prints a measurement, is willing to let an agent edit the files in scope, and will review the resulting commits. The project credits karpathy/autoresearch as its inspiration, and the pi extension is the packaging of that idea into a terminal agent with a dashboard and session files.

How the loop actually runs: tools, measure script and log

Three extension tools form the mechanism. `init_experiment` configures the session once, with a name, a metric, a unit and a direction. `run_experiment` executes any command, times wall-clock duration and captures the output. `log_experiment` records the result, auto-commits, and updates both the widget and the dashboard. The split matters: measurement and recording are separate steps, so the agent has to commit to a judgement after seeing the number.

Session state lives in a single `.auto/` folder at the working-directory root. The README lists `.auto/prompt.md` as the session document holding the objective, metrics, files in scope and what has been tried, and says a fresh agent can resume from that file alone. `.auto/measure.sh` is the benchmark script: it runs pre-checks, executes the workload, and outputs `METRIC name=number` lines. `.auto/log.jsonl` is an append-only log of experiments.

One design detail is worth calling out. The core loop has no hook awareness, according to the README. Hooks are an optional skill, `autoresearch-hooks`, that helps author `.auto/hooks/before.sh` and `.auto/hooks/after.sh`. The skill ships with ten reference scripts in `skills/autoresearch-hooks/examples/`, covering external search, a learnings journal, native notifications, anti-thrash and idea rotation. So the extension stays small and the opinionated extras are opt-in.

The UI reports a confidence score after three or more runs, comparing the best improvement to the session noise floor: at or above 2.0x is shown green as likely real, between 1.0x and 2.0x is yellow as above noise but marginal, and below 1.0x is red as within noise. That is the project's own answer to the obvious failure mode of optimization loops, which is chasing noise.

Installing pi-autoresearch and running a first session

The README gives two install paths. The persistent one installs the package into pi, then starts the agent. The second loads the package for a single session only.

bash
pi install npm:pi-autoresearch
pi
bash
pi -e npm:pi-autoresearch

Inside pi, the loop starts with the `/autoresearch` command followed by a plain-language objective. The README's first example asks the agent to optimize unit test runtime while monitoring correctness.

text
/autoresearch optimize unit test runtime, monitor correctness

A second documented example targets training: run five minutes of train.py and note the loss ratio as the optimization target. The `autoresearch-create` skill asks a few questions about goal, command, metric and files in scope, or infers them from context, then writes the session files and starts the loop immediately. Expect `.auto/prompt.md` and `.auto/measure.sh` to appear at the working-directory root.

To watch progress, `/autoresearch export` opens a live dashboard in the browser that auto-updates as experiments run, and `/autoresearch dashboard` opens a fullscreen scrollable overlay in the terminal. The overlay is navigated with arrow keys or `j`/`k`, `PageUp`/`PageDown` or `u`/`d`, `g`/`G` for top and bottom, and `Escape` or `q` to close. `/autoresearch off` leaves autoresearch mode while keeping `.auto/log.jsonl` intact, and `/autoresearch clear` deletes the log, resets state and turns the mode off for a clean start.

Nothing is bound by default, and that is deliberate

The extension ships with no keyboard shortcuts. The README explains why: pi's built-in keymap grows with every release, so any default chord eventually collides, and it cites a past conflict between `ctrl+shift+f` and pi's transcript search. Every action is reachable as a `/autoresearch` subcommand instead, so the default install cannot hijack a built-in key.

Opting in means writing chords into `<agent-dir>/extensions/pi-autoresearch.json`, where `<agent-dir>` is usually `~/.pi/agent` or `PI_CODING_AGENT_DIR` when set. Omitted or null keys stay unbound.

json
{
  "shortcuts": {
    "fullscreenDashboard": "ctrl+shift+y",
    "export": "alt+shift+e",
    "off": null
  }
}

Each key maps to its subcommand: `fullscreenDashboard` to `dashboard`, `export` to `export`, `off` to `off`. The README warns that extension shortcuts win conflicts, so a clashing chord hijacks the built-in action rather than failing loudly. It also gives a verification script that imports `KEYBINDINGS` from the installed pi and reports whether a chord is already taken, noting that `ctrl+shift+y`, `ctrl+shift+u`, `alt+shift+f` and `ctrl+alt+d` were free as of pi 0.84.x, with the caveat to re-verify because the keymap grows.

This is a reasonable trade. The cost is that new users get no shortcut muscle memory, and the benefit is that upgrading pi will not silently break a chord the extension claimed.

Where the loop breaks: measurement, scope and reverts

The extension has no way to know whether your metric means anything. `.auto/measure.sh` is written by the agent from your description, and if it measures the wrong thing, the loop will happily optimize the wrong thing and commit the result. The confidence score compares the best improvement to the session noise floor after three runs, which catches jitter, but it cannot catch a metric that is stable and irrelevant. A test suite that passes because a test was weakened, or a bundle that shrinks because a feature was dropped, will look like progress.

The second boundary is scope. The loop edits files and auto-commits through `log_experiment`, and reverts regressions. The README points out that the `.auto/` folder is the one thing to preserve across reverts, gitignore and clean up, which implies the loop is operating on a branch where reverts happen. If your working tree has uncommitted work you care about, this is the wrong tool until that work is committed or stashed elsewhere.

The third is that the loop is only as good as the ideas the agent generates. The optional `autoresearch-hooks` skill includes an anti-thrash script and an idea rotation script among its ten examples, which is an admission that a naive loop can repeat failed variants or stall. Those are examples to adapt, not defaults: the README states the core loop has no hook awareness.

Finally, `autoresearch-finalize` exists because the loop produces a noisy branch. It splits the work into independent branches, one per logical change, each starting from the merge-base, with the constraint that groups must not share files. If your optimization touches one file repeatedly, that constraint is a real obstacle to a clean review.

How it differs from running autoresearch by hand or in another agent

The closest reference point is karpathy/autoresearch, which the README names as the inspiration. The difference is packaging. The original is a pattern you wire up yourself; pi-autoresearch is an installable pi package with three named tools, a session folder, a live dashboard and a confidence score. That means less setup and more opinions about where state lives.

Against a general coding agent with no extension, the difference is the record. A chat session leaves no structured history of which variants were tried and what each measured. Here, `.auto/prompt.md` is written so a fresh agent can resume from it alone, and `.auto/log.jsonl` is append-only. That is the part that makes long runs practical.

Against a dedicated experiment tracker such as a hyperparameter sweep tool, the difference is the edit step. A sweep tool varies parameters you defined in advance; pi-autoresearch lets the agent change the code between runs. That is more open-ended and correspondingly harder to bound. If your search space is already expressible as a parameter grid, a conventional sweep is more predictable and easier to reproduce. The extension earns its place when the change itself is the variable.

Maintenance, licence and what you take on

The package is MIT licensed, with a `THIRD_PARTY_NOTICES.md` at the repository root, which is the file to read before redistributing. The repository is not archived, and the last push was on 2026-09-10, with v1.8.1 released on 2026-09-08. The changelog is a top-level file, so upgrade notes are in the repository rather than only in release pages.

Runtime constraints are explicit in `package.json`: `engines.node` is `>=22`, the package is ESM (`"type": "module"`), and the declared package manager is pnpm 10.28.2. The pi packages it builds against, `@earendil-works/pi-ai`, `@earendil-works/pi-coding-agent` and `@earendil-works/pi-tui`, are peer dependencies with `*` ranges and dev dependencies pinned at `^0.74.0`, so the extension tracks pi's development line rather than a fixed release. Upgrading pi is therefore the main compatibility risk, and the README's own note about keymap growth is an example of that risk already materialising once.

The upgrade cost for a user is mostly re-verifying the measure script and any hooks after a pi upgrade, plus re-checking shortcut chords if you opted in. The upgrade cost for a fork is higher: the extension depends on pi internals such as the keybindings module path used in the verification snippet, which is not a stable public surface. Nothing here is legal advice; if you redistribute the package, read the LICENSE and THIRD_PARTY_NOTICES.md yourself.

Editorial conclusion

Adopt pi-autoresearch if you already run pi and can express your target as a command that prints METRIC name=number lines, because the loop is only as trustworthy as that measurement. Skip it if your goal has no automatable metric, or if you cannot review the commits the loop writes. Before starting a session, verify Node 22 or newer, run your measure script by hand, and confirm .auto/ is gitignored, since the loop commits and reverts against your working tree.

Frequently asked questions

What is pi-autoresearch?

It is an extension for pi, a terminal AI coding agent, that runs autonomous optimization loops: try an idea, benchmark it, keep improvements, revert regressions, repeat. It adds three tools, a dashboard widget and an `/autoresearch` command, and it credits karpathy/autoresearch as its inspiration.

How do I install the pi-autoresearch extension?

The README gives two commands: `pi install npm:pi-autoresearch` followed by `pi` for a persistent install, or `pi -e npm:pi-autoresearch` to load the package for one session only. The package requires Node 22 or newer.

Can pi-autoresearch optimize anything, or only code speed?

The README names test speed, bundle size, LLM training, build times and Lighthouse scores as targets, and states it works for any optimization target. The practical requirement is that the target resolves to a command whose `.auto/measure.sh` script outputs `METRIC name=number` lines.

Where does pi-autoresearch store its session data?

All session files live in a single `.auto/` folder at the working-directory root: `.auto/prompt.md` for the session document, `.auto/measure.sh` for the benchmark script, and `.auto/log.jsonl` for the append-only experiment log. Legacy flat `autoresearch.*` files are still read for in-flight sessions.

How does pi-autoresearch decide whether an improvement is real?

After three or more runs it shows a confidence score comparing the best improvement to the session noise floor. At or above 2.0x it is green as likely real, between 1.0x and 2.0x yellow as above noise but marginal, and below 1.0x red as within noise.

Official sources

  1. davebcn87/pi-autoresearch on GitHub
  2. Issues
  3. License: MIT
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/davebcn87-pi-autoresearch.svg)](https://hysenlabs.com/projects/davebcn87-pi-autoresearch)