# autoprompt-skill ships version 2.0.0 with version 1 benchmarks and an unmeasured cost estimate

> Autoprompt is a CLI wrapper that routes a stated goal to one of eleven coding agents, with bounded delegation, a concurrency dial and evidence-backed checks. The install path and the provider table are unusually specific, and the benchmark section is unusually candid about its own limits, which turns out to be the most useful thing in it. The gaps are structural: the measured benefit has one run behind it, the stated cost has none, and the two sections that explain how the thing works contain nothing but images.

**Spielewoy/autoprompt-skill** — Autoprompt is a coding-agent skill that cuts failures by 45% on agentic coding tasks.

- Repository: https://github.com/Spielewoy/autoprompt-skill
- Website: https://www.npmjs.com/package/autoprompt-skill
- Stars: 1,297 · Forks: 88
- Language: JavaScript
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/spielewoy-autoprompt-skill

## The release is 2.0.0 and the benchmarks in it are labelled version 1

The benchmarks section opens by saying these are version 1 benchmarks and that version 2 benchmarks will follow. The manifest is at `2.0.0`, released as v2.0.0 on 2026-09-09, after v1.0.4 on 2026-08-21 and v1.0.3 on 2026-08-20.

So the headline number on the front page belongs to the previous major line, and the project says so in the same sentence as the claim. That is worth crediting, because the alternative is a 2.0 launch page quietly recycling 1.x evidence.

The cost side is where the honesty runs out. The expected trade-off is about 3x the time and 2x the tokens, and the sentence explaining where that comes from says timing and token logs were not retained, so these are planning estimates based on user experience reports and not measured benchmark results. There is a further note that for very small tasks this may differ significantly.

Put plainly: the benefit was measured, the price was not, and the README refuses to blur the two. Budget the cost yourself before turning the mode on for routine work.

## The measured table is internally consistent, and the adjacent number is not comparable

The one measured comparison is hidden inside a collapsed details block:

| Track | Solved | Score | Failed |
|---|---:|---:|---:|
| OpenCode | 60/89 | 67.42% | 29 |
| OpenCode + Autoprompt | 73/89 | 82.02% | 16 |
| Change | +13 solves | +14.61 points | 45% fewer |

The arithmetic holds together. Sixty of eighty-nine is 67.42%, seventy-three of eighty-nine is 82.02%, the difference is thirteen solves and 14.60 points, and thirteen fewer failures out of twenty-nine is 44.8%, which rounds to the 45% in the headline. Nothing has been massaged.

What sits next to it is a third figure. DeepSeek's 82.7% is mentioned, and the project states immediately that it used its own test setup and is therefore a reference point rather than a comparable third run. That is the right call, and it also means the reader sees 82.7% above 82.02% and has to hold the caveat in mind.

The scope is narrow enough to name: one provider, one run, eighty-nine tasks, Linux only, and no spread or repeat count anywhere in the table.

## `--concurrency wide` is the only concurrency setting without a ceiling

There are three concurrency modes. `--concurrency tokensaver` runs at most six subagents at once, which is a real cap. `--concurrency custom --max-subs N` sets your own limit. `--concurrency wide` starts ready, independent work up to the host limit.

That third one is the default-shaped choice and the one with no number in it. Whether wide spawns four agents or forty depends entirely on what the machine and the provider tolerate, and nothing in the tool caps it.

The model controls have a similar shape. `configure PROVIDER --agents off` falls back to the provider's configured model. `configure PROVIDER --agents MODEL` picks one, with `--effort LEVEL` added where supported, so the flag exists and its availability varies by agent. `configure PROVIDER --agents auto --model-map FILE` chooses from a measured model registry, and the note says a comma-separated model list also requires `--model-map`, so the registry file is not optional on that path.

One thing the visible text does not say is who measured the registry. Everything is passed after `--` and before the quoted goal:

```bash
autoprompt activate codex -- --concurrency custom --max-subs 4 "add retries and tests"
```

## The documented install pulls a tarball from a GitHub release, not from npm

The install command is a URL:

```bash
npm install -g https://github.com/Spielewoy/autoprompt-skill/releases/download/v2.0.0/autoprompt-skill-2.0.0.tgz
```

Meanwhile the manifest sets `publishConfig` to `access: public` with `registry: https://registry.npmjs.org/`, and the project's homepage field is the npm package page. So the package is published to npm and the README still tells you to install a release tarball from GitHub. Both channels carry the same name.

There is a third path too, phrased as downloading an installer from GitHub Releases. That makes three possible artefacts for one CLI: the registry tarball, the release tarball, and whatever the release installers are.

The tarball URL is the one line in the whole install section that pins anything, because the version is in the path. That is arguably the reason to prefer it over `npm install -g autoprompt-skill`, and it is also the line most likely to rot, since it has to be edited with every release.

From source it is `git clone`, `cd autoprompt-skill`, `npm install -g .`, then `autoprompt` to launch the installer and choose your agent.

## Bash 4.3 and Python 3.11 with PyYAML, wrapped around one CommonJS entry script

The requirements are Node.js 20 or newer, Python 3.11 or newer available as either `python3` or `python` with PyYAML installed, and Bash 4.3 or newer on macOS or Linux. Git is needed only for the GitHub checkout method.

The package itself is `"type": "commonjs"` with a single bin target, `bin/autoprompt.cjs`, registered under both the names `autoprompt` and `autoprompt-skill`. So a Node CLI is the surface, and Python plus a YAML parser is a runtime requirement underneath it. PyYAML being called out by name rather than left to a transitive dependency is the kind of detail that saves an afternoon.

The platform line is narrower than the payload. Bash is required on macOS or Linux, yet the published files include `scripts/wsl-runtime.cjs` and `scripts/wsl-runtime.ps1`, a PowerShell script, alongside `darwin-runtime-setup.cjs` and five `lima-runtime` files. So there is a Windows path through WSL and a virtual machine path through Lima in the shipped scripts, and the requirements list does not mention either.

If you are on Windows, the requirements section is not the place that will tell you.

## Eleven agents are all marked Working, and one of them is pinned to a release candidate

The support table has one row per agent and every status reads Working. The keys and the versions they were verified against are `claude` at 2.1.263, `codex` at 0.148.0, `opencode` at 1.18.29, `kilo` at 7.5.15, `vscode` at 1.136.1, `prime` at 0.7.2, `omp` at 18.1.14, `deepseek` at 0.1.2-rc.1, `reasonix` at 1.30.0, `hermes` at 0.21.1 and `grok` at 1.0.13.

Two things stand out. There is no partial or untested row, so an all-green table carries the full weight of the claim. And DeepSeek Harness is pinned to `0.1.2-rc.1`, a release candidate, inside a table where everything else is a final version.

The qualifier under the table is that these versions passed Linux runs, and that model and platform availability varies by provider. So every one of the eleven statuses means verified on one operating system.

There is also a small mismatch in the run controls, which claim the same controls apply to all eleven providers. The table underneath mixes `activate` flags such as `--concurrency` with `configure` subcommands such as `--agents`, so it is describing two commands rather than one control surface.

## The published file list names about two dozen scripts individually instead of shipping the directory

The `files` array in the manifest opens with `bin/` as a directory glob and then switches to naming individual paths: `scripts/codex-configure.cjs`, `scripts/darwin-runtime-setup.cjs`, `scripts/provider-closure.cjs`, `scripts/hermes-runtime-closure.cjs`, five `lima-runtime` files, `scripts/wsl-runtime.cjs`, `scripts/wsl-runtime.ps1`, `scripts/codex-runtime-identity.cjs`, `scripts/local-only-safety.cjs`, `scripts/harness-provider-config.cjs`, a `scripts/harness-v2-*.cjs` pattern, plus the `harness-v2-trust/`, `harness-v2-bridge/` and `install/` directories and `scripts/runtime-payload.cjs`.

A script added later does not ship until somebody edits this list. That is a deliberate-looking trade, and it does keep `scripts/` from dragging test helpers into the tarball, but it makes the manifest a second place where the file inventory has to be maintained.

The `agents/` entries are per provider, and the visible part enumerates nine of them: `agents/claude/`, `agents/codex/`, `agents/opencode/`, `agents/kilo/`, `agents/vscode/`, `agents/prime/`, `agents/omp/`, `agents/deepseek/` and `agents/reasonix/`, with a matching `agents/manifests/<provider>-runtime.json` for the same nine. The eleven supported keys include `hermes` and `grok`, and neither appears in the visible entries. The listing does cut off inside `agents/manifests/`, so that may not be the whole story, but it is the whole story you can see.

## How it works is two SVG files, and none of the guides ship inside the package

Two of the README's own headings have no prose in them. How it works contains a single link to `assets/how-it-works-loop.svg`, and The agents contains a single link to `assets/how-it-works-hierarchy.svg`. So the central mechanism of the tool, the review and recheck loop and the agent hierarchy, is documented only as images.

The linked guides are outside the tarball as well. The `files` array covers `bin/`, the named `scripts/` paths and `agents/`. Everything under `docs/` is absent, which includes the custom agent compatibility guide, the support and audit notes, the work paths FAQ, the custom model setup guide, and the benchmark setup and evidence boundaries that the benchmarks section itself points you to.

The README also carries five language versions, with Chinese, Korean, Spanish and Arabic under `docs/translations/`, none of which ship either.

So the practical shape is: an installed CLI with no local copy of the compatibility matrix, the audit notes, or the caveats on the benchmark. Keep the repository open next to the terminal while you evaluate an agent you have not used before.

## Conclusion

Autoprompt is worth a try if you want a single entry point across several coding CLIs and you want the delegation capped rather than left to the agent. Four things to settle first. The 45% figure is one run of 89 tasks against one provider on Linux, and the README says outright that the accompanying 3x time and 2x token costs are planning estimates rather than measurements, so the benefit is measured and the price is not. The support table marks all eleven agents as Working with no partial row, and one of them is pinned to a release candidate. The default concurrency mode has no ceiling. And the compatibility guide, the audit notes and the benchmark evidence boundaries all live in `docs/`, which the published package does not include, so keep the repository open while you evaluate.

## FAQ

### How do I install the autoprompt CLI?

Run `npm install -g https://github.com/Spielewoy/autoprompt-skill/releases/download/v2.0.0/autoprompt-skill-2.0.0.tgz`, then launch the installer by running `autoprompt` and choosing your coding agent. You can also download an installer from GitHub Releases, or install from source with `git clone https://github.com/Spielewoy/autoprompt-skill`, `cd autoprompt-skill`, `npm install -g .` and `autoprompt`.

### What does the autoprompt 45% failure reduction benchmark actually show?

One measured run of 89 tasks on OpenCode: 60 solved and 29 failed at 67.42%, against 73 solved and 16 failed at 82.02% with Autoprompt, a difference of 13 solves, 14.61 points and 45% fewer failures. The section states these are version 1 benchmarks with version 2 to follow, and notes that DeepSeek's 82.7% used its own test setup so it is a reference point rather than a comparable run.

### What are the requirements to run autoprompt?

Node.js 20 or newer, Python 3.11 or newer available as `python3` or `python` with PyYAML, and Bash 4.3 or newer on macOS or Linux. Git is required only for the GitHub checkout method. Note that the published scripts also include a WSL runtime with a PowerShell file and a Lima virtual machine runtime, neither of which the requirements list mentions.

### Which coding agents does autoprompt support?

Eleven, each with a key you pass to commands: claude, codex, opencode, kilo, vscode, prime, omp, deepseek, reasonix, hermes and grok. Every row in the support table is marked Working against a specific tested version, and the table notes that those versions passed Linux runs with model and platform availability varying by provider. DeepSeek Harness is listed at 0.1.2-rc.1, a release candidate.

### How do I control how many subagents autoprompt runs at once?

Pass a concurrency mode after `--` and before the quoted goal. `--concurrency tokensaver` runs at most six subagents at once, `--concurrency custom --max-subs N` sets your own limit, and `--concurrency wide` starts ready independent work up to the host limit, with no ceiling of its own. Model selection is separate, through `configure PROVIDER --agents MODEL` and `configure PROVIDER --agents auto --model-map FILE`.

## Sources

- [License: MIT](https://github.com/Spielewoy/autoprompt-skill/blob/main/LICENSE)
- [Project website](https://www.npmjs.com/package/autoprompt-skill)
- [README](https://github.com/Spielewoy/autoprompt-skill/blob/main/README.md)
- [Releases](https://github.com/Spielewoy/autoprompt-skill/releases)
- [Spielewoy/autoprompt-skill on GitHub](https://github.com/Spielewoy/autoprompt-skill)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/spielewoy-autoprompt-skill
