agent-qa retries a broken step in the same run and pins eight transitive packages
Open-source self-improving QA agent for software teams. A test harness with memory. Write tests in natural language for web and mobile. agent-qa learns from every run, adapts to UI changes, and catches regressions before you ship.
At a glance
- What is it?
- A test harness that takes YAML tests written in natural language, re-observes the interface when a sub-action fails and tries a different path, then reuses what worked. It is a pnpm and turbo monorepo with Node 24 as its floor, seven custom release scripts, and a curated list of exact pins on dependencies it does not own.
- Who is it for?
- agent-qa is a reasonable bet if your failure mode is selectors drifting out of date. The self-healing behaviour is specific rather than general: when a click, fill or select fails, the harness re-observes the interface and tries a different path inside the same run rather than failing on the first broken action.
- Can I use it commercially?
- Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
- Is it still maintained?
- Yes. The repository last received commits 62 days ago.
- What is it written in?
- Mainly TypeScript, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 4, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The installed package and the repository root are different names
The install command in the readme is short:
npm install -D agent-qaThe package you install is called `agent-qa`, unscoped. The manifest at the root of the repository is a private workspace root called `agent-qa-monorepo`, marked `"private": true`, with the actual published packages living under `packages/`.
There is a second package as well, installed only when you want to authenticate with a Codex or Claude Code subscription:
npm install -D @vostride/agent-qa-subscription-authSo there are three scopes in play. The public CLI is unscoped `agent-qa`, the subscription auth helper is scoped to the vendor, and the internal namespace that the workspace guard protects is `@agent-qa/`. The result is that a lockfile entry for this project can appear under two different scopes depending on which package pulled it in.
The monorepo is wired with pnpm workspaces and turbo. Build, test, typecheck, lint, dev and clean are all single turbo invocations, and the workspace pins its package manager at pnpm 10.6.1.
Docker is required for hooks, and hooks are how state gets set up
The feature list describes hooks as running Node, Bun, Python or Bash inside isolated Docker containers, and lists what they are for: setting up environments, calling APIs, seeding fixtures, tearing down state, and passing structured outputs back into the active test run.
So the harness has two runtimes. The agent drives a browser or a device, and the hooks run arbitrary code you wrote, in a container. The readme says plainly that Docker is required for the Node, Bun, Python and Bash hook containers, and that Docker has to be installed before using hooks.
That is a real dependency on a daemon for what is otherwise a test runner, and it is the part most likely to matter on a locked-down CI image. There is no stated fallback for environments where a Docker socket is not available, and the isolation claim is the reason the requirement exists.
The setup commands are three, and the third is conditional:
npx agent-qa init
npx agent-qa install-browsers --chromium
# Mobile projects:
npx agent-qa install-mobile-drivers --allBrowser installation is scoped to Chromium in the documented command, and mobile drivers are all-or-nothing with `--all`. Starting the dashboard is `npx agent-qa dashboard --open`, and tests are run from the CLI against a YAML file, with the readme's example being `tests/hacker-news-top-story.yaml`.
A failed step is retried in the same run, not in the next one
The self-healing claim has a specific trigger and a specific scope, and both are worth being precise about.
The trigger is a failing sub-action. The examples given are click, fill and select, which are the three interactions that break when a label or a role moves. The response is that the harness re-observes the interface and tries a different path in the same run. The stated consequence is that tests recover from UI drift and flaky interactions instead of failing on the first broken action.
Recovery inside the run rather than after it is what makes this different from a selector library that fails once and is then fixed by hand. It also means a run that passes may have taken a path you did not write down, which is the thing to look at when a test goes green for the wrong reason.
The memory side is described in two parts. Every run builds execution memory from product, suite and test observations, and that context is added to future runs. Separately, memory is curated from steps that were healed during execution, so a recovery becomes something later runs avoid rather than repeat. Both halves are described as living in version-controlled code, alongside tests, configs, hooks and suite logic, which means the accumulated knowledge is diffable rather than opaque.
Eight transitive dependencies are pinned to exact versions
The workspace manifest carries a `pnpm.overrides` block with eight entries, and every one is an exact version rather than a range.
`@hono/node-server` at 1.19.14, `basic-ftp` at 5.3.1, `brace-expansion` at 5.0.5, `hono` at 4.12.16, `lodash` at 4.18.1, `monaco-editor>dompurify` at 3.4.2, `path-to-regexp` at 8.4.2 and `webdriver>undici` at 6.25.0.
Three of those names, `brace-expansion` and `path-to-regexp` among them, are the sort of package that appears in advisories often enough that an override list is the normal way to respond. Two of the entries are written as scoped overrides, pinning a transitive dependency of a specific package rather than the package globally.
The trade-off is that this list is now the project's security update mechanism. A patched version of any of these eight does not arrive by widening a range, because there is no range; it arrives by someone editing this block, which means every transitive security update is a deliberate diff in a file that also has nothing else in it.
There is a companion list for native builds. `onlyBuiltDependencies` names exactly three: `better-sqlite3`, `esbuild` and `sharp`. Everything else that needs a compiler is blocked from running install scripts by the package manager.
The namespace guard is a grep with a shell negation in front of it
One script in the root manifest is worth reading in full rather than skimming.
It is `lint:namespace`, and it runs a recursive grep over `packages/` for the string `@agent-qa/`, restricted to TypeScript, JSON and JavaScript files and excluding `node_modules` and `dist`. What makes it a check rather than a search is the leading exclamation mark, which in a POSIX shell negates the exit status.
So the grep succeeds when it finds nothing, the negation makes that into a failure, and the script exits non-zero when a match is present. The intended effect is that no file under `packages/` may import the internal `@agent-qa/` scope, which keeps the published packages from depending on each other through a namespace that consumers cannot resolve.
It is an unusual choice for a project with turbo, typescript and prettier already configured, and it has the usual fragility of a text match: it fires on the string appearing in a comment, a string literal or a documentation snippet, and it does not understand which imports are real.
The three packages named in `onlyBuiltDependencies` are a fair companion detail. Native compilation is opted in per package rather than allowed globally, which is the current default posture and also why those three names are worth knowing.
Node 24 is the floor and the readme does not say so
The workspace manifest sets `engines.node` to `>=24`, and there is a `.nvmrc` at the root of the repository.
None of that appears in the readme. The requirements the page does state are Docker for hooks, a browser or mobile driver installed by the tool, and a model endpoint you bring yourself. The Node version is left to the manifest.
Node 24 is a high floor for a tool people will try quickly. It means an install on an older long-term-support release fails or warns at the resolver rather than at runtime, which is the better failure, but it also means the harness cannot be dropped into a repository whose CI image is pinned to an earlier major without changing that image first.
The model requirement is the other one the readme leaves to the reader, and it is more open than a single endpoint. The feature list says you can bring your own model through OpenAI and Anthropic compatible endpoints, through Gemini, through local or open-source models, and through subscriptions like Codex and Claude Code, with the separate `@vostride/agent-qa-subscription-auth` package for the subscription case.
The practical shape is that there is no model dependency to install, but there is a choice to make, and the two subscription paths add a second package that only matters if you want to bill against a plan you already hold.
Seven release scripts and two validators for the agent-facing files
The scripts block is longer than the feature list suggests. Beyond the six turbo passthroughs there are ten more.
Three of them validate rather than build. `validate:skills` runs a script over the `skills/` directory, `validate:agents` runs one over the repository's agent instruction files, and `validate:publish` runs a script named for the publish surface. The first two are the interesting ones, because they mean the MCP and skills artefacts that the feature list advertises for coding agents are checked by the same repository that ships them, and the agent instruction file at the root has its own validator.
Seven of them handle releases, and none of them delegates to a single tool. There is a Docker release, a Docker release check that runs locally, a dry run, a verify step, a publish step, and a GitHub release step, each a separate `.mjs` file under `scripts/release/`.
That is a lot of machinery for a repository with no GitHub releases, which is worth noting as a fact rather than a criticism: the release pipeline is built and documented in code, and the published artifacts are the npm packages instead.
The repository root also carries `glama.json` for MCP directory listing and `knip.json` for unused-code detection, plus a `docker/` directory matching the Docker release scripts and a `demo-project/` directory that the readme links to as the demo.
Editorial conclusion
agent-qa is a reasonable bet if your failure mode is selectors drifting out of date. The self-healing behaviour is specific rather than general: when a click, fill or select fails, the harness re-observes the interface and tries a different path inside the same run rather than failing on the first broken action. Combined with an action cache that reuses validated plans across similar runs, that targets the real cost of browser testing, which is maintenance rather than authoring.
The parts that will shape your setup are the prerequisites. Node 24 or newer is the engine floor in the workspace manifest and is not stated in the readme. Docker is required for hooks, and hooks are how a test sets up an environment, calls an API, seeds fixtures or tears state down. Browsers and mobile drivers are installed by the tool itself with separate commands. So the harness needs more of a machine than a script would.
Two smaller decisions worth making deliberately. Self-improvement means memory accumulates from product, suite and test observations, including from steps that were healed, and that memory is described as version-controlled code, so what the harness learned about your application becomes something you review and diff. And the model is yours to choose, through OpenAI and Anthropic compatible endpoints, Gemini, local models, or a Codex or Claude Code subscription via a separate auth package.
The project ships a LICENSE.md and a NOTICE.md rather than a conventionally named licence file, and the repository platform reports no recognised licence, so check those two files before you vendor it.
Frequently asked questions
How do I install and start agent-qa?
Install the package with npm install -D agent-qa, then run npx agent-qa init, install a browser with npx agent-qa install-browsers --chromium, and start the dashboard with npx agent-qa dashboard --open. Mobile projects also need npx agent-qa install-mobile-drivers --all, and Docker is required before using hooks.
What format are agent-qa tests written in?
YAML files, run from the CLI with a command of the form npx agent-qa run followed by the path, with the readme's example being tests/hacker-news-top-story.yaml. Actions and assertions inside are written in natural language.
What does self-healing mean in agent-qa?
When a sub-action such as a click, fill or select fails, the harness re-observes the interface and tries a different path within the same run, so the test recovers instead of failing on the first broken action. Recoveries are also curated into memory for later runs.
Can I use my own model with agent-qa?
Yes. It supports OpenAI and Anthropic compatible endpoints, Gemini, local or open-source models, and Codex or Claude Code subscriptions. The subscription path needs the separate @vostride/agent-qa-subscription-auth package installed.
What Node version does agent-qa need?
Node 24 or newer, set as the engine floor in the workspace manifest, with a .nvmrc at the root. The readme does not state the version requirement. The repository root package is a private pnpm and turbo workspace called agent-qa-monorepo.
Why does agent-qa pin specific transitive dependency versions?
The workspace manifest carries eight exact pins in its overrides block, including brace-expansion and path-to-regexp, so a patched release of any of them has to be pushed through that list rather than arriving by widening a version range.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/vostride-agent-qa)