# webcmd: a self-learning browser layer for AI agents

> webcmd is a TypeScript CLI that gives agents a live Playwright-style browser plus a local memory of the sites they have already visited. It is Apache-2.0, installs from npm, and claims up to 90% lower browser-agent token spend.

**agentrhq/webcmd** — Self-learning agent browser

- Repository: https://github.com/agentrhq/webcmd
- Website: https://webcmd.dev
- Stars: 2,623 · Forks: 1,539
- Language: TypeScript
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/agentrhq-webcmd

## The problem webcmd targets: agents paying to rediscover the same sites

Every browser agent run against a known site repeats the same reconnaissance. The agent loads the page, reads the accessibility tree, guesses which control does what, clicks something wrong, and tries again. On a site you visit daily, that reconnaissance is pure overhead, and it is paid in tokens and wall-clock time.

webcmd's README puts the goal plainly: "stop making agents rediscover the same sites on every run and cut browser-agent token spend by up to 90%." The project describes itself as "self-learning browser infra for AI agents", and the description in package.json is broader still: turn websites, browser sessions, desktop apps, and local tools into deterministic CLI surfaces for humans and AI agents.

The audience is narrow and specific. It is engineers running agent harnesses (the README names Claude and Codex as supported harnesses) who need an agent to operate a real, often logged-in browser and who are tired of paying for the same exploration twice. It is not a scraping library and not a headless testing framework, even though the underlying control surface looks like Playwright.

## Two layers: live browser control and a sitemap memory

The README splits the system into two layers. Layer 0 is live browser control: when the site is unfamiliar, the agent uses webcmd browser to inspect, click, type, extract, capture network calls, and finish the task in a real browser. Layer 1 is sitemap memory: when the site is familiar but the action space is not fully mapped, webcmd captures an "agent-facing sitemap" of observed pages, states, actions, workflows, APIs, pitfalls, and fallback paths.

That second layer is the interesting one. The memory is not a cache of page HTML; it is a structured record of what was observed, including the things that went wrong (pitfalls) and the paths that worked when the primary one did not (fallback paths). The README is explicit that learning is conservative: "the live browser is always truth, Webcmd never explores just to learn, and a memory failure never blocks the task." First access may use a Webcmd Cloud seed; subsequent learning stays local.

That last sentence is the design constraint worth pausing on. A cloud seed means the first visit to a site can start from knowledge someone else gathered, but the project claims learning after that point stays on the local machine. The README does not describe how the seed is produced, how it is shipped, or what happens if you want no cloud contact at all, so treat the seed as an undocumented dependency until you check the docs site.

## Profiles, sessions, and why session IDs are immutable

The concurrency model is the part most likely to bite you, so it is worth reading the README's own words. Profiles are cookie jars. Sessions are independent browser windows within a profile. Session IDs are immutable, Profile-scoped, and safe to reuse for that Session's lifetime. Parallel agents should create separate Sessions. Raw browser commands require an explicit readable Session ID.

Each of those sentences has an operational consequence. Reusing one session across parallel agents is not supported by design; you create a session per agent. The ID is scoped to a profile, so the same ID string under two profiles is not the same window. And the CLI refuses raw browser commands without a session, which prevents the common failure mode of an agent driving whatever window happens to be open.

The README also shows that agents can send a single sandboxed Playwright-style program to a named session instead of issuing many small commands. For multi-step local exploration, that is fewer round trips and less back-and-forth in the agent transcript.

## Installing webcmd and running your first session

webcmd requires Node.js 20.6+. The README gives the global npm install and then a skills step:

```bash
npm install -g @agentrhq/webcmd
webcmd skills add
```

When prompted, choose Claude, Codex, another supported harness, or a custom skills path. According to the README, that installs exactly one skill, webcmd-browser. Load or tag that skill only for live browser work; the README states that installation and setup commands do not require it.

There is also an agent-driven path, in which you hand the setup prompt to your harness:

```text
Fetch and follow https://raw.githubusercontent.com/agentrhq/webcmd/main/start.md to set up Webcmd end to end.
```

For a first real session, create a profile and a session, then list the open tabs. The README's example uses the profile name work and the returned session ID work-project-k7:

```bash
webcmd --profile work session create "Work Project" -f json
# id: work-project-k7
webcmd --profile work --session work-project-k7 browser tabs
```

To run a multi-step program, write a Playwright-style script and pass it with --file, or pipe a single snippet through --stdin. The README shows both forms:

```bash
webcmd --profile work --session work-project-k7 browser run --file explore.js
printf 'return await page.title();' \
  | webcmd --profile work --session work-project-k7 browser run --stdin
webcmd --profile work session close work-project-k7
```

What you should see: session create with -f json prints an id, browser tabs lists the windows in that session, and the run commands return whatever the script evaluates. Close the session when the task is done, as the last line does. If you skip the session flag on a raw browser command, expect the CLI to reject it rather than pick a window for you.

## Where webcmd is the wrong tool

The self-learning layer only pays off on sites you revisit. On a one-off scrape of a site you will never touch again, you pay the live-browser cost anyway and the memory has nothing to amortize against. The README's own framing supports this: learning happens as agents use sites, and webcmd never explores just to learn.

The second limitation is the shape of the memory. The README lists what a sitemap captures, but it does not document how a stale entry is invalidated when a site changes its layout, nor how you inspect or edit what was learned. Sites that ship UI changes weekly are the worst fit for any memory layer, and the README does not say what happens when a remembered path no longer exists. It does say a memory failure never blocks the task, which is a safety property rather than a freshness guarantee.

The third is scope. This is a CLI plus a skill for agent harnesses, not a hosted browser service. If you need a fleet of browsers behind an HTTP endpoint, or per-tenant isolation with quotas, webcmd's model of local profiles and sessions is the wrong abstraction. The README also does not document rollback or downgrade behaviour for the CLI itself, so version pinning is your problem.

## How webcmd differs from browser-use

The README positions webcmd against browser-use directly, using the BU Bench V1 benchmark from browser-use's own repository. The claim is that on this 100-task browser automation benchmark, webcmd recorded the highest accuracy and lowest estimated controller cost per completed task, and fewest agent turns per completed task in the comparison, with 67% accuracy, $0.255 cost per completed task, and 9.8 agent turns per completed task.

The methodological note matters more than the numbers. All tools in that comparison used the same Pi controller, controller model, a Codex gpt-5.4 judge, and the CloakBrowser engine, and the README flags that this is a stronger judge than the original BU Bench setup, whose current runner uses Gemini 2.5 Flash. Holding the controller and judge constant is the right instinct, but it means the comparison measures the browser layer, not the whole agent stack.

The architectural difference is the memory. browser-use is a library that drives a browser for a task; webcmd adds a persistent local record of what was observed on each site and reuses it on later runs. Whether that matters depends entirely on whether your workload revisits sites. For a benchmark of 100 distinct tasks, the memory layer has little to amortize; for a daily workflow against ten known sites, it is the whole point.

## Licence, maintenance and upgrade cost

webcmd is Apache-2.0, with a NOTICE file in the repository root, which is the standard Apache arrangement: you may use, modify and redistribute it, including commercially, provided you keep the licence and notice. Apache-2.0 also carries an explicit patent grant, which matters if you embed this in a product. This is a description of the licence text, not legal advice; if you are redistributing a modified webcmd, read LICENSE and NOTICE yourself.

The repository is not archived, and the last push was on 2026-09-10. Releases are frequent and small: webcmd-v0.8.4 on 2026-09-09, webcmd-v0.8.3 on 2026-09-08, webcmd-v0.8.2 on 2026-09-07. The presence of release-please-config.json and .release-please-manifest.json indicates automated release tooling, which fits that cadence.

Frequent patch releases have a cost. A global npm install means every engineer's webcmd version can drift independently, and the README does not document a compatibility policy between CLI versions and installed skills. The skills themselves are generated: the Makefile builds skill-src/cli into skills/ and skill-src/mcp into mcp-skills/, and it warns that source files are never named SKILL.md because installers match that literal filename. If you fork the skills, you inherit that build pipeline, which uses litprompt and is separate from the TypeScript build that still runs through npm.

## Conclusion

Adopt webcmd if you run agents against a small set of sites repeatedly and want the navigation knowledge to stay on your machine: profiles are cookie jars, sessions are isolated windows, and raw browser commands require an explicit session ID. Do not adopt it if you want a hosted, multi-tenant browser service or a tool that explores on its own; the README states webcmd never explores just to learn, so an unfamiliar site still costs live browser turns. Before rolling it out, verify two things yourself: that your harness has a place for the single webcmd-browser skill, and that your Node version satisfies the engines field of >=20.6.0.

## FAQ

### How do I install webcmd?

Install Node.js 20.6 or newer, then run npm install -g @agentrhq/webcmd followed by webcmd skills add. The skills step prompts you to choose Claude, Codex, another supported harness, or a custom skills path, and installs exactly one skill, webcmd-browser.

### What is the difference between a profile and a session in webcmd?

The README describes profiles as cookie jars and sessions as independent browser windows within a profile. Session IDs are immutable and profile-scoped, and the README advises parallel agents to create separate sessions.

### Does webcmd work with Claude and Codex?

Yes. The README lists Claude and Codex among the supported harnesses you can pick during webcmd skills add, and the repository contains .claude-plugin/ and .codex-plugin/ directories. Load the webcmd-browser skill only for live browser work.

## Sources

- [agentrhq/webcmd on GitHub](https://github.com/agentrhq/webcmd)
- [License: Apache-2.0](https://github.com/agentrhq/webcmd/blob/main/LICENSE)
- [Project website](https://webcmd.dev)
- [README](https://github.com/agentrhq/webcmd/blob/main/README.md)
- [Releases](https://github.com/agentrhq/webcmd/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/agentrhq-webcmd
