# Argus takes the Driver seat and splits it into four roles that are not allowed to grade each other

> Argus is a Python and TypeScript runtime for long-running autonomous research and engineering work that continues past a single model turn. Its central mechanism is a four role split, Manager, Planner, Engineer, Reviewer, with an explicit column for what each may not do, plus a hard rule that credentials, payment, irreversible actions, and publication always stop for a human. The surrounding packaging has its own complications, including a repository that is a preview of another one.

**lbx154/Argus** — A self-evolving multi-agent system for autonomous research, operating 24/7 to explore, learn, and improve.

- Repository: https://github.com/lbx154/Argus
- Website: argusbot.cn
- Stars: 309 · Forks: 34
- Language: Python
- License: MIT
- Published: 2026-09-17 · Updated: 2026-09-17 · Language: en
- Canonical page: https://hysenlabs.com/projects/lbx154-argus

## The repository you clone is a preview of another repository

The most important paragraph on the page is a callout, not a feature list. `microsoft/ArgusAgent` is named as the official source repository, and `lbx154/Argus` is described as the development preview. Changes reach the official repository through synchronization, and the page makes the distinction explicit: installing source from `main` is not the same as installing a published Desktop release.

Branch discipline follows from that. `main` is the stable source branch and `dev` is the development branch, development pull requests target `dev`, and only reviewed, validated changes are promoted to `main`. Anyone who wants to contribute is pointed at `CONTRIBUTING.md` for the process.

Packaging is described as two separate channels, source updates and packaged desktop previews, with the version banner naming v0.1.6.

There is also a Tauri desktop directory in the tree alongside the Python package, the TypeScript workspace, `argus_doctor.py`, an `argus_skill/` directory, and a `technical_report/` directory. So there are three deployable shapes in one repository, and the documentation is careful to say which one you are installing.

For an evaluator, that means the code you review here may not be the code that runs, which is worth knowing before drawing conclusions about it.

## Four roles carry an explicit list of what they may not do

The Driver and harness model is the conceptual frame. A model is an engine that burns compute and emits tokens, a harness is the drivetrain coupling those tokens to files, shells, compilers, GPUs, and tests, and the Driver is the seat that chooses what to do next, judges whether the last result was any good, and knows when to stop and ask. The argument for taking that seat is stated bluntly: in every other agent system that seat holds a human, which is why the work stops when they go to bed.

The seat is split across four roles, and each has an owning column and a prohibition column.

| Role | Owns | May not |
| :--- | :--- | :--- |
| Manager | Stage transitions, and where an admitted lesson is kept | Perform the work it is admitting |
| Planner | The next task, and the evidence it must produce | Move the campaign to the next stage |
| Engineer | Implementation, research, experiments, artifacts | Declare its own work complete |
| Reviewer | The verdict on correctness, evidence, limitations | Edit anything; it runs read-only |

Each prohibition closes a specific failure mode. The Manager cannot implement what it just admitted as a lesson, so admitting a rule is not the same as proving it. The Engineer cannot mark its own work complete, and the Reviewer, which may return `blocked`, runs read-only so a verdict cannot be quietly repaired in the same step.

The pipeline is written as `Manager` to `Planner` to `Engineer` back and forth with `Reviewer`, so review is a loop rather than a final gate.

## Four categories of action always stop for a human

Autonomy has an explicit boundary, and it is short: credentials, payment, irreversible actions, and publication always stop for a human.

That list is the operational content of the whole design. Everything else, planning, execution, verification, pausing, resuming beyond a single model turn, is delegated. The four exceptions are the ones where a wrong action is either unrecoverable or externally visible, and they are stopped rather than gated by a threshold.

The supporting claim about supervision is given as a measurement from the project's own campaigns. Across 27 campaigns and 1,548 hours it needed a human research decision about once every 310 hours, at a 95 to 99 percent duty cycle, and the rest is said to be in the technical report on arXiv.

The stated reason the review role matters is the reason this ratio is claimed at all: because the worker cannot grade its own work, nobody has to watch it. That is a design argument rather than a benchmark, and the duty cycle range is the kind of number worth reading alongside the source rather than quoting on its own.

The community route for questions is a WeChat group, where the page notes that group 1 is full and directs people to group 2, and says to open an issue if the printed QR expiry date has passed.

## Improvement comes from scoping lessons, not from retraining

Argus is described as improving without retraining. The mechanism is a scope ladder: admitted Skills and source-linked Wiki findings are scoped `project`, then `vertical`, then `global`, by how far they were shown to hold.

Read that ladder as an evidence threshold. A finding confirmed inside one repository stays at project scope, one that holds across a domain is promoted to vertical, and only a finding that survives everywhere becomes global. Nothing is retrained; the lesson is stored with the scope that matches its evidence.

New domains ship as verticals against a core that does not change. Seven verticals are built in and seventeen more live in the community package `argus-verticals`, installed from its git URL and registered through an entry point group. The package comment adds that the pre-rename entry point group is still read for one release, which is a deprecation window stated in the packaging rather than only in release notes.

The notable constraint is that there are zero references to the authority boundary in any of the seventeen verticals. The authority rules, the prohibitions and the human escalation points, are core code and not something a domain package can redefine, so adding a vertical cannot quietly widen what the system is allowed to do.

## Nine backends, and you reuse the CLI you already have

Installation is a table of agent CLIs, and the stated advice is to reuse the one you already work in, because Argus does not require a separate account. The documented backends are GitHub Copilot CLI, Pi, OpenAI Codex CLI, Claude Code, Cursor CLI, OpenCode, Grok Build, Qoder, and DeepSeek Harness.

Each row gives an install command and an authentication step, and the commands are mostly npm globals:

```bash
npm install -g @anthropic-ai/claude-code
```

```bash
npm install -g --ignore-scripts @earendil-works/pi-coding-agent
```

Authentication is per backend rather than for Argus: `copilot login`, `codex login`, running `claude` and then `/login`, `agent login` or `CURSOR_API_KEY`, `opencode auth login`, `grok login`, `qodercli login`, and for DeepSeek either `DEEPSEEK_API_KEY` or its Models page. The backend name in the second column, `copilot`, `codex`, `claude`, `cursor`, `pi`, `opencode`, `grok`, `qoder`, `dsh`, is the handle Argus uses internally.

Two platform rules are stated. All platforms need Node.js 22.12 or newer and one authenticated agent CLI, and the instructions say not to mix commands between platforms. Docker is not required for a normal installation, and is only an optional prerequisite for the separate Harbor evaluation integration. The preferred path is described as letting the code agent you already use install and verify Argus.

## Node 22 with type stripping sits on top of a Python 3.11 package

The repository is two toolchains in one tree, and the TypeScript side depends on a recent Node feature.

The root manifest is `argus-workspace`, private, with npm workspaces for `packages/contracts`, `packages/runtime`, and `packages/api`, and an engine floor of Node 22.12. The scripts run TypeScript directly with `node --experimental-strip-types` for contract generation and checking, and the Pi extension generator is checked the same way. `build` chains the two check scripts and then builds the three packages in order, `typecheck` adds `tsc --noEmit` over the test configuration, and `test` runs `node --import tsx --test` across the three test directories.

So type stripping is not a preference, it is the mechanism, which is why the Node floor is where it is.

The Python side is a hatchling package named `argus`, requiring Python 3.11 or newer, with classifiers for Windows, Linux, and macOS and Python 3.11 to 3.13. Its dependency list is short and tells you what the runtime actually does: `psutil` for processes, `portalocker` for file locking, `jsonschema` and `pydantic-settings` for configuration, `pypdf` for reading PDFs, `rich` for terminal output, `fastapi` with `uvicorn[standard]` and `websockets` for the API, and `mcp` for the model context protocol bridge, with `tzdata` pulled in only on Windows.

The verticals package is deliberately not a dependency, with a comment saying the framework must run without it.

## The version number appears in three places and they disagree

Pick any one of these and you will be quoting a different number from the others.

The README banner names v0.1.6. The Python package metadata declares version 0.1.8. The most recent published release is v0.1.9, titled Argus 0.1.9 macOS, and the two before it are v0.1.8 Windows and v0.1.7 Windows.

The release titles carrying a platform name is the second signal. A tag that says which operating system it targets is a packaging decision, and it means the same version number exists for several platforms rather than one universal artefact. Combined with the statement that source updates and packaged desktop previews are separate channels, the practical reading is that the source tree, the Python package version, and the desktop release are three different numbering surfaces that are not kept in step.

There is also a Tauri desktop directory, a `deploy/` directory, and an `update/` directory in the tree, which is consistent with a project that ships several installable artefacts.

If you need to reproduce someone's results, ask which of the three numbers they ran rather than assuming the tag name identifies it.

## Argus-Pi draws the Driver and Harness line explicitly

The optional component is Argus-Pi, a lightly customized fork of Pi with small Argus-focused improvements to task prompts, PDF reading, and execution and retry status reporting. Three improvements, described as small, and available as a source preview.

What matters is the division the page states: Argus remains the Driver for scheduling, role assignment, and task lifecycle, while Argus-Pi focuses on the Harness for model and tool execution. So the fork does not become the agent; it stays underneath it, and any other supported backend can be used instead.

The Harness side also learns at runtime. Ordinary Pi tasks can validate a reusable JSON transformation, use it in the current task, and retain its Skill and Wiki for later tasks, all under the existing task budget rather than a new one.

Three integration surfaces are documented alongside it. The Harbor Framework can invoke the complete bounded Manager, Planner, Engineer, and Reviewer runtime as a custom agent. A coding-agent plugin uses a packaged MCP bridge and host-specific Skills without changing the core runtime. And a Counterexample Lab provides a live research loop with an isolated Jacobian MCP bridge and safe in-app source updates.

The common thread is that each one adds capability without touching the role prohibitions, which is the constraint the architecture is built around.

## Conclusion

Argus fits a team that wants unattended agent campaigns with an audit trail rather than a chat interface, since the reviewer is read-only and the escalation rules are narrow and named. It does not fit someone who needs the latest code from this repository and assumes it is canonical, because the official source lives elsewhere and the versions disagree across the README, the Python package, and the release tags. Before adopting it, install from the official repository, pick a backend you are already authenticated with, read which actions always stop for a human, and treat the campaign duty cycle figures as the project's own report rather than an independent measurement.

## FAQ

### What is the purpose of Argus?

Argus is a runtime for persistent, reviewed autonomy in research and engineering, letting long agent work plan, execute, verify, pause, and continue beyond a single model turn. It takes the Driver role that other systems leave to a human, splitting it into Manager, Planner, Engineer, and Reviewer, and it runs against nine agent backends.

### What does the Argus Reviewer role do?

It owns the verdict on correctness, evidence, and limitations, may return blocked, and runs read-only so it may not edit anything. The Engineer cannot declare its own work complete, so grading is never done by the worker.

### When does Argus stop for a human?

Four categories always stop: credentials, payment, irreversible actions, and publication. The project reports that across 27 campaigns and 1,548 hours it needed a human research decision about once every 310 hours at a 95 to 99 percent duty cycle.

### How do I install Argus?

All platforms need Node.js 22.12 or newer and one authenticated agent CLI, and the advice is to reuse the CLI you already work in since no separate Argus account is required. Docker is not needed for a normal install, only for the optional Harbor evaluation integration.

### Which coding agents can Argus drive?

Nine are listed: GitHub Copilot CLI, Pi, OpenAI Codex CLI, Claude Code, Cursor CLI, OpenCode, Grok Build, Qoder, and DeepSeek Harness. Most are installed as npm globals and authenticated with their own login command.

### Where is the official Argus source code?

microsoft/ArgusAgent is named as the official source repository, and lbx154/Argus is the development preview, with changes reaching the official repository through synchronization. Within the preview repository, main is the stable branch and dev is the development branch.

## Sources

- [Issues](https://github.com/lbx154/Argus/issues)
- [lbx154/Argus on GitHub](https://github.com/lbx154/Argus)
- [License: MIT](https://github.com/lbx154/Argus/blob/main/LICENSE)
- [README](https://github.com/lbx154/Argus/blob/main/README.md)
- [Releases](https://github.com/lbx154/Argus/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/lbx154-argus
