Model or dataset
boshu2/agentops avatar
boshu2/agentops

AgentOps judges every change from a session that did not write it

The operations layer for agentic engineering — portable skills and contracts connecting intent, agents, software factories, and independent judgment.

447 stars41 forksGoApache-2.0

At a glance

What is it?
A Go CLI called ao plus a set of SKILL.md files that port between Claude Code, Codex, Cursor, OpenCode and others. The interesting part is not the install, it is the verdict set: a change is judged by a fresh session against the same Given/When/Then scenarios, and the third verdict is NOT_PROVEN.
Who is it for?
AgentOps fits a team that has already tried telling an agent to review its own work and found that useless, because the separation here is structural rather than a prompt asking nicely. It is also a fair fit if your failure mode is vocabulary drift, since the Given/When/Then scenarios double as the review contract.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 1 day ago.
What is it written in?
Mainly Go, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 2, 2026, and from our analysis. They are not legal advice.

Editorial analysis

One installer per agent, or you get every skill twice

The quickstart opens with a warning rather than a command, and it is the first thing to get right: pick one method per agent, because a plugin plus npx on the same agent gives you every skill twice.

Claude Code and Codex each have their own plugin path:

bash
claude plugin marketplace add boshu2/agentops
claude plugin install agentops@agentops-marketplace
claude plugin details agentops@agentops-marketplace

The Codex variant swaps `plugin install` for `plugin add` and checks with `codex plugin list --json`. The Claude Code bundle carries skills, four agents and tool call guards, and skills there are invoked as `/agentops:<skill>` while Codex uses `$agentops:<skill>`.

Everything else goes through one command from your project directory, with Node.js installed:

bash
npx skills@latest add boshu2/agentops

The installer targets are `cursor`, `opencode`, `gemini-cli`, `antigravity`, `pi`, `grok` for Grok Build, and `openclaw`. Grok Bot has no installer target, so you add the same `SKILL.md` folders by hand through its own skill settings.

For scripts, name the agents explicitly rather than letting the installer decide:

bash
npx skills@latest add boshu2/agentops -g -a cursor opencode -y

`-g` makes it a user level install, and `-y` on its own can install into every agent the installer knows about.

Where to look after installing is a documentation question rather than a tooling one. `docs/SKILL-ROUTER.md` is linked from the README header and is the routing table for skills. `docs/install-day2-ops.md` covers the extra tools some skills need, the Codex specific roles and read limits, and updating afterwards. And `docs/contracts/multi-runtime-tier-charter.md` holds the host coverage and limits mapping, which is the document to read if you care whether a skill has actually been exercised on the host you are using rather than merely being installable there.

The unit of work is a scenario in the issue, not a file

Intent is written as behaviour, in Gherkin, before anything is built. The example in the documentation is a job redelivery feature:

gherkin
Feature: Job redelivery is idempotent
  A Job is one unit of queued work. Delivering it again never repeats its side effect.

  Scenario: A completed Job is delivered again
    Given Job "J-42" completed and charged the customer $20
    When the worker receives Job "J-42" again
    Then it returns the completed result of "J-42"
    And the customer has been charged $20 exactly once

  Scenario: A Job that failed before charging is delivered again
    Given Job "J-43" failed before charging the customer $20
    When the worker receives Job "J-43" again
    Then Job "J-43" completes

Two rules sit around that snippet. Each scenario gets concrete data, one action, and an observable result, so there is no ambiguity about what a step means. And the scenarios live in the issue or the conversation, with no `.feature` file required, so this is not a Cucumber suite you have to wire into a test runner.

The vocabulary part matters as much as the syntax. The stated aim is one domain term per concept, carried through intent, code and tests, and a system that calls queued work a Job should call it a Job everywhere. Three names for one concept is listed in the project's own table as a failure mode worth designing against, and Given/When/Then is the tool it reaches for.

Three verdicts, and the author never approves its own work

The step table is four rows and the third is the interesting one.

Shape uses the `plan` skill, turning a request into scenarios for one small change, and it is explicitly skippable when the intent is already clear. Build uses `implement`, which makes the change and tests both scenarios. Judge uses `validate`, described as a new session that did not write the change returning `PASS`, `FAIL` or `NOT_PROVEN` against the same scenarios. Learn uses `memory`, which is optional and produces reviewed `.context/` pages that later work can query.

`NOT_PROVEN` is the verdict to sit with. A green test run and a self-reported pass are not the same claim, and this loop distinguishes a change that demonstrably meets the scenarios from one where nothing established it either way. A judge that cannot produce evidence returns the third option instead of quietly passing or quietly failing.

Two operational rules follow. You can enter at whatever step you need, and an existing change goes straight to Validate, so this is not only for new work. And the author never approves its own work, which is the whole reason validate spawns a fresh session rather than asking the implementing session to look harder.

Merging and releasing are explicitly left to your repository's own rules.

rpi runs the same loop without stopping to ask you

The `rpi` skill is Plan, Implement and Validate for one outcome, run end to end without check ins. Your agent's own permission prompts still apply, which is the boundary: AgentOps does not suppress the security layer of the tool you are driving.

What it does define is where the loop stops. Three conditions are named: acceptance, a blocker, or a spent limit. A spent limit is the interesting one, because it means the skill is expected to run against some ceiling rather than until the tokens run out, and the documentation does not say what that ceiling is set to.

The acronym is the loop itself, and the naming is consistent with how the rest of the project is built. The skills are files, not code paths, so the same file can be invoked under different names by different agents, and the invocation prefix changes with the host while the content does not.

One more detail from the quickstart is easy to skim past and useful to know: most skills need nothing but your coding agent, but Validate also needs the `ao` CLI. So a partial install is coherent. You can plan, implement and review work without the binary and only lose the judging step.

The binary itself is the Go half of the repository. There is a `cli/` directory that the Makefile builds and tests with recursive make calls, a `bin/` directory, a `homebrew-tap/` for distribution, and `.goreleaser.yml` for the release builds. That matters because the skills are portable text and the CLI is the one part that is compiled and versioned: the release cadence runs at minor bumps, with v3.6.0 on 2026-08-18, v3.7.0 on 2026-09-14 and v3.8.0 on 2026-09-22, and the last push was on 2026-10-01.

Goals are experimental and they need Beads

Work larger than one change is supposed to become a goal made of many RPI cycles, tracked in Beads, which the documentation describes as a dependency-aware issue tracker from the gastownhall project.

The setup is a prerequisite rather than a suggestion. Beads is installed with `brew install beads`, then initialised in your repository with `bd init`, and the goals feature is labelled experimental until you have done both. On a machine without Homebrew, that is a wall you have to clear before anything else works.

Once it is there, the graph is what holds the larger picture: intent, dependencies between pieces of work, and verdicts. That is the answer the project gives to losing the thread on work bigger than one session, and it is a different mechanism from the session-level skills, because the state outlives any one conversation.

The first step in that flow is an interview, and the interview skill asks one question at a time. The stated purpose is to settle a half-formed goal before agents go autonomous, which is a coherent position given the rest of the design: if a change is judged against written scenarios, the scenarios themselves had better be right.

council runs judges in fresh contexts and keeps the dissent

The project has an answer for the case where one model's opinion is not enough, and the skill is called `council`.

Judges run in fresh contexts. You assign each one a model, an effort level and a perspective, and the assignment can be one model family or several vendors. Then you pick how they interact: compare, duel, which scores each other's ideas, or debate to your majority. Dissent is kept rather than discarded, which is the detail that separates this from a majority vote written as a conclusion.

The council can also answer an interview for you, so a multi-model deliberation can settle a goal definition before work starts. A separate skill, Idea Genie, is the brainstorming counterpart for options.

Compare the shape of this against validate, which is a single fresh judge against written scenarios. Council is for open questions where there is no scenario to check against yet, and validate is for closed ones where there is. Keeping those two apart is what stops the expensive multi-model path from being the default, and it also means a council verdict is a judgement while a validate verdict is a result.

Memory writes lessons to .context/ after a disclosure check

The failure this addresses is repeating the last session's investigation, and the mechanism is a `memory` skill that turns reviewed, disclosure-checked lessons into `.context/` pages described as safe to commit.

Two qualifiers in that sentence carry the weight. Reviewed, because the output is meant to pass through a human before it becomes shared project state. And disclosure-checked, because an agent writing down what it learned about your codebase can easily write down something you did not intend to publish. The repository root has a `.context/` directory to match, sitting alongside `.agents/` for research and idea reports, and `.out-of-scope/` for whatever the project has decided not to take on.

The rest of the durability story is scattered across those directories by design. Plans and decisions are saved on the bead or the issue, so Plan, Interview and Navigate outputs live with the work item rather than in a chat log. Research and idea reports go under `.agents/`. Council reports are written wherever you choose, which is the one deliberately unopinionated location in the set.

So there are four places state accumulates: the issue or bead, `.agents/`, `.context/`, and your choice for council output. Two of those are committable, which is what makes the memory useful to someone who was not in the session.

The default make target is a release gate

The Makefile treats the repository as something with generated artefacts that can drift, and the default goal is the strictest thing available.

`.DEFAULT_GOAL` is `local-ci`, which runs `./scripts/ci-local-release.sh`, and a comment notes that the script already covers build, test and release binary validation. `local-ci-fast` is the same script with `--quick`, and `ci` is an alias for the full gate. With that as the default, a bare `make` is not a build, it is the check you would run before tagging.

Three targets handle derived artefacts, and the split between them is the useful part:

bash
make regen-all
make regen-check
make docs-check

`regen-all` runs `scripts/regen-all.sh` to regenerate every derived artefact after adding a skill or command, described as a one command finaliser. `regen-check` runs the same script with `--check`, writes nothing, and is described as a pre-push gate. `docs-check` runs a CLI reference generator in check mode and a documentation release validator, so a stale reference page fails the build.

The root shows the same discipline outside the Makefile: `.goreleaser.yml` for the Go binary, `.gitleaks.toml` for secret scanning, `renovate.json` for dependency updates, `.codecov.yml` for coverage, `.githooks/` for hooks, a `homebrew-tap/` and a `bin/` for distribution, `mkdocs.yml` with `requirements-docs.txt` for documentation, and `skills.sh.json` plus `registry.json` as the registries the installers read.

The rest of the tree is where the project's ambitions are visible. There are `evals/` and `evidence/` directories, which suggests agent behaviour is measured and the measurements are kept, plus `tests/`, `spec/`, `schemas/`, `packs/`, `plugins/`, `workflows/`, `deploy/`, `lib/`, `scripts/` and `examples/`. Alongside the README sit four prose files that look like the project's constitution: `AGENTS.md`, `CLAUDE.md`, `PRODUCT.md`, `PROGRAM.md` and `GOALS.md`, plus `PRACTICE-REGISTRY.md` and a `NOTICE` file. `goals-affects-files.yaml` at the root suggests goals are wired to the files they are allowed to touch.

Editorial conclusion

AgentOps fits a team that has already tried telling an agent to review its own work and found that useless, because the separation here is structural rather than a prompt asking nicely. It is also a fair fit if your failure mode is vocabulary drift, since the Given/When/Then scenarios double as the review contract. Skip it if you have no appetite for writing scenarios before the code, because the loop's whole value depends on having something to judge against, and skip the goals half until you have Beads working, since that path is marked experimental and needs a separate issue tracker installed first. Verify three things: that you install by exactly one method per agent, since a plugin plus npx on the same agent gives you every skill twice; that you start a new session so the skills load, and that Validate can find the `ao` binary, since most other skills do not need it; and that `.context/` pages, which memory writes for you to commit, get reviewed for disclosure before they land. The project is Apache 2.0, v3.8.0 shipped on 2026-09-22, and the last push was on 2026-10-01.

Frequently asked questions

What is AgentOps?

A set of optional `SKILL.md` skills plus a Go CLI named `ao` that bring DevOps discipline to coding agents. You state intent as behaviour, the skills carry it through Plan, Implement and Validate, and the change is judged by a fresh session that did not write it.

What does AgentOps do?

It turns a request into Given/When/Then scenarios for one small change, has the agent make the change and test both scenarios, then hands the same scenarios to a new session that returns PASS, FAIL or NOT_PROVEN. Larger work becomes a goal made of many such cycles, tracked as a dependency graph in Beads.

Do I need the ao CLI to use the AgentOps skills?

Not for most of them. The quickstart says most skills need only your coding agent, and Validate is the one that also needs the `ao` CLI. That makes a partial install coherent: you can plan, implement and review without the binary and only lose the judging step.

Which coding agents can I install AgentOps skills into?

Claude Code and Codex have their own plugin marketplaces. Everything else goes through `npx skills@latest add boshu2/agentops`, with installer targets for cursor, opencode, gemini-cli, antigravity, pi, grok for Grok Build and openclaw. Grok Bot has no installer target, so its `SKILL.md` folders are added by hand.

How do I avoid installing every AgentOps skill twice?

Pick one install method per agent. The README opens the quickstart by warning that a plugin plus npx on the same agent gives you every skill twice, so do not run the Claude Code or Codex plugin commands and the npx installer against the same agent. A new session is needed afterwards for the skills to load.

Official sources

  1. boshu2/agentops on GitHub
  2. License: Apache-2.0
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/boshu2-agentops.svg)](https://hysenlabs.com/projects/boshu2-agentops)