CLI tool
first-fluke/oh-my-agent avatar
first-fluke/oh-my-agent

oh-my-agent: a vendor-agnostic agent harness that gates on artifacts, not summaries

Portable, vendor-agnostic agent harness for project-specific skills, workflows, and agent teams aligned with your codebase, conventions, and engineering standards.

1,323 stars151 forksTypeScriptMIT

At a glance

What is it?
oh-my-agent keeps one .agents/ directory as the source of truth for skills, workflows and gates, then projects it into a dozen agent runtimes. Its selling point is mechanical verification: a Stop hook that refuses to end a session until your own typecheck, test and lint scripts exit 0.
Who is it for?
Adopt oh-my-agent if your team already runs parallel coding agents across more than one runtime and you want a single .agents/ directory plus gates that read files on disk instead of trusting an agent's summary. Skip it if you work inside one vendor's native tooling and have no need for portability, or if your project has no typecheck, test or lint script for the Stop hook to call.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 5 days ago.
What is it written in?
Mainly TypeScript, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 25, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The problem: agents narrate success, artifacts do not

Spawning parallel agents is cheap. Knowing whether they finished is not. An agent can write "tests pass, all criteria met" and nothing inside that same session can contradict it, because the session's own context is the only evidence available to it. oh-my-agent's README states the premise plainly: "Agents narrate success. oh-my-agent checks the artifacts."

The project is aimed at engineers running multi-agent workflows on a real codebase, the kind where a skipped refactor or a silently regressed test costs more than the tokens spent. It is not a chat client and not a model. It is a harness: a directory of skills, workflows, rules and gates that sits alongside your code and is projected into whichever agent runtime you happen to be using. The README frames the target as "project-specific skills, workflows, and agent teams aligned with your codebase, conventions, and engineering standards."

How the gates work: exit codes, artifact files and an append-only log

The verification layer is deliberately mechanical. No LLM is asked whether the work looks correct; a command exits 0 or it does not, a file is on disk or it is not.

The Stop hook lives at `.agents/hooks/core/persistent-mode.ts`. According to the README, it blocks session termination while a persistent workflow is active and runs the configured gate script before allowing a stop. The important constraint is the allowlist: only `typecheck`, `test` and `lint` are executable, so an agent that writes some other command into the state file gets it ignored rather than run. The hook is capped at 5 reinforcements, which means a gate that stays red cannot trap the session forever.

The Anti-Circumvention Gate is the second layer. `oma ralph:verify --json` checks four artifacts that a shortcut cannot fake: ultrawork's phase records, the plan JSON, a distinct QA agent's result file, and a distinct refactor agent's result file. If an artifact is missing, the phase did not run, whatever the summary says. The verdict is JSON, and the README is explicit that the JSON verdict, not the agent's summary, is the result.

Verification is separated from implementation. The judge is spawned as a separate agent with fresh context and is briefed on the criteria only, never on what the implementer claims to have fixed. It re-verifies every criterion each iteration, including criteria that already passed, on the reasoning that fixing C2 is how C1 silently regresses.

Every gate pass, gate failure and decision appends one JSON line to `.agents/state/sessions/{sid}/events.jsonl`, stamped with vendor and runtime session id. The log is append-only and cross-vendor, so a run can be audited after the fact rather than reconstructed from memory. Budgets are enforced through the same machinery: `session.quota_cap` caps tokens, spawn count and per-vendor spend, and the orchestrator refuses the next spawn when a dimension is exceeded. When the wall-clock budget runs out, the README says the Stop hook stops honestly with partial status recorded on the event log.

Installing oh-my-agent and running a first verification

The install scripts auto-install bun, uv and serena if they are missing. On macOS or Linux the README gives a curl pipe into bash:

bash
curl -fsSL https://raw.githubusercontent.com/first-fluke/oh-my-agent/main/cli/install.sh | bash

On Windows the equivalent is a PowerShell invocation:

powershell
irm https://raw.githubusercontent.com/first-fluke/oh-my-agent/main/cli/install.ps1 | iex

If you would rather manage the prerequisites yourself, the manual path requires bun, uv and serena to be present already:

bash
bunx oh-my-agent@latest

There is also an APM distribution. `apm install first-fluke/oh-my-agent` deploys all skills to every detected runtime, and a single skill can be pulled with a path such as `apm install first-fluke/oh-my-agent/.agents/skills/oma-frontend`. Be aware of the scope difference: APM ships skills only, while workflows, rules, `oma-config.yaml`, keyword-detection hooks and the `oma agent:spawn` CLI come from `bunx oh-my-agent@latest`. The README advises picking one distribution per project to avoid drift.

After the interactive setup you choose a preset. Backend, Content, DevOps, Frontend, Fullstack, Fullstack Mobile, Fullstack Web, Mobile, Research, or All for every agent and skill. Once a project is configured, the per-agent check battery runs through one command:

bash
oma verify <agent>

That runs a shared core (scope violation, charter alignment, hardcoded secrets, TODO scan, declared outputs) plus type-specific checks such as TypeScript strict, tests, raw SQL, Flutter analyze and inline styles. Expect a pass or fail per check, not a prose report. For the workflow-level gate, `oma ralph:verify --json` returns the JSON verdict described above.

Skill quality is measured rather than assumed. `oma skills eval` measures utility lift on held-out tasks, treatment against baseline, and `oma skills opt` keeps only edits that improve the measured lift.

The portability claim and what it costs you

The README's argument for keeping `.agents/` as the single source of truth is that verification is worth little if it is locked to one vendor. oh-my-agent projects that directory into each runtime's native layout, so the same skills, workflows, rules and gates are shared, and switching vendors is described as a config change rather than a migration. The repository layout backs this up: `.claude-plugin/`, `.cursor-plugin/`, `.mcp.json`, `mcp.json`, `integrations/` and `action/` all sit at the top level alongside `cli/` and `web/`.

The cost is a projection layer you now own. Every runtime you add is another layout the harness has to keep in sync, and the README's own warning about picking one distribution per project to avoid drift applies to this too. If your team uses exactly one runtime, the indirection buys you little and gives you one more thing to debug when a skill does not appear where you expect it.

There is a second, subtler cost. The gates only know what you tell them. The Stop hook can run `typecheck`, `test` and `lint` and nothing else. A project whose quality bar lives in manual review, design critique or data correctness has no exit code for the hook to read, and the harness will happily let the session end. The mechanism is strong within its lane and silent outside it.

Where the design is thin

The README documents the happy path well: install, pick a preset, verify, audit the event log. It does not document rollback. If a preset turns out to be wrong, or a projected layout conflicts with files you already have in `.claude/` or `.cursor/`, there is no described procedure for undoing it. The repository has an `uninstall` shaped hole that the README leaves open, which matters because the install scripts write into multiple runtime directories at once.

The reinforcement cap of 5 is a real trade-off, not a bug. It prevents a permanently red gate from trapping a session, but it also means an agent can outlast the gate by simply failing five times. The event log will record the partial status, so the failure is auditable, but nothing in the described mechanism blocks the session from ending. If your workflow depends on a hard stop, you are relying on the log rather than the hook.

The judge protocol re-verifies every criterion each round, which is the right call for catching regressions and the expensive one. On a workflow with many criteria and several iterations, that is a lot of fresh-context agent invocations. `session.quota_cap` exists precisely because this can run away, and you should expect to tune it rather than accept a default.

Alternatives and the actual difference in approach

The closest alternative in the README's own framing is Microsoft's Agent Package Manager, invoked as `apm install first-fluke/oh-my-agent`. The difference is scope, not philosophy. APM distributes skills to detected runtimes; it does not carry workflows, rules, `oma-config.yaml`, keyword-detection hooks or the `oma agent:spawn` CLI. If all you want is to share skill definitions across editors, APM is the smaller dependency and the README says so directly. If you want the Stop hook and the workflow gates, APM alone will not give them to you.

The other alternative is the native tooling of whichever runtime you already use. Claude Code, Cursor, Codex and OpenCode each have their own plugin and hook systems, and each is maintained by the vendor that ships the runtime. A single-vendor setup avoids the projection layer entirely and gets new runtime features the day they ship. What it does not give you is the cross-vendor event log or a gate definition that survives a vendor switch. That is the actual trade: oh-my-agent sells portability and mechanical verification, and charges you a sync layer plus a dependency on a TypeScript CLI whose workspace is pinned to `node >=26.0`.

Maintenance, licence and upgrade cost

The repository is not archived, and the last push was on 2026-08-28. Recent releases follow the `cli-v12.8.0` and `web-v4.2.5` naming pattern, with the CLI and the documentation site versioned separately through release-please. The workspace `package.json` is at version 14.10.0 and is private, so the published artefact is the CLI rather than the workspace root.

Licence is MIT, which permits commercial use, modification and redistribution provided the copyright notice and permission notice are retained. That is a statement about the licence text, not legal advice; if your organisation has rules about bundled agent skills or about piping an install script into a shell, check them before rollout.

Upgrade cost is concentrated in two places. The CLI is distributed via `bunx oh-my-agent@latest`, so the version you get is whatever is latest at invocation time unless you pin it. And because `.agents/` is projected into runtime-specific directories, a CLI upgrade can change the projected layout, which is the moment drift between distributions becomes visible. The repository ships `mise.toml`, `biome.json`, `commitlint.config.mjs` and a `bench:oma-hook` script for hook latency, so the project does measure its own hook overhead, though the README does not publish a figure.

Editorial conclusion

Adopt oh-my-agent if your team already runs parallel coding agents across more than one runtime and you want a single .agents/ directory plus gates that read files on disk instead of trusting an agent's summary. Skip it if you work inside one vendor's native tooling and have no need for portability, or if your project has no typecheck, test or lint script for the Stop hook to call. Before committing, verify three things: that the preset you pick covers the agents you actually need, that your gate commands exit 0 on a clean checkout, and that the .agents/state/sessions/{sid}/events.jsonl log is readable by whatever you use to audit runs. The Stop hook is capped at 5 reinforcements, so a permanently red gate ends the session with partial status rather than trapping you.

Frequently asked questions

How do I install oh-my-agent?

On macOS or Linux, the README gives a curl script piped into bash that auto-installs bun, uv and serena if they are missing. On Windows there is a PowerShell equivalent, and if you already have the prerequisites you can run bunx oh-my-agent@latest instead.

How do I uninstall oh-my-agent?

The README does not document an uninstall or rollback procedure, even though the install scripts write into multiple runtime directories such as .claude, .cursor and .codex. If you need a clean removal, plan for it before installing.

What runtimes does oh-my-agent work with?

The README says it runs across a dozen agent runtimes from one portable .agents/ directory, and the APM install path lists .claude, .cursor, .codex, .opencode, .github and .agents as the detected runtime layouts. The full supported list is in the README table, which is truncated in the repository listing.

What exactly does the Stop hook check in oh-my-agent?

It blocks session termination while a persistent workflow is active and runs the configured gate script before allowing a stop. Only typecheck, test and lint are executable, and an agent that writes any other command into the state file gets it ignored rather than run.

Where does oh-my-agent store its audit trail?

Every gate pass, gate failure and decision appends one JSON line to .agents/state/sessions/{sid}/events.jsonl, stamped with vendor and runtime session id. The log is append-only and cross-vendor, so it can be read after the run.

Official sources

  1. Official documentation
  2. Official README
  3. Project repository
  4. Release notes
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/first-fluke-oh-my-agent.svg)](https://hysenlabs.com/projects/first-fluke-oh-my-agent)