# AXI argues the CLI was built for humans and agents pay for it

> A set of ten design principles for command line tools, a benchmark comparing agent success rates and token cost against raw CLIs and MCP, and one SDK package that most people will never install.

**kunchenguid/axi** — Design principles for agent ergonomics. Higher accuracy with lower token cost than both MCP and regular CLI.

- Repository: https://github.com/kunchenguid/axi
- Website: https://axi.md/
- Stars: 2,201 · Forks: 194
- Language: TypeScript
- License: MIT
- Published: 2026-10-07 · Updated: 2026-10-07 · Language: en
- Canonical page: https://hysenlabs.com/projects/kunchenguid-axi

## The premise: output shape is the cost driver

The repository description states the thesis in one line: higher accuracy with lower token cost than both MCP and regular CLI. The README expands it into a framing about two dominant paradigms. CLIs were originally built for humans, and structured tool protocols like MCP impose overhead. AXI is offered as a third option, agent-native CLI tools built from ten design principles that treat token budget as a first-class constraint.

The argument is easy to sympathize with and worth restating precisely. When an agent runs a command, it pays for the output twice: once in tokens and once in the reasoning needed to work out what the output means. A command that prints a formatted table of thirty columns spends tokens on six of them. A command that prints nothing when there is nothing to report makes the agent guess whether the query succeeded. Neither failure is a bug in either tool; both are consequences of a human-first interface.

That framing also explains why the project is a set of principles plus a catalog rather than a framework. There is no runtime here that intercepts your agent's tool calls. What AXI provides is a checklist, a specification document, and a growing set of reference implementations that follow it.

## Ten principles, generated rather than maintained by hand

The principles table in the README is not written by a person editing markdown. It sits between generated markers and states that it is generated from `principles.yaml`, with the full specification of each principle living in a skill file under `.agents/skills/axi/SKILL.md`. The repository root also carries a `catalog.yaml` and a `VISION.md`, and both the principles table and the catalog tables are regenerated by scripts rather than edited.

That structure is itself one of the ten principles applied to the project itself. The table makes several of them concrete enough to argue with. Token-efficient output says to use TOON format for around 40 percent token savings over JSON. Minimal default schemas asks for three or four fields per list item rather than ten or more. Content truncation asks for truncated large text with size hints and a `--full` escape hatch, which is the detail most implementations get wrong: an escape hatch matters more than the default.

The remaining principles are about the interaction shape rather than the data shape. Pre-computed aggregates exist to eliminate round trips. Definitive empty states ask for an explicit zero results instead of ambiguous empty output. Structured errors and exit codes cover idempotent mutations, no interactive prompts, and failing loudly on unknown flags. Content first says running with no arguments should show live data rather than help text, which is the single most disruptive of the ten for anyone used to conventional Unix behaviour. Ambient context describes installing opt-in session integrations first and then offering an on-demand skill.

## Two benchmarks, one model, and a caveat the authors put in writing

The README publishes two benchmark tables, and they are the most interesting part of the repository because the authors included their own limitations.

The browser benchmark covers 490 runs, which the README breaks down as 14 tasks times 7 conditions times 5 repeats, using Claude Sonnet 4.6. The AXI implementation, chrome-devtools-axi, records 100 percent success at an average cost of $0.074 and 21.5 seconds, against 99 to 100 percent for dev-browser, agent-browser, and four variants of chrome-devtools-mcp ranging from $0.091 to $0.120. The cost gap is under half a cent. The turn gap is more interesting: 4.5 turns against 6.2 to 7.6 for the MCP variants.

The GitHub benchmark covers 425 runs, 17 tasks times 5 conditions times 5 repeats. Here the gap widens. gh-axi records 100 percent success at $0.050 and 15.7 seconds, the plain `gh` CLI records 86 percent at $0.054, GitHub MCP records 87 percent at $0.148 and 34.2 seconds, and GitHub MCP with ToolSearch drops to 82 percent while costing more.

Then the caveat, which deserves to be quoted rather than paraphrased: Claude Sonnet 4.6 is the model the published runs used, not a limit of the harness, both benchmarks accept a `--model` flag passed straight through to the agent CLI, and results for other models are simply not published here and may respond differently to output format. Read the tables as evidence about one model's sensitivity to output density, and as a claim rather than a result for anything else. There is no independent reproduction in the repository either, and no benchmark harness is published beyond the two `bench-browser` and `bench-github` directories.

## What you actually install, and from where

This repository is not the tool you install. It is the principles, the catalog, and the SDK. The quick start makes that explicit by pointing at reference implementations in their own repositories and then running two install commands:

```sh
npm install -g gh-axi
npm install -g chrome-devtools-axi
```

The second half of the quick start is an instruction rather than a command: add a line to your `CLAUDE.md` or `AGENTS.md` telling the agent which tool handles which domain. That is the ambient context principle applied to adoption. An agent will not find these tools on its own, so the design includes the step where you tell it, in the file it already reads.

The catalog lists official implementations as reference implementations maintained by the AXI project, validating the principles across different domains: gh-axi for GitHub, chrome-devtools-axi wrapping chrome-devtools-mcp, lavish-axi for turning agent-generated HTML artifacts into a review surface where humans annotate and send feedback back, and quota-axi. Both the principles table and the catalog tables are generated from yaml, and the README documents a CONTRIBUTING.md path for adding an AXI to the catalog, which is the mechanism by which this becomes a registry rather than one person's preference list.

Worth noting that the reference tools wrap what already exists. gh-axi wraps the official `gh` CLI and chrome-devtools-axi wraps chrome-devtools-mcp. The claim is not that the underlying interface is wrong, but that its default output is.

## A private monorepo with one published package

The package.json at the root is named `axi-repo-tools` and marked private, so nothing installs from this directory. It pins pnpm as its package manager at version 11.1.1, declares the workspace as ES modules, and defines scripts that are worth reading as a description of the project's workflow: docs:gen runs a generate-docs script, docs:check runs the same script with a check flag, and docs:test runs a node test file for the generator itself. The release automation is release-please, with a matching config file and a manifest.

The published artifact is a single package, axi-sdk-js, and its release tags all carry the package name as a prefix: axi-sdk-js-v0.1.10 on 2026-08-07, axi-sdk-js-v0.1.11 on 2026-08-21, and axi-sdk-js-v0.1.12 on 2026-09-16. The version numbers show what the SDK is doing. v0.1.10 added a dependency-free version fast path. v0.1.11 added a project-scope session-start hook with install, status and uninstall subcommands, and routed initialize and resolveContext failures through the AXI error contract. v0.1.12 handles stdout EPIPE gracefully.

That last one is a small bug with a lot of information in it. A broken pipe on stdout is what happens when a tool's output is piped into something that stops reading early, which is a common agent pattern and a classic source of noise and spurious failures. Fixing it in a patch release suggests the author has watched that failure happen.

The SDK is also pre-1.0 across roughly nine months of releases, so treat its interface as unstable. Eight open issues against 2,170 stars is an unusually clean ratio, though it also reflects a project where most discussion happens elsewhere rather than on the tracker.

## Conclusion

AXI is doing something narrower and more useful than it first appears to be. It is not proposing a new transport between agents and tools, and it is not asking you to abandon MCP; it is arguing that the shape of a command's output decides how many tokens an agent spends reading it, and that most output shapes were chosen when a human was at the keyboard. The honest weak spot is the evidence. Two benchmarks, both on a single model, both run by the project that wrote the tools being compared, with a stated caveat that other models may respond differently to output format. If the idea holds, it will hold because the principles are cheap to apply to a tool you already own, so the useful first step is not installing anything from the catalog, it is picking one command you call often and reshaping its output to match the ten items in the table.

## FAQ

### What is Axi used for?

AXI is a set of ten design principles for building command line tools that AI agents drive, plus a catalog of implementations that follow them. The premise is that both classic CLIs and structured tool protocols cost an agent more tokens than the task requires, so reshaping a tool's output reduces cost and improves success rate.

### Does AXI replace MCP?

No. The official chrome-devtools-axi reference implementation wraps chrome-devtools-mcp, and gh-axi wraps the official `gh` CLI. The claim is that the default output of those interfaces is expensive for an agent to read, not that the interfaces themselves are wrong.

### What do the benchmarks actually show?

The browser table covers 490 runs of 14 tasks across 7 conditions with Claude Sonnet 4.6, where chrome-devtools-axi reports 100 percent success at $0.074 against $0.091 to $0.120 for MCP variants. The GitHub table covers 425 runs and shows the wider gap: gh-axi at 100 percent and $0.050 versus GitHub MCP at 87 percent and $0.148. The README states no other model's results are published.

### What package do I install from this repository?

The root package.json is named axi-repo-tools and is private, so nothing installs from this repository. The only published package is axi-sdk-js, tagged as axi-sdk-js-v0.1.x and most recently released as v0.1.12. The tools themselves, gh-axi and chrome-devtools-axi, live in their own repositories.

### Which principle would change my existing CLI the most?

Content first, because it inverts the convention that running a command with no arguments prints help. AXI argues that running with no arguments should show live data instead. Several others are cheaper to adopt first, such as definitive empty states, pre-computed aggregates, and a `--full` flag that lifts content truncation.

## Sources

- [kunchenguid/axi on GitHub](https://github.com/kunchenguid/axi)
- [License: MIT](https://github.com/kunchenguid/axi/blob/main/LICENSE)
- [Project website](https://axi.md/)
- [README](https://github.com/kunchenguid/axi/blob/main/README.md)
- [Releases](https://github.com/kunchenguid/axi/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/kunchenguid-axi
