CLIARE: measuring CLI agent readiness from the released binary
CLI agent-readiness measurement, command-shape inference, and CI scorecards
At a glance
- What is it?
- CLIARE probes a command-line tool as a black box, records what it does, and emits a command index, issue ledger and CI scorecards. It is aimed at maintainers who ship CLIs and at teams wiring those CLIs into agent harnesses.
- Who is it for?
- Adopt CLIARE if you maintain a CLI whose surface changes between releases, or if you are writing a harness that has to call someone else's CLI without guessing. Skip it if your tool is a single command with no subcommands, no flags worth documenting and no side effects to catch.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 60 days ago.
- What is it written in?
- Mainly Rust, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The contract gap CLIARE was built to close
A CLI written for humans has one contract: the help text, plus whatever the docs say. An agent harness needs a different one. It needs to know which subcommands exist in the binary you actually shipped, which flags and positionals are safe, which commands produce parseable JSON or YAML, and which paths need auth, a project directory, fixtures, network access or a local daemon. The README frames this as a drift problem: docs, --help and the released binary can tell different stories, and humans route around that while agents discover it by trial and error. CLIARE's answer is to stop reading documentation and start measuring the binary. The README describes the project as "like OpenAPI/Swagger for CLIs, but generated from runtime evidence rather than hand-written docs", which is the clearest statement of intent in the repository. The audience is three groups named in the README: maintainers who want a fix queue before a release, security reviewers who want to know what a safe-looking command writes to disk, and harness authors who want a map instead of a guessing loop. It is not a linter for your argument parser and it does not read your source.
How the probe loop turns runtime evidence into a command index
The mechanism is black-box probing under bounded controls. You point CLIARE at an executable and a context, it runs the binary, snapshots filesystem state around each probe, and records evidence rather than assumptions. The workspace layout in Cargo.toml shows the pipeline split into named crates: cliare-measure, cliare-inspect, cliare-shape, cliare-inference, cliare-evidence, cliare-context, cliare-policy, cliare-issues, cliare-score, cliare-report, cliare-eval and cliare-guidance, with cliare-cli and cliare-app on top. That is a reasonable decomposition for this problem. Measurement is separate from inference, inference is separate from scoring, and reporting reads from the evidence store rather than re-running anything. The traversal is budgeted: the justfile sets a default profile of deep with max_depth := "12", max_probes := "5000" and concurrency := "8", which tells you the authors expect large surfaces and do not want an unbounded walk. Contexts are explicit (clean, repository, authenticated, host, fixture-backed, CI, per the README), and the README states that CLIARE reports persistent side effects from commands that look read-only, including help, version, invalid-flag and output-mode probes. That last point is the most interesting design decision here: the tool treats a help command as a suspect, not a no-op.
Installing CLIARE and running a first measurement
The README gives a Cargo install as the entry point, and the repository also ships install.sh at the top level if you prefer not to build from source. The first command below installs the binary and prints its own metadata as JSON, which is a cheap way to confirm the tool works before you point it at anything else.
cargo install cliare
cliare metadata --format jsonFor a first real run, measure one of your own CLIs. The README uses mycli as the placeholder and writes results into .cliare/mycli, with a profile and a refresh flag. The refresh flag matters: without it, you may be reading a previous run's evidence.
cliare measure mycli --out .cliare/mycli --profile standard --refreshOnce the run finishes, the summary command reads the evidence directory and prints a one-command overview. The README shows both a plain form and a JSON form; use the JSON form when you want to pipe it into something else.
cliare summary --out .cliare/mycli
cliare summary --out .cliare/mycli --format jsonTo get a usable artifact out of the run, write the command index and the harness-facing files. The README states that the harness can then load .cliare/mycli/command-index.json, .cliare/mycli/AGENT_SKILL.md and .cliare/mycli/persona-harness.json.
cliare describe .cliare/mycli --write
cliare report harness --out .cliare/mycli --writeIf you are auditing rather than documenting, the security report and the issue list are the two outputs to read first, because they are where undocumented side effects and agent-blocking gaps surface.
What the issue ledger and scorecards actually give a maintainer
The maintainer-facing output is a queue, not a grade. The README lists the categories it sorts into: missing help, confusing diagnostics, parseable-output gaps, unsafe discovery side effects, precondition blockers and command-shape drift. Each entry is described as carrying an evidence-backed assessment, its meaning for agents and harnesses, the associated commands, a suggested remedy and a verification command. That last field is the one worth checking when you evaluate the tool. A finding that ships with the command to re-verify it is a finding you can close; a finding that only says "help coverage is low" is not. The justfile cheatsheet lays out the intended workflow in ordered steps: a local dogfood check, a standard maintainer review (measure, review, issues, then drill into a report area by name such as output-contracts, then a single issue by id), a deep measurement for large surfaces, a host-state review for security, and an authenticated-context comparison. The presence of report-area and report-issue as separate verbs suggests the design assumes you will triage rather than read the whole ledger top to bottom. The README also mentions release gates for help coverage, diagnostics, parseable output and unsafe discovery behaviour, which is the CI story: the same measurement run on every release, compared against a threshold.
Command indexes are not skills, and the README is right about that
One section of the README draws a distinction that is easy to skim past: skills teach intent, workflow and policy, while a command index describes what the CLI supports right now. The argument is that agents need both, instruction for judgment and evidence for navigation, and that a skill alone will not tell a harness that a flag was renamed or that a subcommand now requires a project directory. This is the strongest conceptual claim in the project and it also defines the boundary of the tool. CLIARE cannot tell an agent what a command is for, whether calling it is a good idea in a given situation, or what the output means. It can tell the agent that the command exists, what shape it takes, what it emits and what it touched. If you were hoping for an automated skill generator, that is not this. The persona-harness.json artifact is the closest thing to it, and the README does not describe its schema, so treat it as an output to inspect rather than an interface to build against until you have looked at a real file.
Where CLIARE is the wrong tool
The probing model has costs the README does not hide but does not dwell on either. Anything that requires real credentials, a live daemon, a paid API call or a fixture set has to be arranged by you before the run; the tool measures in a context you choose, and a context you set up badly produces evidence about the wrong thing. A standard-profile run in a clean context will not see behaviour that only appears with auth state present or on the host, and the README's own examples switch to --context authenticated --auth-state present --execution-mode host --profile deep for exactly that reason. Second, a CLI with side effects that are expensive or irreversible is a poor measurement target. The tool snapshots filesystem state around probes, which is how it detects writes, but that is detection, not prevention. If a subcommand deploys, migrates or deletes, running it under a probe harness is a decision you make deliberately. Third, the whole approach assumes a released binary you can execute. A library, a service, or a CLI that only works inside a container you have not built yet is out of scope. Fourth, the repository is young: the most recent release listed is v0.1.9 from 2026-07-02, and the last push to main was on 2026-08-02. Version numbers below 1.0 in a tool that changes its output artifacts are a real consideration if you plan to parse those artifacts in CI.
OpenAPI generators, help-text parsers and the difference in evidence
The natural comparison is an OpenAPI or Swagger generator, and the README makes that comparison itself. The difference is where the specification comes from. An OpenAPI generator typically reads annotations in source code, a framework's route table, or a hand-written spec file, and emits a document describing an HTTP interface. CLIARE executes the compiled binary and infers the surface from what it observes, which means it can describe a CLI whose source you do not have, and it can catch drift between the documented surface and the shipped one. The trade-off runs the other way too: a source-derived spec is usually more complete and more stable, because it does not depend on a probe budget or a context choice. A help-text parser sits between the two. It is cheap and deterministic, but it inherits whatever --help says, which is precisely the artifact the README argues can disagree with the binary. CLIARE's position is that only execution settles the question. If your CLI is generated from a framework that already emits a machine-readable schema, you may not need the probe loop at all.
Editorial conclusion
Adopt CLIARE if you maintain a CLI whose surface changes between releases, or if you are writing a harness that has to call someone else's CLI without guessing. Skip it if your tool is a single command with no subcommands, no flags worth documenting and no side effects to catch. Before trusting a run, check the profile and context you passed: a standard-profile run in a clean context will miss daemon and auth behaviour that a host-context run would surface, and the README does not document rollback of anything the probes write.
Frequently asked questions
What license does CLIARE use?
The repository and the workspace Cargo.toml both declare Apache-2.0.
What does CLIARE stand for?
The README expands it as CLI Agent Readiness Evaluation, and notes that it is pronounced like "Claire".
How do I install CLIARE?
The README shows cargo install cliare, and the repository also contains an install.sh at the top level.
Which Rust version does CLIARE need to build?
The workspace Cargo.toml sets rust-version to 1.89 and edition to 2024.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/modiqo-cliare)