Model or dataset
caliber-ai-org/ai-setup avatar
caliber-ai-org/ai-setup

Caliber keeps agent context files current, and scores them without a model

Continuously sync your AI setups with one command. Codebase tailor suited agent skills, MCPs and config files for Claude Code, Cursor, and Codex.

1,295 stars125 forksTypeScriptMIT

At a glance

What is it?
An MIT CLI that writes CLAUDE.md, Cursor rules, AGENTS.md and Copilot instructions for your repository, keeps them refreshed by a pre-commit hook, and audits them with deterministic checks rather than a language model. Its newest feature, Jev compaction, is TypeSafe's technology and needs a TypeSafe or AI Gateway key to run.
Who is it for?
Caliber fits a team that has already committed to agent context files and finds them stale within weeks, since the pre-commit hook plus the no-model audit address exactly that. Check what you are buying before you rely on the compaction feature, because Jev is TypeSafe's scoring technology reached through your own AI Gateway or TypeSafe key, and it falls back to built-in compaction when unavailable rather than failing loudly.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 9 days ago.
What is it written in?
Mainly TypeScript, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 3, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The files agents read, refreshed by a pre-commit hook

The problem being addressed is that an agent's context files go stale. You write CLAUDE.md describing how the repository works, and two months later the paths it names have moved, the commands no longer exist, and the file is worse than nothing because it is confidently wrong. Caliber's answer is to generate those files and then keep them honest automatically.

Setup is one command:

bash
npx @rely-ai/caliber bootstrap

After that, in Claude Code or the Cursor CLI you run the /setup-caliber command, and the agent writes CLAUDE.md, Cursor rules, AGENTS.md and Copilot instructions, shows you the diff, and installs hooks that keep them current. If you have no agent to hand, `caliber init` does the same job through a wizard.

The hook is the load-bearing part. It reads your diff at commit time and refreshes the context files your agents read, so the files track the code instead of trailing it. A separate `caliber refresh` command updates configs from recent changes on demand, and `caliber hooks` manages the refresh and sync hooks when you want to see or change them.

Requirements are modest: Node 20 or newer. It runs on your existing Claude Code or Cursor seat, or against your own Anthropic, OpenAI, MiniMax or Vertex key. Bootstrap, scoring and sync are described as running fully local, so the key matters for the parts that call a model rather than for setup.

caliber score audits without calling a model

The audit command is the part that distinguishes this from the many tools that offer to review your agent configuration. caliber score checks the context files against the repository with no LLM involved, answering three questions: do the paths exist, are the commands real, has anything drifted.

That is a deterministic check, and it is a meaningfully different proposition from asking a model whether the instructions look right. A model review tells you the file is plausible. This tells you a directory was renamed. The comparison mode points it at a branch:

bash
caliber score --compare main

The commands table describes score as auditing configuration quality with no LLM, sitting alongside refresh, which updates configs from recent changes. So the division of labour is deliberate: refresh is the mutating command driven by the hook, score is the read-only verification you can run in review or in continuous integration without spending anything or shipping your repository to a provider.

That last property is the one to lean on. A config audit in CI that costs no tokens and sends nothing anywhere is easy to adopt, because it costs nothing to be wrong about and nothing to run. If your context files have rotted, this is what tells you.

Sync degrades instead of failing

Write a skill, rule, plugin or MCP server in one agent and `caliber sync` puts it into the others. The README describes the operation as deterministic, and makes two guarantees that matter more than the copying.

The first is that it degrades where an agent lacks a feature, with the concrete example that a skill becomes a Copilot instruction file. So the five supported targets do not need to be feature-equivalent, and a GitHub Copilot user without a skills mechanism still ends up with something usable rather than an error. The second is that it never overwrites a file you edited by hand, which is the guarantee that makes it safe to leave enabled.

That second property is doing a lot of work. Any tool that writes into your configuration directory is a candidate for clobbering work you did deliberately, and the usual answer is documentation telling you to back things up. Here the behaviour itself is the answer.

Status is available as a command:

bash
caliber sync --status

What sync does not do is reconcile. It mirrors outward from the source agent, and the routing between a plugin, an MCP server, a rule and a skill is decided per target format. There is a related command, `caliber learn`, which learns patterns from your sessions into CALIBER_LEARNINGS.md, which is the piece that captures knowledge rather than configuration.

Every write is shown as a diff, and backed up

The safety story has three layers and they are worth separating. Every write is displayed as a diff before it is applied, so nothing changes without being shown. Every write is backed up into a .caliber/backups directory, so a bad change is recoverable rather than merely visible. And telemetry, which is anonymous and described as collecting command names but never code, can be turned off by setting CALIBER_TELEMETRY_DISABLED=1.

The command set backs this up with an explicit escape hatch: `caliber undo` reverts changes and `caliber uninstall` removes everything Caliber added. Having both, with uninstall described as removing everything the tool added rather than just unlinking a bootstrap, is the behaviour you want from something that writes into four different agents' configuration directories.

The remaining commands fill in the loop. `caliber plugin install` installs the Jev plugin into the current project. `caliber config` chooses provider, key and model, which is where the Anthropic, OpenAI, MiniMax and Vertex options are selected. `caliber compact` produces a Jev compaction report, with a --write flag for generic transcripts. And `caliber learn` writes learned patterns into CALIBER_LEARNINGS.md.

Taken together the ten commands describe a lifecycle: set up, verify, refresh, share, compact, learn, manage hooks, configure, and undo. That last pair is the one most tools in this category leave out.

The compaction feature needs TypeSafe's key

The newest feature is described as Jev compaction, credited to TypeSafe, and it is worth reading closely because the licensing story is not what the rest of the project implies. Everything above runs under MIT with no vendor required. This does not.

The motivation is a specific and reasonable complaint about /compact: summaries forget the file path, the error, and the constraint you gave an hour ago. Jev's approach is to score every tool call, drop the stale ones, and keep everything else byte-for-byte, which is a much stronger guarantee than summarisation.

Installing it in Claude Code means enabling function hooks, exporting a key, and adding the plugin from this repository as a marketplace:

bash
export CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1
export AI_GATEWAY_API_KEY=...        # or TYPESAFE_API_KEY
claude plugin marketplace add caliber-ai-org/ai-setup
claude plugin install caliber-jev-compaction@caliber

For Cursor, Codex and anything else the same scoring runs from the CLI against a transcript:

bash
caliber compact --provider cursor  --transcript <jsonl>
caliber compact --provider generic --transcript ./session.json --write

Two things follow. You bring your own Vercel AI Gateway or TypeSafe key, so the open-source project routes you to a commercial service for its most interesting capability. And when Jev is unavailable, Claude Code falls back to its built-in compaction, so the failure mode is a quieter experience rather than an error. Setup and verification for the feature are in FUNCTIONAL_CHECKLIST.md rather than the main README, which is where the operational detail went.

Five agents in the header, three in the description

The support list is inconsistent between the metadata and the documentation. The repository description names Claude Code, Cursor and Codex. The README header lists Claude Code, Cursor, Codex, OpenCode and GitHub Copilot. The body of the README then talks about Copilot instruction files as a degraded target and Cursor as a first-class one, which supports the longer list.

OpenCode appears only in the header line and nowhere else in the visible documentation, so whether it has first-class support or degradation handling is unclear from what is written.

The naming is similarly split three ways. The GitHub organisation is caliber-ai-org, the npm scope is @rely-ai, and the binary the package installs is plain caliber. The plugin marketplace name used in the install command is caliber, and the marketplace source is this repository. None of that is wrong, but if you are pinning the tool in a script or a Dockerfile, the string you need is @rely-ai/caliber for the package and caliber for the command, and the two are not derivable from each other.

The project is MIT licensed, has 1,293 stars, 126 forks and 19 open issues, and the last push to main was on 2026-09-24. The release tags track closely: v1.55.0 and v1.55.1 both on 2026-09-19, then v1.55.2 on 2026-09-24, with the tag published a second after that push.

The repository carries an action.yml nobody documents

The top level holds directories that map onto the feature set, and one file that the README does not mention at all. There are .claude/, .cursor/ and .agents/ directories, which are the per-agent outputs, and a .context/ directory that appears to be the canonical copy all of them are generated from. A plugin/ directory holds the Jev plugin source, which is compiled separately from the CLI with its own tsconfig. skills/ and scripts/ hold the skill definitions and the generators that produce them.

The undocumented file is action.yml, a GitHub Actions workflow definition. Its presence means there is a CI integration for running Caliber in a pipeline, and the score command is the natural fit given that it costs nothing and sends nothing out. It is not mentioned in the README, in the commands table, or in the contributing notes, so treat its contents as unknown until you read it.

Three other root files describe the project rather than build it: TODOS.md for the maintainers' own backlog, LANDING_HERO.md for marketing copy, and FUNCTIONAL_CHECKLIST.md for the Jev setup and verification. Alongside them sit the conventional files, AGENTS.md, CLAUDE.md, CHANGELOG.md, CONTRIBUTING.md, CODE_OF_CONDUCT.md and SECURITY.md. A .husky directory and an index.cjs file sit next to a package.json declaring type module, so the CommonJS entry point is a deliberate compatibility shim rather than a leftover.

Generated skills and a live Jev test sit in the build gate

The build script is where the project's discipline shows. Running prebuild executes build:skills:check and then build:plugin:check, and both are check modes of generators that have a write mode. build:skills runs generate-skills.mjs; build:skills:check runs the same script with a flag that verifies instead of writing. build:plugin and build:plugin:check are the equivalent pair for the Jev plugin, and validate:plugin runs a validation script followed by the plugin check.

So skills and plugin assets are generated artefacts with a drift check wired into prebuild, and a stale generator cannot be merged silently. The main build then runs tsup, copies the hook runner, and builds the plugin. prepare runs husky, which installs the git hooks on install, and lint, format, typecheck, test and test:coverage are all separate scripts using eslint, prettier, vitest and tsup.

The test scripts are worth a look. Alongside test and test:watch there are two end-to-end targets, e2e:jev and e2e:gateway, pointing at files named jev-live.test.ts and jev-gateway-live.test.ts. Those names say they are live tests, meaning they contact the real service, which is consistent with the feature needing a key. So the compaction path has real integration coverage rather than only mocked unit tests, and the cost of that is that part of the suite is not runnable without credentials.

The runtime dependencies explain the provider options: the Anthropic SDK, the Anthropic Vertex SDK with google-auth-library for Vertex, and the OpenAI SDK, plus the inquirer set for the wizard, commander for the CLI, chalk, diff and glob.

Contributing is one line, which is a good sign for a project that expects patches:

bash
git clone https://github.com/caliber-ai-org/ai-setup.git && cd ai-setup
npm install && npm run test

Editorial conclusion

Caliber fits a team that has already committed to agent context files and finds them stale within weeks, since the pre-commit hook plus the no-model audit address exactly that. Check what you are buying before you rely on the compaction feature, because Jev is TypeSafe's scoring technology reached through your own AI Gateway or TypeSafe key, and it falls back to built-in compaction when unavailable rather than failing loudly. Confirm the sync degradation rules match how your agents store skills, because a skill becoming a Copilot instruction file is a downgrade rather than a translation. And note the package is published under the @rely-ai scope while the repository sits under caliber-ai-org, which matters if you pin it in automation.

Frequently asked questions

What does Caliber write into my repository?

After npx @rely-ai/caliber bootstrap you run /setup-caliber in Claude Code or the Cursor CLI, and the agent writes CLAUDE.md, Cursor rules, AGENTS.md and Copilot instructions, shows you the diff, and installs hooks that keep the files current. caliber init does the same as a CLI wizard.

Does the Caliber score command use an LLM?

No. caliber score checks the context files against the repository with no LLM involved, verifying that paths exist, commands are real and nothing has drifted. caliber score --compare main runs it against a branch.

Will Caliber sync overwrite a file I edited by hand?

No. The documentation states that sync never overwrites a file you edited by hand. It also degrades where an agent lacks a feature, so a skill can become a Copilot instruction file rather than failing.

Does Caliber's Jev compaction work without a paid service?

You bring your own Vercel AI Gateway or TypeSafe key, since Jev is TypeSafe's scoring technology. If Jev is unavailable, the Claude Code plugin falls back to its built-in compaction.

How do I uninstall Caliber?

caliber undo reverts its changes and caliber uninstall removes everything Caliber added. Every write is shown as a diff first and backed up into .caliber/backups, and anonymous usage analytics can be disabled with CALIBER_TELEMETRY_DISABLED=1.

Official sources

  1. caliber-ai-org/ai-setup on GitHub
  2. License: MIT
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/caliber-ai-org-ai-setup.svg)](https://hysenlabs.com/projects/caliber-ai-org-ai-setup)