Model or dataset
ooples/token-optimizer-mcp avatar
ooples/token-optimizer-mcp

Token Optimizer MCP: a local ledger for agent context, with a knowledge graph attached

Measure token savings per AI coding agent, optimize context, and share a live local knowledge graph across 16 CLI clients.

538 stars65 forksJavaScriptMIT

At a glance

What is it?
Token Optimizer MCP is an MCP server and Claude Code plugin that intercepts expensive reads, keeps a per-project knowledge graph, and measures how many transport tokens it actually avoided. It is for teams running several CLI coding agents against the same repositories.
Who is it for?
Adopt it if you run Claude Code, Codex, Gemini or another of the sixteen supported clients against repositories where the same files get re-read every session, and you want a local, telemetry-free record of what that costs. Do not adopt it if you only use one agent occasionally, if you cannot run a Node 22+ process alongside your client, or if you need a hosted dashboard your whole team can open.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository received new commits within the last day.
What is it written in?
Mainly JavaScript, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The problem: agents pay full price for conclusions they already reached

The README frames the waste in four categories: re-reading files that have not changed, dumping an entire file to see three lines, running unbounded searches, and re-deriving conclusions from a previous session that were never written down. The last one is the expensive case. The project puts a finding at roughly 150 tokens to carry and 5k to 50k tokens to re-derive. That asymmetry is the whole argument for the tool.

The intended user is someone running one or more CLI coding agents against a repository over many sessions. The README's own example is an auth file: a session that discovers a clock-skew fix was reverted once, that per-host retry budgets were chosen over a global budget to avoid deadlock, and that git shows 47 changes in 90 days. None of that lives in the repository. It lives in a transcript that gets discarded. Token Optimizer's pitch is that the next session should not have to find it again.

How the interception works, and why the plugin matters more than the server

Two layers do the work. The MCP server exposes tools and a knowledge graph. The plugin layer is what denies calls. The README is explicit that installing the bare server only gives the model tools it can ignore, while the plugin makes a built-in Read of a 200 KB file fail, with the refusal naming the cached, diffed replacement. The same treatment applies to Grep, Glob, Edit and Write, and to cat, head and grep -r issued through the shell.

The mechanism that size-based rules cannot replicate is session diffing. Re-reading a file already read in the same session returns only a diff. A rule that says "block files over N kilobytes" will never catch a small file that was read ten minutes ago and has not changed since.

The graph is built as a side effect of work rather than by an ingestion job. Nodes represent files, symbols, tasks and findings. Edges carry typed relations: derived_from, contains, supersedes, contradicts, related. When the agent touches a relevant file, accumulated findings are fed back. There is no embedding model, no rebuild step and no query to formulate, which is a deliberate design choice: the graph is a byproduct, not a pipeline you maintain.

Attribution is per client. Returned context, optional cost equivalents and net transport avoided are grouped by operation and MCP handshake identity, so Codex, Claude Code and Gemini get separate rows. Records that predate identity capture stay explicitly unattributed rather than being assigned to whichever agent happens to be connected.

Installing the Claude Code plugin and running a first audit

For Claude Code the README says to install the plugin rather than the bare MCP server, because the plugin is what enforces the denials. The installation is three commands typed into the client:

text
/plugin marketplace add ooples/token-optimizer-mcp
/plugin install token-optimizer@token-optimizer
/reload-plugins

After the reload, the enforcement layer is active for that client. The README describes the result as the entire installation, with the other fifteen clients documented under the Installation section of the repository.

The package is also published on npm as @ooples/token-optimizer-mcp and declares Node.js 22 or newer. Its bin entries include token-optimizer-mcp, token-optimizer-daemon, token-optimizer-install, token-optimizer-doctor, token-optimizer-diagnostics and token-optimizer-proxy, and package.json defines a postinstall script plus doctor and diagnostics scripts. The README does not present an npm install as the Claude Code path, so treat the plugin route as the documented one.

Once installed, the first real use is the audit tool:

text
token_audit

The README describes the output as one ranked queue of what is costing the most per session, with each line naming how to fix it. A monthly cost equivalent appears only after you configure your own effective rate. It is deliberately not a dashboard and not six separate reports.

What the measurement actually covers, and what it excludes

This is the part to read carefully, because the project is unusually candid about its own gaps and that candour is easy to skim past. The headline number is net verified MCP transport tokens avoided: gross reduction minus later expansions, debited from the same net. In the README's captured dashboard that is 43,491 net, from 54,037 gross reduction minus a deliberate 10,546-token expansion.

Two other figures sit outside that headline. Modeled graph substitutions, 6,332 tokens of potential in the capture, are excluded while the causal graph-reuse study remains marked Collecting. The randomized control arm for downstream graph effects is also separate. A third figure, 486,074,740 historical or tool-reported tokens, is quarantined as repository scan volume that never entered model context. It is reported, and it is not counted as savings.

The dashboard reads native CLI usage receipts and prices uncached input, cache reads, cache writes and output using the captured provider, model, route, request-time tier and a versioned official source. Ambiguous model ids stay Not priced rather than receiving a blended guess. API list-price equivalents are kept separate from provider-reported charges and are never labeled as a subscription invoice. The accounting rules live in docs/TOKEN_ACCOUNTING.md.

The practical consequence: if you quote the headline number, you are quoting MCP transport only. The graph, which is the feature the project is most proud of, is not yet in the verified total.

Where it is the wrong tool

The enforcement model is the first constraint. Denying a Read, Grep, Glob, Edit, Write or shell cat is a strong intervention. If your workflow depends on an agent seeing a file in full before touching it, or if you use an agent that does not go through the supported hook path, the plugin layer will not help and the bare server gives you tools the model can ignore. The README states that plainly.

The second constraint is scope. Everything is local: no account, no telemetry, no hosted service. That is a feature for a single developer and a limitation for a team that wants a shared view. The graph is per-project and lives on the machine where it was built. The README's own capture covers 11 local projects, which tells you the unit of value is one workstation, not an organization.

The third constraint is that the strongest claim is still collecting. If your reason for adopting this is the knowledge graph rather than the read interception, you are adopting ahead of the evidence. The read-diff and denial behaviour is what the verified numbers describe. The graph reuse study is not finished.

Finally, the project is young in release terms. Version 6.0.0 landed on 2026-08-27, with 6.0.1 and 6.0.2 following on 2026-08-28. Three patch releases inside roughly a day suggests active churn. The last push to the repository was on 2026-09-10, and the repository is not archived.

How it differs from a context-compression proxy

The obvious comparison is a compression or summarization layer that sits in front of the model and shrinks whatever passes through it. That approach is content-agnostic: it reduces tokens by rewriting text, and it applies the same treatment whether the content is a stale file or a fresh one.

Token Optimizer takes the opposite route. It does not compress the payload. It refuses to make the call and substitutes a cached, diffed result. That means it can only save on calls it recognizes, and it saves nothing at all on content the agent has never seen. In exchange, the saving is structural rather than lossy: a diff of an unchanged file is not a summary of it, it is the absence of a change.

A second difference is the memory layer. A compression proxy has no notion of what the agent learned. The graph here is a separate store that persists across sessions and is keyed to files and symbols. Whether that store pays for itself is exactly the question the project has not yet answered with a randomized control arm.

Licence, maintenance and upgrade cost

The licence is MIT, and the README notes that commercial use is allowed. For an internal engineering tool that is about as permissive as it gets: no copyleft obligation on your own code, no seat counting. This is not legal advice, and if you redistribute the package or bundle it into a product you should read LICENSE in the repository yourself.

Upgrade cost is the more interesting question. The package ships a postinstall script and a set of diagnostics and doctor entry points, and it installs hooks into your CLI clients. That means upgrades touch client configuration, not just a node_modules directory. The repository provides an uninstall-hooks script and a token-optimizer-diagnostics binary, which suggests the authors expect hook state to need cleaning up.

Given three releases in two days in late August, pinning a version and reading CHANGELOG.md before moving is reasonable. The README does not document rollback, so if you need a documented downgrade path you will not find one there.

Editorial conclusion

Adopt it if you run Claude Code, Codex, Gemini or another of the sixteen supported clients against repositories where the same files get re-read every session, and you want a local, telemetry-free record of what that costs. Do not adopt it if you only use one agent occasionally, if you cannot run a Node 22+ process alongside your client, or if you need a hosted dashboard your whole team can open. Before trusting any number, run token_audit and read the token accounting contract in docs/TOKEN_ACCOUNTING.md, because the headline figure covers MCP transport only while graph substitutions and the randomized control arm are still marked Collecting.

Frequently asked questions

Does Token Optimizer MCP increase token usage?

The project measures expansions as well as reductions and debits them from the same net figure, so its verified headline is 43,491 net tokens avoided after subtracting a deliberate 10,546-token expansion. Carrying a finding into context costs roughly 150 tokens according to the README, against 5k to 50k to re-derive it.

What is token optimization in AI, as this project defines it?

Token Optimizer treats it as four specific problems: re-reading unchanged files, dumping a whole file to see a few lines, unbounded searches, and re-deriving conclusions from earlier sessions. Its answers are denial with a cached diff, session-level diffing, and a per-project knowledge graph.

What are tokens in MCP, and which ones does the dashboard count?

The dashboard counts MCP transport tokens avoided, computed as gross reduction minus later expansions. Repository scan volume that never entered model context is quarantined separately, and modeled graph substitutions are excluded from the verified headline while the causal study remains Collecting.

How can I optimize tokens in Claude Code with this project?

The README says to install the plugin rather than the bare MCP server, because the plugin is what denies expensive Read, Grep, Glob, Edit, Write and shell calls and substitutes a cached, diffed replacement. It adds three commands: /plugin marketplace add ooples/token-optimizer-mcp, /plugin install token-optimizer@token-optimizer, and /reload-plugins.

Which clients does Token Optimizer MCP support?

The README and package description both state sixteen CLI clients, with Codex, Claude Code and Gemini shown as separately attributed rows in the dashboard capture. The full list is under the Installation section of the repository.

Official sources

  1. Issues
  2. License: MIT
  3. ooples/token-optimizer-mcp on GitHub
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/ooples-token-optimizer-mcp.svg)](https://hysenlabs.com/projects/ooples-token-optimizer-mcp)