# token-diet has to hack the loader, because a skill body only loads on demand

> token-diet is a shell installer plus a handful of markdown files that push a concision directive into every session of a coding agent, across Claude Code, Codex, Cursor, Windsurf and Cline, at four selectable levels. The claims are specific and the README is careful with them, including the note that the 31 percent average is unweighted and that the 54 percent figure is best case rather than typical.

**Kulaxyz/token-diet** — Always-on token-efficiency skill for coding agents (Claude Code, Codex, Cursor, Windsurf, Cline). ~31% lower bill on average, no loss of correctness.

- Repository: https://github.com/Kulaxyz/token-diet
- Stars: 472 · Forks: 4
- Language: Shell
- License: not declared
- Published: 2026-09-18 · Updated: 2026-09-18 · Language: en
- Canonical page: https://hysenlabs.com/projects/kulaxyz-token-diet

## A skill only loads when asked, so the installer hijacks the loader

The mechanism is the interesting part, and it comes last in the README because it is the least like a feature.

A skill body in these coding agents loads on demand, which means a skill that only loads when invoked is not an always-on skill. So to make the directive apply to every session, the installer injects it through a channel that fires every time. For Claude Code that is a `SessionStart` hook. For Codex it is the `AGENTS.md` file, which is always loaded. For Cursor, Windsurf and Cline it is their rule files, which are also always loaded.

So the same text reaches five agents through three different mechanisms, and the tool is honest that this is a trick, crediting another project for using the same approach. What you install is not a plugin with a manifest and a capability list. It is text placed where the host will read it every session.

That has a practical consequence before you install anything. The value of this project depends entirely on writing into files that belong to other tools, and the repository contains six entries: a README, a SKILL file, two activation files for the default and ultra levels, a benchmark directory and the install script. There is no license file, no CI configuration and no test suite in the tree.

The last push to main is dated 4 July 2026, and there are no tagged releases, so the version you install is whatever the branch holds.

## Concision applies to output and never to reasoning

The rules are grouped into eight areas, and the eighth one is the one that makes the rest defensible.

Replies lead with the answer, with no preamble of the sort that starts with a pleasantry and no postamble offering further help, and they report deltas rather than narration. Docs, memory files, hand-offs, plans and comments get the minimum words that still say everything, with a rule about commenting the non-obvious reason rather than the code. Code gets built only when asked, stays idiomatic, carries no dead branches and is never cryptic. Context work greps before it reads, reads the lines it needs rather than whole files, batches reads, and never re-reads a file it has just edited.

Then the guardrail line, which is the one to check before you install: concision applies to output and never to reasoning. Correctness is off-limits, so is critical test coverage, and so are verbatim code, commands and errors. In other words the instruction set explicitly refuses to trade away the things you would not want traded away.

The test rules are the most concrete version of that. Only key and critical or edge paths, grouped, at most ten per session, and never skip money, auth or data loss paths. Ten tests is a number a reviewer can argue with, which is the point: a rule you can check is a rule that can be corrected, and it also means a session that needs thirty tests is a session where the rule is wrong rather than the tests being wrong.

## Four levels, and only the chat turns become telegraphic

There is a level on the command line and a matching slash command, and the four settings are not simply more or less strict.

`on` is the default and applies everything above. `lite` is narrowed to communication and artifacts. `ultra` is the one with a scope limit rather than a severity limit: the chat and progress reporting become telegraphic, while code, tests and docs stay precise. And `off` turns the directive off.

The slash command takes the same four values, `/token-diet [on|lite|ultra|off]`, which means the level can be changed mid session without reinstalling. That matters more than it sounds, because ultra is the level where you would want it off for a code review and on for a long planning conversation.

The installer scopes the install in two other ways. A target flag names one agent or all of them, and the accepted values are claude, codex, cursor, windsurf, cline, all and print, where the last is a way to see what would be done rather than do it. And a project flag installs into the current repository instead of globally, which is the difference between a rule that follows you everywhere and a rule a team can commit and share.

There is also an uninstall flag, which for a tool that writes into other tools' configuration is the flag you most want to know exists.

## The 102 to 34 token example is the whole pitch

The before and after section is one answer rewritten twice, and the token counts are measured rather than estimated, against the o200k_base tokenizer.

The normal version runs to 102 tokens. It opens with a pleasantry, explains that Stripe computes the signature over the exact raw bytes of the request so passing an already parsed JSON body into the event constructor fails, suggests reading the raw body instead with a specific method in a specific framework, and then offers to write a code snippet.

The ultra version runs to 34 tokens, a 66 percent cut. It keeps the constructor name, the raw bytes point, the method, and the file path, and drops the rest.

The line under the example is the claim that matters: same fix, same identifiers, same path, only the filler is gone. That is a narrower claim than being smarter or being right, and it is the one this tool can actually support, because a directive about output style cannot change which function the model decides to call.

The reductions in that example are larger than the headline average, and the README does not pretend otherwise. The average for a bill across its own benchmark runs is about 31 percent, with a range from 17 to 54 depending on session type, and output savings running from 30 to 81 percent.

## The 31 percent is an unweighted mean of three scenarios

The numbers section is three rows and four sentences, and the sentences are the reason the numbers are usable.

The scenarios are output heavy work such as advice, planning and explanation, which shows 81 percent less output and a 54 percent lower bill; a code change with tests run against a large real repository of 1673 files, at 49 percent output and 22 percent bill; and read heavy comprehension, at 30 percent output and 17 percent bill. The average row shows 53 percent output and 31 percent bill.

Then the caveats. The average is unweighted across the three scenarios, and the 54 percent is explicitly called best case rather than typical. Correctness is stated as having held in every run. And the mechanism behind the spread is explained in one line: output savings are consistent, while the bill win depends on how much of the session is output versus unavoidable file reading.

That last sentence is the honest version of the product. A session that is mostly the agent reading your code cannot be made much cheaper by making the agent talk less, so the ceiling on the bill saving is set by the shape of the work rather than by the discipline. The worst case in the table, 17 percent, is the number to plan around.

The method is not hidden. Full tables are in the benchmark results file, and the whole thing can be re-run with a command that takes an Anthropic API key and a node script. Running your own sessions through it is how you get a number that applies to your work rather than to the author's.

## Install is a curl pipe with two documented shapes

The install section shows the command twice, and the second copy is the one that matters.

```bash
# one-liner — auto-detects Claude Code, Codex, etc.:
curl -fsSL https://raw.githubusercontent.com/Kulaxyz/token-diet/main/install.sh | bash

# pass options through `bash -s --`, e.g. the telegraphic level:
curl -fsSL https://raw.githubusercontent.com/Kulaxyz/token-diet/main/install.sh | bash -s -- --ultra
```

The second form exists because a pipe into bash normally swallows its arguments, so the extra `bash -s --` is what lets the ultra level through. Anyone who writes the first form and then wonders why the level did not change has their answer in that detail.

The alternative is to clone the repository and run the script directly, which is the route to take if you want to read it first, and this is a tool that writes into other tools' configuration, so reading the script is not a paranoid move. The flags the README names are the telegraphic level, the uninstall, the target agent with its list of values including a print mode, and the project scope.

Auto-detection is the default, so the script looks for the installed agents and writes to the ones it finds. That is convenient and it is also the part with the largest blast radius: the tool is deciding which of your editors to modify on the first run.

## Conclusion

token-diet fits someone who reads a lot of agent output, is annoyed by preamble and postamble, and wants a rule installed once rather than a reminder repeated in every prompt. It does not fit someone who reads a large codebase through an agent, since a read heavy session is where the percentage win is smallest, and it does not fit anyone who needs the directive to stay out of their agent's configuration. Before running the installer, read what it writes, because a tool whose value depends on injecting a hook into another tool's config deserves a look at the script first, and note that the repository has no licence file and no release history.

## FAQ

### What does the token-diet skill actually change in an agent's behaviour?

It installs a directive that covers eight areas: replies drop preamble and postamble, docs and comments drop to the minimum words that still say everything, code is built only when asked, context work greps before reading and never re-reads a just edited file, independent tool calls are batched, and broad bounded search is delegated to a cheaper model. Concision applies to output and never to reasoning.

### How does token-diet stay on in every session?

A skill body loads on demand, so the installer injects the directive through a channel that fires every session instead: a SessionStart hook for Claude Code, an always-loaded context file such as Codex AGENTS.md, and the rule files for Cursor, Windsurf and Cline. The same text reaches five agents through three mechanisms.

### How much does token-diet save?

The README reports about 31 percent lower bill on average with output savings from 30 to 81 percent, across three scenarios on Sonnet 5 runs, and says correctness held in every run. It also says the average is unweighted across the three scenarios, that the 54 percent figure is best case rather than typical, and that the bill win depends on how much of a session is output versus unavoidable file reading.

### What are the levels in the token-diet slash command?

Four: on, which is the default and applies everything, lite, which covers communication and artifacts only, ultra, which makes chat and progress reporting telegraphic while code, tests and docs stay precise, and off. The level can be changed in a session with the slash command without reinstalling.

### How do I install token-diet for one agent or one project?

The one-liner pipes install.sh into bash and auto-detects installed agents. A target flag with the value claude, codex, cursor, windsurf, cline, all or print selects one, a project flag installs into the current repository instead of globally, and an uninstall flag removes it. To pass options through the pipe, the script needs the extra bash -s -- so the argument is not swallowed.

## Sources

- [Issues](https://github.com/Kulaxyz/token-diet/issues)
- [Kulaxyz/token-diet on GitHub](https://github.com/Kulaxyz/token-diet)
- [README](https://github.com/Kulaxyz/token-diet/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/kulaxyz-token-diet
