Model or dataset
drona23/claude-token-efficient avatar
drona23/claude-token-efficient

claude-token-efficient: a single CLAUDE.md that shortens Claude Code output

One CLAUDE.md file. Keeps Claude responses terse. Reduces output verbosity on heavy workflows. Drop-in, no code changes.

6,074 stars472 forksPythonMIT

At a glance

What is it?
drona23/claude-token-efficient is one instruction file you drop into a project to cut Claude's preamble, restatement and over-engineering. It saves output tokens only when output volume is high, and it costs input tokens on every message.
Who is it for?
Adopt claude-token-efficient if you run Claude Code in output-heavy, repeated workflows where terse, parseable answers matter more than discussion, and if you are willing to measure your own output-token change rather than trust the headline figure.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 109 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 3, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What claude-token-efficient actually changes about a Claude Code session

The project targets a specific annoyance: the words Claude produces around the answer. According to the README, the default behaviour includes opening with "Sure!" or "Great question!", closing with "I hope this helps!", restating your question before answering, adding unsolicited suggestions, and agreeing with incorrect statements. Each of those costs output tokens and none of them carries information you asked for.

The fix is not code. It is a CLAUDE.md file placed at the root of your project, which Claude Code reads as context. The README frames the trade-off plainly: the file adds input tokens on every turn, and the savings come from reduced output tokens, so the net is positive only when output volume is high enough to offset the persistent input cost. That single sentence is the whole decision. The project is for people running resume bots, agent loops, code generation pipelines and other repeated structured tasks where verbosity compounds across hundreds of calls. It is not for someone asking one question a day.

Rules in chat versus the CLAUDE.md file: two ways to run the same idea

The README offers two options rather than one. The quick version is a line pasted into a session:

code
Rules: Read files first. Write complete solution. Test once. No over-engineering.

That works immediately with no setup and suits one-off tasks. The second option is dropping the CLAUDE.md file into the project tree, which the README shows as a single entry alongside your source. The file applies automatically on every message, which the README calls better for regular work and more efficient at scale.

The comparison table in the README lists setup as none versus one file, cost as higher versus lower, and best-for as quick sessions versus regular work and pipelines. Note what the table does not claim: the pasted rules are not described as worse in quality, only as higher cost because you re-supply them each time. The repository layout reflects the same split, with RULES.md, a profiles/ directory and an examples/before-after.md alongside the root CLAUDE.md.

Installing claude-token-efficient and a first real use

There is no package to install and no build step. The README describes the file as drop-in with zero setup and no code changes, so installation means placing CLAUDE.md where Claude Code will read it. Clone the repository and copy the file into your project root:

bash
git clone https://github.com/drona23/claude-token-efficient
cp claude-token-efficient/CLAUDE.md your-project/CLAUDE.md

Your project tree should then show the file at the top level, matching the layout in the README:

code
your-project/
└── CLAUDE.md

After that, start a session as usual. The README says the rules apply automatically on every message, so the observable change is in the shape of the answers: no preamble, no closing pleasantries, no restating of your question. If you want the stricter rules that the README associates with the 63% word-reduction figure, the repository keeps alternative rule sets under profiles/ and RULES.md, which you would copy in place of the root file.

To reproduce the measured numbers rather than the word-count table, the README points at the benchmark directory and this command:

bash
python3 benchmark/run.py -n 5 --model opus

That runs five iterations against the named model and writes a per-model report; the README directs you to benchmark/SUMMARY.md and benchmark/SEMANTIC.md for the output-token results.

The benchmark numbers do not say what the headline says

The README leads with a 63% reduction across five prompts, from 465 words to 170. It then qualifies that heavily in the methodology note: five prompts, one run each, no variance controls, and word counts rather than tokens. The note calls it a directional indicator, not a controlled study.

The repository's own later benchmark contradicts the headline for the current file. Measured output-token reduction with the minimal CLAUDE.md is about 4% on haiku, 12% on sonnet and 7% on opus. The README states that the 63% figure is reachable only with a stricter rules profile. So the number a reader remembers is the one the project's own measurement does not support for the default configuration.

The semantic evaluation adds a second, sharper finding: baselines on current models already show 0% preamble, sycophancy, "as an AI" and smart quotes. Rules aimed at those behaviours therefore add input cost without changing output. That is a real limitation of the rule set as shipped, and the README's own advice is to trim accordingly. Anyone adopting this should read the semantic report before assuming the rules are earning their context.

Where claude-token-efficient is the wrong tool

The README is unusually direct about the cases where the file loses. Single short queries lose, because the file loads into context on every message and on low-output exchanges it is a net token increase. Casual one-off use loses for the same reason. Pipelines that start a fresh session per task lose, because the README says fresh sessions do not carry the CLAUDE.md overhead benefit the same way persistent sessions do.

Two exclusions matter more than the rest. The first is parser reliability: if you need guaranteed parseable output, the README says to use structured outputs built into the API, such as JSON mode or tool use with schemas, and calls that a more robust solution than prompt-based formatting rules. A markdown instruction file cannot enforce a schema. The second is deep failure modes. Hallucinated implementations and architectural drift are not fixed here; the README states those require hooks, gates and mechanical enforcement. A prompt cannot gate anything.

The exploratory case is the subtlest. The README notes that the override rule lets you ask for debate and alternatives at any time, but that if exploratory or architectural work is your primary workflow, the file will feel restrictive. The rules bias toward terse answers, and terse is the opposite of what design discussion needs.

How it compares with structured outputs and with leaving Claude alone

The closest alternative is not another CLAUDE.md. It is the API's own structured output features. The difference in approach is enforcement versus suggestion. A schema in JSON mode or a tool definition constrains what the model can emit, so a malformed response fails loudly. claude-token-efficient asks the model to be brief in prose, and the README itself concedes that prompt-based formatting rules are the weaker option when parseability is the requirement.

The other alternative is doing nothing and accepting the default verbosity, then controlling cost by choosing a smaller model or reducing the number of calls. That is a reasonable position for low-volume users, and the README's own arithmetic supports it: the at-scale table estimates roughly 9,600 tokens saved per day at 100 prompts per day, about $0.86 per month on Sonnet, rising to about $25.92 per month across three projects at 1,000 prompts per day. Those are small absolute sums. If your usage sits near the low end, the persistent input cost is not worth chasing.

What distinguishes this project from both alternatives is that it is a file, not a dependency. There is nothing to version-pin, nothing to break at runtime, and nothing to remove beyond deleting one markdown file.

Maintenance cost, licence and what the repository does not document

The repository is MIT licensed, which permits commercial use, modification and redistribution provided the copyright notice and permission notice are retained. That is the standard permissive arrangement; if you redistribute the file inside a product, keep the LICENSE text with it. Nothing here constitutes legal advice, and the LICENSE file in the repository is the authoritative text.

Maintenance is light by construction. The last push to the default branch was on 2026-06-16, and the repository is not archived. The artefact is a markdown file, so there is no dependency graph, no runtime and no upgrade path in the usual sense. The real maintenance burden falls on you: the README warns that instruction files add input tokens on every turn and that if the file grows too much it can cost more than it saves. Keeping it short is an ongoing editorial job, not a one-time install.

Several things are not documented. The README does not describe a rollback procedure beyond removing the file, does not give a versioning scheme for the rule sets, and does not state how the profiles/ directory relates to the root CLAUDE.md in terms of precedence. Model support is explicitly limited: benchmarks were run on Claude only, and the README says local models such as llama.cpp and Mistral are untested. There are no retrieved releases, so there is no changelog to consult for behavioural changes between rule sets.

Editorial conclusion

Adopt claude-token-efficient if you run Claude Code in output-heavy, repeated workflows where terse, parseable answers matter more than discussion, and if you are willing to measure your own output-token change rather than trust the headline figure. Do not adopt it for single short queries, casual one-off use, pipelines that start a fresh session per task, or any case where you need guaranteed machine-parseable output, because the README points at structured outputs (JSON mode, tool use with schemas) for that instead. Before committing, verify three things: the current CLAUDE.md line count, since the README warns that a file which grows too much can cost more than it saves; the per-model numbers in benchmark/SUMMARY.md, since the README reports roughly 4% reduction on haiku, 12% on sonnet and 7% on opus with the current minimal file; and whether your workflow is exploratory, because the README states that if debate and alternatives are your primary mode, the file will feel restrictive.

Frequently asked questions

How do I use claude-token-efficient to reduce Claude Code token usage?

Copy the repository's CLAUDE.md into your project root, where Claude Code reads it as context on every message. The README says the rules then apply automatically with no code changes. For a one-off session you can instead paste the short rules line into the chat.

Is Claude Code more token efficient with claude-token-efficient installed?

Only when output volume is high. The README states the file adds input tokens on every message while savings come from reduced output tokens, so on single short queries it is a net token increase. Its own measured reduction with the current minimal file is about 4% on haiku, 12% on sonnet and 7% on opus.

Which Claude model does claude-token-efficient save the most tokens on?

The README reports the largest measured output-token reduction on sonnet, at roughly 12%, against about 4% on haiku and 7% on opus. It does not claim a most efficient model overall, since the benchmark measures the effect of the instruction file rather than comparing models on their own.

How does Claude token consumption work with an instruction file like claude-token-efficient?

The CLAUDE.md file is loaded into context on every turn, which adds input tokens, while the terse rules reduce the tokens Claude generates as output. The README frames the net result as positive only when output volume is high enough to offset that recurring input cost.

How do I make Claude token efficient with claude-token-efficient?

The README gives a short rules line for pasting into a session and a CLAUDE.md file for automatic application on every message. It also keeps stricter rule sets under profiles/ and RULES.md for cases where the minimal file does not reduce output enough.

Official sources

  1. drona23/claude-token-efficient on GitHub
  2. Issues
  3. License: MIT
  4. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/drona23-claude-token-efficient.svg)](https://hysenlabs.com/projects/drona23-claude-token-efficient)