Open-source project
JetBrains/benjamin-plus-skill avatar
JetBrains/benjamin-plus-skill

benjamin-plus: A Measured Token-Efficiency Skill for Coding Agents

Benjamin-Plus: a measured token-efficiency skill for coding agents (−17.9% cost median, quality unchanged). Inject it, don't install it.

329 stars14 forksShellMIT

At a glance

What is it?
benjamin-plus is a short Markdown instruction set from JetBrains that injects five behavioral habits into a coding agent to reduce unnecessary tool calls and token usage. Measured against 80 paired SkillsBench tasks with Claude Code, it cut median cost by 17.9% and token count by 22% with no statistically significant change in task quality.
Who is it for?
benjamin-plus is worth testing for any team running Claude Code or Codex CLI at volume on tasks similar to SkillsBench. The integration is one shell command and one settings.json edit; if it does not help on your specific workload, reverting is equally simple.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 33 days ago.
What is it written in?
Mainly Shell, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What benjamin-plus teaches a coding agent

Coding agents accumulate costs in two ways: the tool calls they make, and the growing conversation context those calls append to every subsequent call. benjamin-plus addresses both by injecting five behavioral rules into the agent's context at session start.

Rule one is recon in one pass: gather all needed facts from the repository in a single combined step instead of making several separate lookups. Before adopting a code format or convention, look at two real examples rather than one. Rule two is keyhole reads: when the agent only needs to inspect a file, read 50 lines rather than the entire file; data the agent will actually transform should never be truncated. Rule three is probe the environment once: check all dependencies in a single command and install missing ones in a batch, rather than discovering them one failure at a time. Rule four is that the task's own verification check defines done: if the task specifies how to verify success, that command is the endpoint; a check that fails twice signals a wrong approach, not a symptom to retry. Rule five is that polling is a step: check a running build every 30 seconds rather than every second.

The README reports that polling alone accounted for nearly half of all agent steps on some platforms.

How the measurement was done and what the numbers mean

The README describes the evaluation as a paired A/B test. Eighty SkillsBench tasks ran with Claude Code 2.1.201 in Docker sandboxes using Sonnet 5 at low effort. The only difference between arms was the presence of the injected skill text. The primary metric was cost per task; the secondary metric was task reward (whether the agent solved the task correctly).

Median cost fell 17.9% in the treated arm. Token count fell 22%. On task quality, the result was 7 better, 5 worse, 68 ties by sign test, giving p=0.77: the quality change is not statistically significant, and the README explicitly states it was not powered as an equivalence test. Large quality effects are ruled out; small ones are not.

The same skill was also tested on 675 paired Java SWE-bench tasks using Codex CLI and gpt-5.6-luna. In that setting, cost fell 4.4% (confidence interval -7.5% to -1.5%, p=0.003) and tool calls fell 20%, with solve rate unchanged (p=0.22).

The README notes that savings scale with baseline bloat. A leaner-running baseline measured -10% median cost in an earlier run; the treated arm stayed flat while the control drifted upward by 10.5%. Expect roughly 10% to 18% depending on session behavior.

Injecting the skill into Claude Code or Codex CLI

Clone the repository to a local path:

bash
git clone https://github.com/JetBrains/benjamin-plus-skill ~/.benjamin-plus

For Claude Code, add a SessionStart hook to ~/.claude/settings.json:

json
{ "hooks": { "SessionStart": [ { "matcher": "startup|resume|clear|compact",
  "hooks": [ { "type": "command", "command": "cat ~/.benjamin-plus/injected-instruction.md" } ] } ] } }

For a per-project installation without a hook, append the instruction to CLAUDE.md in the project root:

bash
cat ~/.benjamin-plus/injected-instruction.md >> CLAUDE.md

For Codex CLI, AGENTS.md is loaded automatically into every session, so appending there requires no hook:

bash
cat ~/.benjamin-plus/injected-instruction.md >> ~/.codex/AGENTS.md

For any other agent, append the contents of injected-instruction.md to the system prompt. The README states this is the whole integration, approximately 3KB.

The README is explicit about a finding from testing: as a discoverable skill folder, the rules saved nothing (-0.5%, not significant). Agents burned steps just finding SKILL.md. Injection is the method that produces the measured savings.

What the skill does not cover and where it may underperform

The five rules address lookup and waiting behavior. They do not change what the agent builds, how it structures code, or which tools it calls for functional work. A task where most of the cost comes from writing and testing code rather than from exploratory file reads will see smaller gains.

The README is transparent about the limits of the measurement. The benchmark is SkillsBench (80 tasks) and Java SWE-bench (675 tasks), both of which favor tasks similar to typical software engineering work. The saved steps from polling (nearly half of all steps on some platforms) will not appear on a benchmark that does not involve long-running builds.

The README also notes that medians are the honest unit: a few hard-task tails can give back part of the aggregate savings. If your agent sessions are predominantly long, complex tasks where each tool call has high informational value, the behavioral rules may conflict with necessary lookups rather than eliminate wasteful ones.

benjamin-plus compared to writing your own agent instructions

Writing custom instructions for a coding agent directly in CLAUDE.md or AGENTS.md is the most common alternative. A team can encode project-specific conventions, tool preferences, and workflow rules that are more targeted than a generic efficiency skill.

The practical difference is that generic efficiency rules like those in benjamin-plus can be adopted immediately without knowing anything specific about the project, while custom instructions take time to write and test. The README suggests using both: appending injected-instruction.md to CLAUDE.md means the efficiency rules combine with any project-specific rules already in that file.

A custom agent instruction file can also be tuned to specific workloads and updated as the team learns what the agent does inefficiently. benjamin-plus is a fixed rule set developed against one benchmark; it will not adapt to project-specific waste patterns that do not appear in SkillsBench. The README encourages users to open issues with regression reports and paired numbers if the skill loses money on a specific workload.

The repository is licensed under MIT and the last push was on 2026-08-27.

The SHA256SUMS.txt file in the repository lets users verify the integrity of the injected-instruction.md file before appending it to any agent configuration.

Editorial conclusion

benjamin-plus is worth testing for any team running Claude Code or Codex CLI at volume on tasks similar to SkillsBench. The integration is one shell command and one settings.json edit; if it does not help on your specific workload, reverting is equally simple. The README makes the measurement methodology and its limits explicit: results came from 80 paired tasks on a specific benchmark with Sonnet 5, and the savings vary from roughly 10% to 18% depending on baseline session bloat. Teams with tightly tuned agents may see smaller gains than teams running verbose unoptimized sessions.

Frequently asked questions

What does benjamin-plus inject into a coding agent?

It injects five behavioral rules from the injected-instruction.md file: gather facts in one pass, read only the lines needed, probe dependencies in one command, treat the task's own verification check as the endpoint, and poll builds at 30-second intervals rather than continuously. The full text is in RULESET.md and is approximately 745 tokens.

Are the token savings from benjamin-plus consistent across all workloads?

The README states that savings scale with baseline bloat. The measured range is roughly 10% to 18% median cost reduction depending on how inefficient the agent's default session behavior is. Tasks where most cost comes from functional work rather than exploratory lookups will see smaller gains.

Does benjamin-plus work with agents other than Claude Code?

The README documents integration for Claude Code (settings.json hook or CLAUDE.md append), Codex CLI (AGENTS.md append), and any other agent via system prompt injection. Cross-platform testing was done on Java SWE-bench with Codex CLI and gpt-5.6-luna, where cost fell 4.4% with p=0.003.

Official sources

  1. Issues
  2. JetBrains/benjamin-plus-skill on GitHub
  3. License: MIT
  4. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/jetbrains-benjamin-plus-skill.svg)](https://hysenlabs.com/projects/jetbrains-benjamin-plus-skill)