Model or dataset
headroomlabs-ai/headroom avatar
headroomlabs-ai/headroom

Headroom: compressing tool output and JSON before it reaches the model

Compress tool outputs, logs, files, and RAG chunks before they reach the LLM. 20% fewer tokens for coding agents, 60-95% fewer tokens for JSON, same answers. Library, proxy, MCP server.

74,088 stars5,723 forksPythonApache-2.0

At a glance

What is it?
Headroom is an Apache-2.0 Python project that compresses tool outputs, logs, files and RAG chunks before they reach an LLM, available as a library, a local proxy and an MCP server. The savings on JSON are large, but the reversible cache and the wrapped-agent setup are the parts to check before adopting.
Who is it for?
Adopt Headroom if your agents spend most of their context budget on JSON tool results, log dumps or retrieved chunks, and you are willing to run a local proxy and keep the CCR cache on disk. Do not adopt it if your workload is ordinary prose conversation, if you cannot accept a lossy step between the tool and the model, or if you need a stable API rather than a 0.37.x series that shipped three releases in the last week of August 2026.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository received new commits within the last day.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

DEEP OPEN-SOURCE ANALYSIS

The context budget problem Headroom is aimed at

Coding agents are expensive for a reason that has little to do with the quality of the model: they read a lot. A single code search returns a hundred results, an incident debugging session pulls in log lines, a RAG pipeline hands over chunks that are mostly boilerplate. All of that lands in the prompt, and the prompt is what you pay for on every turn.

The README frames the target plainly. Headroom claims 60 to 95 percent fewer tokens for JSON data and 15 to 20 percent fewer tokens for coding agents, with the same answers. The distinction matters. A JSON tool result is highly structured and highly redundant, so a content-aware compressor can drop most of it. A coding session is mostly source code and instructions the model actually needs, so the ceiling is far lower.

The audience is developers running agents rather than people building chat products. The README lists Claude Code, Cursor, Codex, LangChain, Agno and Strands among the callers it expects to sit in front of, and the repository ships example files for several of them, including a LangChain demo, an MCP demo and a set of Strands scripts. If you are not piping tool output through a model, there is little here for you.

ContentRouter, CacheAligner and CCR: what actually happens to a payload

The architecture diagram in the README describes a three-stage pipeline. CacheAligner runs first. It detects volatile content that would break a provider's KV cache prefix and warns about it; the README is explicit that it never rewrites prompts. That is a deliberate restraint. A tool that silently edits prompts to improve cache hit rates would be very hard to trust, and Headroom does not do it.

ContentRouter comes next. It detects the content type and picks a compressor. There are three named ones: SmartCrusher for JSON, CodeCompressor, which the README describes as AST-based, and Kompress-v2-base, a text model hosted on Hugging Face under chopratejas/kompress-v2-base. A JSON array of search results and a Python file go down different paths.

The last stage is CCR, which the README expands as reversible compression. Originals are cached locally and the model can call headroom_retrieve to get them back. This is the design decision that makes the whole thing defensible: compression is lossy, and the escape hatch is a tool call rather than a re-fetch. The trade-off is that the cache is now part of your runtime. If the process that holds it goes away, retrieval goes with it.

The README also mentions output token reduction, which trims what the model writes back by dropping ceremony and restated code and skipping deep thinking on routine steps. That is a different lever from input compression and a more opinionated one, since it changes the model's behaviour rather than its input.

Installing Headroom and wrapping Claude Code or Cursor

The README gives three install paths and they are not equivalent. The CLI ships only through PyPI. The npm package headroom-ai is the TypeScript SDK, a library you import, and it provides no headroom command. Installing it with npm and then looking for the binary is a common way to lose ten minutes.

The recommended route is uv, which installs the CLI as a global tool in a self-contained virtual environment. Python 3.10 or later is required.

bash
uv tool install --python 3.13 "headroom-ai[all]"

After that, the headroom command is on your PATH. The README suggests running a wrapped agent session each time rather than configuring an agent by hand, because wrapping performs the setup for you. For Claude Code:

bash
headroom wrap claude

Wrapping does more than set an environment variable. According to the README, it starts a local proxy, installs Serena for semantic code navigation, and launches the agent session pointed at that proxy. Serena is registered at user scope, which for Claude Code means ~/.claude.json, so it stays available in your other projects until you run headroom unwrap. If you do not want it, wrap with --code-memory none.

If you would rather not wrap anything, the proxy is a drop-in OpenAI-compatible endpoint on port 8787:

bash
headroom proxy --port 8787

The docker-compose.yml file publishes that same port as 8787, alongside Qdrant on 6333 and 6334 and Neo4j on 7474 and 7687, all bound to 127.0.0.1. The compose file notes that none of the three services authenticates inbound callers by default and that the proxy's /v1/* data plane is open unless HEADROOM_PROXY_TOKEN is set. If you want to reach the proxy from another machine, the file says to set the token and publish the port together, never one without the other.

Where Headroom is the wrong tool

The headline numbers apply to JSON. The README's own figures put coding agents at 15 to 20 percent, and that is the honest range for a workload where most of the tokens are code the model needs to read. If your prompts are ordinary conversation, there is nothing for a content-aware compressor to find, and you have added a proxy hop and a cache directory for a rounding error.

Reversibility is a real limitation, not a footnote. The model has to decide to call headroom_retrieve. A compressor that drops a field the model did not realize it needed produces a confidently wrong answer, and that failure is quieter than a token overage. The README states that originals are cached for retrieval on demand; it does not describe what happens when the cache is unavailable or when the model simply does not ask.

There is also a maintenance cost that the release history makes visible. Three releases (v0.36.4, v0.36.5 and v0.37.0) landed between 2026-08-22 and 2026-08-27, and pyproject.toml still classifies the project as Development Status 4 - Beta. The last push to main was on 2026-08-27. Fast iteration on a component that sits between your agent and your model provider means you should expect to pin a version and read changelogs, not float on latest.

How Headroom differs from a generic LLM gateway

The obvious alternative is a proxy or gateway that sits in the same place in the request path and handles routing, retries, logging and spend tracking. Those tools generally treat the request body as opaque. They will count the tokens and forward the payload unchanged.

Headroom's difference is that it parses the payload and rewrites it based on what it finds. SmartCrusher sees a JSON structure and removes redundancy within it; CodeCompressor works on the AST rather than the text. That is why the JSON figures are an order of magnitude better than the coding-agent figures: the gains track how much structure the compressor can recognize.

The cost of that approach is that Headroom has to understand your content, and understanding is where it can be wrong. A gateway that forwards bytes cannot corrupt a tool result. Headroom can, which is why the CCR retrieval path exists and why the CacheAligner stage is careful to warn rather than rewrite. If you want observability and routing, a gateway is the simpler answer. If the bill is dominated by structured payloads, the parsing is the point.

The repository also contains a Rust workspace with crates for headroom-core, headroom-proxy, headroom-simulators, headroom-py and headroom-parity. The Cargo.toml comments describe a parity crate and a set of invariants around byte-faithful passthrough on unmutated bytes. That suggests the project is porting compression logic to Rust and testing the two implementations against each other, which is a stronger guarantee than a single implementation would give.

Licence, packaging and what a Headroom upgrade costs

Headroom is Apache-2.0 in both pyproject.toml and the Cargo workspace, and the repository carries a NOTICE file, which is the usual companion to that licence. Apache-2.0 is permissive and includes a patent grant, so the licence itself is unlikely to be the blocker for commercial use. The NOTICE file matters if you redistribute the software: read it and carry it forward. None of this is legal advice, and if you are embedding Headroom in a product you ship, have someone check the NOTICE and the third-party dependencies.

Upgrade cost is dominated by the extras system. The base install is light, but [all] pulls in the proxy, MCP, ML, code, memory, vector, relevance and image paths, and the README notes that [vector] needs a C++ toolchain and is not part of [all]. The Dockerfile builds native extensions with maturin and installs a Rust toolchain pinned to 1.95.0, so a from-source install is not a quick pip. If you deploy the container, you inherit that build chain.

The bigger upgrade question is the wrapped-agent configuration. headroom wrap writes to files outside your project, including ~/.claude.json for Serena. The README gives headroom unwrap as the undo, but it does not document what happens if a wrap is interrupted partway through, or how unwrap behaves when the agent has been reconfigured by hand in the meantime. Test wrap and unwrap on one machine before you run it across a team.

Editorial conclusion

Adopt Headroom if your agents spend most of their context budget on JSON tool results, log dumps or retrieved chunks, and you are willing to run a local proxy and keep the CCR cache on disk. Do not adopt it if your workload is ordinary prose conversation, if you cannot accept a lossy step between the tool and the model, or if you need a stable API rather than a 0.37.x series that shipped three releases in the last week of August 2026. Verify first that your agent is on the supported wrap list, that you can run the whole thing on Python 3.13 with the [all] extra, and that headroom unwrap restores your original agent configuration before you roll it out to a team.

Frequently asked questions

How do I install Headroom?

The README gives two PyPI routes: uv tool install --python 3.13 "headroom-ai[all]" installs the CLI as a global tool, and pip install "headroom-ai[all]" installs it into the current environment. Python 3.10 or later is required. The npm package headroom-ai is the TypeScript SDK and provides no headroom CLI.

How do I use Headroom with Claude Code?

The README recommends launching a wrapped session with headroom wrap claude, which starts a local proxy, installs Serena for semantic code navigation and launches the agent pointed at that proxy. You can undo it later with headroom unwrap. To skip Serena, wrap with --code-memory none.

How do I use Headroom with Cursor?

Cursor is on the wrap list, so the documented path is headroom wrap cursor, which starts the local proxy and launches the agent through it. The README does not give a Cursor-specific configuration beyond that command.

How do I use Headroom with Codex?

headroom wrap codex is the documented command. The README also notes that if Codex cannot inherit a shell PATH reliably, you should install Headroom as a persistent uv tool, run command -v headroom, and use that absolute path in the MCP config as command = "/absolute/path/from/command-v/headroom" with args = ["mcp", "serve"].

What is Headroom in AI?

In this project, Headroom is a context compression layer that sits between your agent and the LLM. It compresses tool outputs, logs, RAG chunks, files and conversation history before they reach the model, and exposes the same functionality as a Python and TypeScript library, a local proxy, and an MCP server.

Official sources

  1. Official documentation
  2. Official README
  3. Project repository
  4. Release notes
For maintainers

Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/headroomlabs-ai-headroom.svg)](https://hysenlabs.com/projects/headroomlabs-ai-headroom)
Community notes

Community notes