Model or dataset
open-compress/claw-compactor avatar
open-compress/claw-compactor

Claw Compactor: a 14-stage token pipeline that rewrites context before the model sees it

14-stage Fusion Pipeline for LLM token compression — reversible compression, AST-aware code analysis, intelligent content routing. Zero LLM inference cost. MIT licensed.

2,007 stars183 forksPythonMIT

At a glance

What is it?
Claw Compactor is a Python token compression engine from open-compress that chains fourteen content-aware stages over prompts, code and logs. It runs locally with no model inference, and its main selling point, reversible rewinding, is also where the sharpest limits sit.
Who is it for?
Adopt Claw Compactor if you send large, structured payloads (code, JSON, logs, diffs) to an LLM and want the reduction to happen before the request leaves your process. Do not adopt it if your context is short conversational text, or if you need a guarantee that every compressed span can be recovered, because recovery depends on the RewindStore still holding the original.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Activity is slowing. The repository last received commits 6 months ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The problem: context you pay for but do not need

Most LLM calls carry more text than the task requires. A workspace scan, a pasted log, a JSON response, or a diff arrives whole, and the model is billed for all of it. Claw Compactor targets that gap directly: it is a pre-processor that sits between your data and the API call, shrinking the payload before it is sent. The README frames the goal as "15-82% compression depending on content" with "Zero LLM inference cost", meaning the reduction is produced by deterministic Python stages rather than by a second model.

The intended audience is developers building agents or tooling that feed structured material into a context window: repository contents, logs, search results, diffs. It is less obviously aimed at chat applications, where the input is short and already dense. The project is MIT licensed, written in Python, requires Python 3.9 or newer, and is published on PyPI as claw-compactor. The repository's pyproject.toml declares no runtime dependency list in the excerpt available, which matches the README's claim of zero required dependencies.

How the Fusion Pipeline moves data through fourteen stages

The architecture is a linear chain. An input enters as a FusionContext, described in the README as a frozen dataclass, and each stage returns a new FusionResult rather than mutating anything in place. That immutability is the design's backbone: it makes stage ordering explicit and keeps a stage from accidentally seeing another stage's side effects.

Before a stage does any work, it calls should_apply(), which inspects content type, language and role. Stages that do not match are skipped at zero cost. Cortex runs early and classifies the payload as code, JSON, logs, diffs or search results, and detects the language for code. Downstream stages then make type-aware decisions. Ionizer handles JSON by statistical sampling. Neurosyntax compresses code through tree-sitter AST analysis. LogCrunch, SearchCrunch and DiffCrunch each fold their own content shape. TokenOpt, Abbrev and Nexus handle formatting, natural-language shortening and token classification near the end of the chain.

The README's own benchmark output shows this selectivity in practice: on a 47-file workspace, Ionizer applied to 8 of 47 files but produced a 71.2% reduction where it fired, while TokenOpt applied to all 47 and produced 4.1%. That spread is the honest picture of a routing pipeline. Most stages contribute little on most files, and one or two stages carry the result.

Installing claw-compactor and running a first benchmark

The package is on PyPI under the name claw-compactor, and the project exposes a console script of the same name. Install it into a Python 3.9 or newer environment:

bash
pip install claw-compactor

After installation, the README's demo shows a benchmark subcommand that scans a directory, reports per-stage application counts and reductions, and prints a token total with a cost estimate. Run it against a workspace you already have:

bash
claw-compactor benchmark ./my-workspace

The output the README presents is a table with one row per stage, showing how many files each stage applied to, the reduction it produced, and the milliseconds it took, followed by a summary of tokens before and after. Treat the dollar figures in that summary as illustrative: the README computes them from a GPT-4 rate it does not restate in the excerpt, so your own cost depends on your model's pricing. What you should actually check is the Applied column. A stage showing 0/47 on your content is telling you the router did not consider that stage relevant, and the compression you get will come from the rows that did fire.

Reversibility is real but stored, not free

Reversible compression is the feature most likely to decide adoption. The README describes a RewindStore, a hash-addressed LRU cache where Ionizer places originals so the LLM "can call" them back. The important word is stored. Reversibility here means the original bytes are kept in a local cache keyed by hash, not that the compressed form mathematically reconstructs the input.

That distinction has consequences the README does not resolve. An LRU cache evicts. If the compressed payload outlives the cache entry, or if the process that holds the store is not the process that later needs the original, the reversible path is gone and you are left with the lossy output. The README does not document eviction limits, persistence of the store across runs, or what happens on a cache miss. If your workflow depends on recovering exact originals, those are the questions to answer from the source before you commit, not after.

A second limit is scope. The comparison table marks Claw Compactor reversible with a plain "Yes", but the pipeline contains stages that clearly are not: abbreviation of natural language and token-level format optimization discard information by definition. The README does not state which stages participate in the rewind path. Assume reversibility applies to the stages that explicitly store originals, and treat the rest as lossy.

Where it beats perplexity-based compressors, and where it does not

The nearest alternative named in the README is LLMLingua-2, which drops tokens by perplexity score. The difference in approach is the whole argument. A perplexity filter has one signal, and it is tuned for natural language; applied to code it will happily delete identifiers, and applied to JSON it will delete keys. Claw Compactor routes by content type instead, so a tree-sitter stage handles Python and a sampling stage handles JSON. The README's comparison table reports a ROUGE-L of 0.653 at a 0.3 compression rate against 0.346 for LLMLingua-2, and notes that LLMLingua-2 pulls in torch and transformers while Claw Compactor lists zero required dependencies.

SelectiveContext is the other reference point, also self-information based, listed at 10-40% compression. gzip plus base64 reaches 60-80% but produces output the README describes as not LLM-readable, which is the right way to think about it: a compressed byte stream saves tokens but the model cannot reason over it, so the saving is only useful if the model never needs to read the content.

The honest counterweight is that content-aware routing is only as good as its classifier. If Cortex mislabels a file, the wrong stage runs, and the README does not describe a confidence threshold or a fallback path for ambiguous input. A mixed-format document, such as prose with embedded code fences, is the case most likely to route badly.

Maintenance status, licence and the upgrade question

The repository is not archived, and the last push was on 2026-04-01. The most recent release in the list is v7.1.0 from 2026-03-20, described as "Modular Architecture + Repo Polish". Before it came v7.0.2 and v7.0.1 on 2026-03-18, both packaging fixes for CI and PyPI. That sequence is worth reading carefully: two patch releases in the same day, both about getting the wheel and the PyPI upload right, immediately before a minor release that restructured the code. It suggests a project that has recently moved its internals and is still settling the packaging around that move.

For upgrade cost, the practical risk is the module layout, not the CLI. A v7.1.0 that advertises a modular architecture is likely to have moved import paths relative to earlier versions, and the CHANGELOG in the repository is the place to confirm that before pinning. The CLI surface shown in the README, claw-compactor benchmark, is the stable part.

The licence is MIT, which permits commercial use, modification and redistribution provided the copyright notice and permission notice are included. That is a permissive arrangement with few obligations, but it is not legal advice, and if you are embedding the library in a product you should have your own counsel read the LICENSE file rather than a summary.

Editorial conclusion

Adopt Claw Compactor if you send large, structured payloads (code, JSON, logs, diffs) to an LLM and want the reduction to happen before the request leaves your process. Do not adopt it if your context is short conversational text, or if you need a guarantee that every compressed span can be recovered, because recovery depends on the RewindStore still holding the original. Verify three things first: that pip install claw-compactor resolves to the version you expect, that the stages listed under claw-compactor benchmark actually fire on your files, and that whatever you do with the compressed output can tolerate the lossy stages, since only the reversible path is documented as recoverable.

Frequently asked questions

What is the purpose of Claw Compactor?

It is an LLM token compression engine that reduces the size of prompts, code, JSON, logs and diffs before they are sent to a model. It runs a 14-stage Fusion Pipeline locally, so the reduction costs no model inference.

How do I install Claw Compactor?

It is published on PyPI as claw-compactor and requires Python 3.9 or newer. Installing the package also provides a claw-compactor console script, which the README uses for the benchmark subcommand.

Does Claw Compactor need an LLM to compress text?

No. The README states zero LLM inference cost and lists zero required dependencies, in contrast to LLMLingua-2, which the comparison table shows depending on torch and transformers. All fourteen stages are deterministic Python.

Is Claw Compactor compression reversible?

The README describes Ionizer storing originals in a hash-addressed RewindStore that the LLM can call back. Because that store is an LRU cache, recovery depends on the original still being present, and the README does not document eviction limits or persistence.

Official sources

  1. License: MIT
  2. open-compress/claw-compactor on GitHub
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/open-compress-claw-compactor.svg)](https://hysenlabs.com/projects/open-compress-claw-compactor)