Model or dataset
jia-gao/leanctx avatar
jia-gao/leanctx

leanctx: drop-in prompt compression that routes code and tool output around the compressor

Drop-in prompt compression for production LLM apps. Cut your token bill 40-60% without changing your code. Python SDK, LLMLingua-2, MIT.

326 stars6 forksPythonMIT

At a glance

What is it?
leanctx wraps the OpenAI, Anthropic and Gemini Python clients and compresses only the parts of a prompt its classifier decides can survive distortion. The repository documents a 1.8 pp accuracy cost against an 18.7 percent token reduction layered on an existing compressor, and it is honest that compression never adds accuracy.
Who is it for?
Adopt leanctx if you run a Python LLM app with large retrieved documents, growing tool-call histories or long chat logs, and you are willing to accept a measured accuracy cost in exchange for a token reduction you can regenerate from the committed per-item records. Do not adopt it if your prompts are mostly code, JSON or stack traces: those route to verbatim by design, so you would pay the sidecar latency and the extra dependency for close to zero savings.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 39 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The bill problem leanctx targets, and the apps it fits

The project exists because input tokens are a recurring line item. The README lists four situations: RAG apps carrying large retrieved documents, long-running conversational agents on LangChain, LangGraph or CrewAI, document-processing pipelines, and coding agents with growing tool-call histories. In all four the prompt grows over the life of a request or a session, and the growth is uneven. A retrieved passage is prose. A tool result is often a stack trace or a JSON blob. Compressing both the same way is the failure mode the project is built around.

The README is direct about where the existing options stop. Provider prompt caching, on Anthropic, OpenAI or Gemini, helps with stable prefixes such as system prompts, tool definitions and retrieved-document pools, but the README states it does not help with dynamic per-query content: chat history, freshly retrieved docs, tool outputs. The project's position is that you compose the two rather than pick one. Naive truncation is dismissed because it drops the middle of a document, and the README says the LongBench v2 numbers show that concretely. Hosted compression APIs are named as requiring you to send context to their servers, and leanctx's counter is that the model runs locally and the only outbound calls go to your existing provider.

That last point decides the audience. If your prompts contain regulated data and you cannot send them to a third-party compression service, the local-model design is the reason to look here. If you already have a hosted compressor you trust, the case is weaker.

Loss-tolerance routing: the mechanism that separates leanctx from a compressor wrapper

The core idea is a classifier that assigns every segment of a prompt to one of three classes before any compression happens. The README calls this loss-tolerance routing, and the flow is: an agent request enters, segments are classified, each class gets a different treatment, the segments are recomposed, invariants are checked, and the result is either sent compressed or, on failure or timeout, sent as the original.

The three classes are the substance. Zero tolerance covers code, stack traces, tool_use_id, tool name and input, and JSON. These are copied verbatim, byte for byte, with 0 percent altered. High tolerance covers documentation, retrieved passages, logs and prior turns, and those go through LLMLingua-2 on-device with roughly 50 percent removed. Conditional covers low-confidence prose and oversized context, and uses a self-LLM path that is opt-in, removing 41 to 49 percent.

The release history shows how seriously the zero-tolerance class is taken. v0.3.1 is titled "preserve tool_result code/traceback verbatim", which means an earlier version did not, and the fix was important enough to cut a release for. If you have ever watched a coding agent lose the line number in a traceback after a middleware pass, that changelog entry is the one to read.

The recompose-and-check step is the safety valve, and the fallback is the design decision worth noting. On a failed invariant check or a timeout, the original prompt goes to the provider. You lose the savings on that request, not the request. That is the right default for a middleware layer, though it also means your savings rate depends on how often the check trips, and the README does not document a metric for that.

Installing leanctx and making a first compressed call

The package is on PyPI. The README gives an extras-based install, and the extras matter because the base install has no dependencies at all: pyproject.toml lists an empty dependencies array, so provider SDKs and the compressor are opt-in.

bash
pip install 'leanctx[openai,lingua]'    # or [anthropic], [gemini]

The lingua extra pulls llmlingua, and the README warns that the first Lingua call loads roughly 1.2 GB of model weights into ~/.cache/huggingface/. Later calls reuse the cache. If you skip the lingua extra, leanctx falls back to passthrough, so a missing extra looks like a working install that saves nothing.

The usage pattern is the drop-in claim in practice. You import OpenAI from leanctx instead of from openai and pass a leanctx_config dictionary with the mode, a token threshold that decides when compression starts, and a routing map.

python
from leanctx import OpenAI

client = OpenAI(
    leanctx_config={
        "mode": "on",
        "trigger": {"threshold_tokens": 2000},
        "routing": {"prose": "lingua"},
    },
)

response = client.chat.completions.create(
    model="gpt-4o-mini",
    max_tokens=512,
    messages=[{"role": "user", "content": LONG_DOCUMENT}],
)

The response carries two extra usage fields, leanctx_tokens_saved and leanctx_ratio, which the README prints as 1841 and 0.49 in its example. Those are the numbers to log first, because they tell you whether the threshold of 2000 tokens is being crossed at all in your traffic.

There is also a verification path that needs no API key, which is unusual and useful for CI. The bench CLI ships seven registered scenarios, and the agent-structural one enforces five invariants.

bash
leanctx bench list                                   # 7 registered scenarios
leanctx bench run agent-structural --workload agent  # 5 invariants enforced, exit 0 = pass

Exit code 0 means pass, so this drops into a pipeline without parsing output. The bench extra installs respx, which mocks provider HTTP, and the pyproject comment says that is why no API key is needed. For a container, the repository ships a Dockerfile with a LINGUA build arg: the default slim image is about 180 MB with the three provider SDKs and tiktoken, while LINGUA=true adds llmlingua, torch and transformers for a much larger image, around 3 GB by the Dockerfile's own comment.

The accuracy cost, and the sub-bucket where it is worst

The headline benchmark is honest in a way that is rare. All 503 LongBench v2 questions, Claude Haiku 4.5 evaluation, temperature 0.1, run as a semantic pass on top of ClawRouter's seven structural compression layers. Raw prompts average 27,865 tokens, ClawRouter alone brings that to 26,397, and adding leanctx as Layer 8 brings it to 21,470. That is 23.0 percent against raw and 18.7 percent against the already-compressed baseline.

Against that, accuracy moves from 45.3 percent to 43.5 percent overall, a loss of 1.8 pp. The README states plainly that compression costs accuracy, it does not add it, and that the claim is only that the cost is small and bounded against the 2 pp go/no-go gate set for the integration. The split explains the number: 275 of the questions routed entirely to verbatim, where leanctx changed nothing and the accuracy delta is exactly 0.0 pp, and 228 routed to lingua, where the loss is 3.9 pp. Because 54.2 percent of Layer-8 input tokens route to verbatim, the compression actually applied to eligible content is 40.8 percent.

The part to read before adopting is the sub-bucket table. short/lingua is minus 17.6 pp on N=68, and Single-Document QA is minus 11.5 pp on N=78. Long Structured Data Understanding goes the other way, plus 9.4 pp on N=32, which is more likely noise than a benefit. A 1.8 pp average is not the number that matters if your traffic looks like short prose queries, because that bucket is where the damage concentrates. The README does not offer a per-bucket gate, only the aggregate one.

Two other details belong here. Sidecar latency is 47 ms p50 on GPU, which is the cost you pay on every compressed request. And at Sonnet input pricing the savings are about $78 per 1,000 requests. The README also records that the benchmark was executed and audited by outside contributors, and that an audit of an earlier draft found a headline "+7.4 % on long context" was eval noise rather than a real compression effect. Publishing that correction is a point in the project's favor.

Where leanctx is the wrong tool

The routing design sets a hard boundary. If your prompts are mostly code, stack traces, JSON or tool-call payloads, nearly everything classifies as zero tolerance, gets copied verbatim, and contributes nothing to the savings by construction. You would still pay the sidecar latency, the 1.2 GB model download if the lingua extra is installed, and the extra dependency surface, for a ratio near zero. The README's own token accounting supports this reading: the verbatim half contributes exactly 0 to both the savings and the accuracy delta.

Short prompts are the second boundary. Compression only starts once the request crosses threshold_tokens, which the README sets at 2000 in its example. A workload of brief questions with brief answers never triggers it, and the short/lingua bucket is also where the documented accuracy loss is largest, so short prose is the worst combination of low savings and high distortion.

The project also labels itself Alpha in pyproject.toml, and the version is 0.3.1. That is not a maturity claim either way, but it tells you the integration surface is still moving: v0.3.0 added OpenTelemetry observability and the bench CLI, and v0.3.1 changed what gets preserved verbatim. If you pin an older minor version, you may be running the behavior that v0.3.1 was released to fix.

Finally, the README does not document rollback. The fallback on a failed invariant check or timeout is to send the original prompt, which covers the per-request case, but there is no described way to disable compression globally at runtime without changing client construction, and no documented kill switch beyond the mode key. Plan for that before an incident, not during one.

How leanctx differs from ClawRouter and from hosted compression APIs

ClawRouter is the natural comparison because the benchmark layers leanctx on top of it, and the two do different work. ClawRouter's seven layers are described as structural compression. leanctx adds a semantic pass as Layer 8, and the measured delta of 18.7 percent is explicitly what leanctx adds to a system that is already compressing. If you already run ClawRouter, the question is whether the extra 18.7 percent justifies 1.8 pp of accuracy, and the README frames that against a 2 pp gate rather than claiming a free win. If you run neither, ClawRouter gives you structural reduction with no accuracy discussion in the README, while leanctx gives you a routing classifier and a local model.

Hosted compression APIs such as Compresr and Token Company are the other axis. The README states they require sending your context to their servers and that their models are closed source. leanctx's difference is that the compressor runs on your infrastructure, the code is MIT licensed, and no outbound calls happen except to your existing provider. That is a data-residency argument, not an accuracy argument, and the README does not compare accuracy between the two approaches, so do not read leanctx as more accurate. It is a different deployment model with a published accuracy cost of its own.

Provider prompt caching is not a competitor at all in the README's framing. It wins on stable prefixes, leanctx works on dynamic per-query content, and the stated position is to compose them. That is the one comparison where the project claims complementarity rather than advantage.

Licence, dependencies and the cost of keeping up

leanctx is MIT licensed, and the repository ships a LICENSE file at the top level. The practical consequence is that you can vendor it, fork it or ship it inside a closed product, which is not true of the hosted alternatives the README names. The LLMLingua-2 model weights and the torch and transformers stack that the lingua extra pulls in carry their own licences, and those are separate from leanctx's MIT grant. This is not legal advice; if you redistribute a container built with the LINGUA build arg, check the licences of everything that image contains, because the slim and lingua images are very different artifacts.

The dependency story is deliberately thin. The base package declares no dependencies. Provider SDKs, tiktoken, llmlingua, respx, datasets and the OpenTelemetry extras are all optional and version-pinned in pyproject.toml: anthropic at 0.40.0 or newer, openai at 1.50.0 or newer, google-genai at 0.3.0 or newer, llmlingua at 0.2.0 or newer, tiktoken at 0.7 or newer. Python 3.10 through 3.12 are the classified versions.

Upgrade cost concentrates in two places. First, the lingua path: llmlingua, torch and transformers move on their own schedules, and a torch upgrade is the kind of change that can force a rebuild of a large image. Second, the routing rules themselves, since v0.3.1 changed what is preserved verbatim, so a minor bump can change your token savings and your accuracy in the same step. The OpenTelemetry extra is described as API-only: leanctx never owns the SDK or exporter, and the application configures TracerProvider, MeterProvider and exporters. That keeps the dependency out of the library but puts the operational work on you. The last push to the repository was on 2026-08-22, and the most recent release in the repository is v0.3.1 from 2026-04-26.

Editorial conclusion

Adopt leanctx if you run a Python LLM app with large retrieved documents, growing tool-call histories or long chat logs, and you are willing to accept a measured accuracy cost in exchange for a token reduction you can regenerate from the committed per-item records. Do not adopt it if your prompts are mostly code, JSON or stack traces: those route to verbatim by design, so you would pay the sidecar latency and the extra dependency for close to zero savings. Before wiring it into production, run pip install 'leanctx[openai,lingua]' and leanctx bench run agent-structural to confirm the five invariants hold on your platform, then check the short/lingua bucket in benchmarks/clawrouter/full_long_bench_evaluation_result.md, which is the sub-bucket where the documented accuracy loss is largest.

Frequently asked questions

How do I use leanctx in an existing Python app?

Install with an extras target such as pip install 'leanctx[openai,lingua]', then import OpenAI from leanctx instead of from openai and pass a leanctx_config dictionary with mode, trigger and routing keys. The README's example sets mode to on, threshold_tokens to 2000 and routes prose to lingua. Without the lingua extra, leanctx falls back to passthrough.

What is leanctx?

It is an MIT-licensed Python SDK for prompt compression in production LLM applications, built around loss-tolerance routing. Segments classified as zero tolerance, such as code, stack traces and JSON, are copied verbatim, while prose, retrieved passages and prior turns go through LLMLingua-2 on-device. The README states prompts and user data never leave your infrastructure by default.

What is a real alternative to leanctx for cutting token costs?

The README names provider prompt caching on Anthropic, OpenAI and Gemini as complementary rather than competing, since it helps with stable prefixes and not with dynamic per-query content. For a direct substitute it names hosted compression APIs such as Compresr and Token Company, which require sending context to their servers and use closed-source models. The README does not compare accuracy between leanctx and those services.

Official sources

  1. Issues
  2. jia-gao/leanctx on GitHub
  3. License: MIT
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/jia-gao-leanctx.svg)](https://hysenlabs.com/projects/jia-gao-leanctx)