Open-source project
teamchong/pxpipe avatar
teamchong/pxpipe

pxpipe: cutting Claude Code input tokens by rendering context as PNGs

cut Fable 5 token usage by rendering text context as images. The reader is the same vision channel that Anthropic's computer use already relies on for screenshots.

7,447 stars652 forksTypeScriptMIT

At a glance

What is it?
pxpipe is a local proxy that rewrites bulky request context into dense images before it reaches the model. It cuts tokens on token-dense content, but it is lossy, and the README says so first.
Who is it for?
Adopt pxpipe if your Claude Code sessions are dominated by bulky, token-dense context (system prompt, tool docs, long file dumps) and you can keep byte-exact identifiers in text turns. Do not adopt it if your workload is sparse prose, or if you need verbatim recall of hashes, IDs and secrets from imaged content: the README reports 13/15 on Fable 5 and 0/15 on Sol for exact 12-char hex strings, with misses described as silent confabulations.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 4 days ago.
What is it written in?
Mainly TypeScript, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 27, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What pxpipe actually changes about a Claude Code request

Claude Code re-sends a lot of the same material on every turn: the system prompt, tool documentation, and older conversation history. That material is billed as input tokens each time. pxpipe sits between the client and the provider as a local proxy and rewrites the bulky parts of each request into PNG pages before the request leaves the machine.

The premise is that an image's token cost is set by its pixel dimensions, not by how much text is packed into it. The README states that dense content (code, JSON, tool output) packs roughly 3.1 characters per image-token against roughly 1 character per text-token on real Claude Code traffic. The reader for those images is the same vision channel Anthropic's computer use already uses for screenshots, so no new model capability is being assumed.

Who this is for: people running Claude Code on long sessions where the system prompt and tool docs dominate the bill. It is not for anyone who needs every identifier in the request to survive byte-for-byte, and the README says as much in a section titled "The honest part".

The proxy path, the imaging rule, and what stays text

The data flow is a request rewrite, not a response rewrite. Responses stream normally; pxpipe only compresses the request. Recent turns stay as text. The system prompt, tool docs and older bulk history are the parts that get imaged.

Not everything is imaged. A profitability gate, described as calibrated on N=391 production rows, decides per request whether the image math wins. The README notes the workload dependency directly: wins on token-dense content at about 1 char per token, loses money on sparse prose at about 3.5 chars per token. That gate is the difference between a token saver and a token adder, and it is the part I would want to see exercised on my own traffic.

Model scope is a default allowlist: claude-fable-5, gemini-3.6-flash, gemini-3.7-flash. Opus 5, Sol, GPT 5.5 and Grok are opt-in through the dashboard chips or the PXPIPE_MODELS environment variable. Anything not on the allowlist passes through byte-identical, and PXPIPE_MODELS=off disables imaging entirely. On the GPT path, tool definitions stay native JSON and no Anthropic cache_control markers are used.

Installing pxpipe and pointing Claude Code at it

The README's quick start is two commands. The first starts the proxy on 127.0.0.1:47821; the second points Claude Code at it through ANTHROPIC_BASE_URL. Nothing is installed globally if you use npx.

bash
npx pxpipe-proxy                                  # proxy on 127.0.0.1:47821
ANTHROPIC_BASE_URL=http://127.0.0.1:47821 claude  # point Claude Code at it

After that, open the dashboard at http://127.0.0.1:47821/. The README says it shows tokens saved, every text-to-image conversion side by side, a kill switch, and live model chips. That side-by-side view is the fastest way to see whether your own prompts are being imaged at all, or whether the profitability gate is leaving them alone.

If you need /remote-control, claude.ai connectors or first-party gates to keep working, the README offers a wrapper instead of the environment variable:

bash
pxpipe warp -- claude          # also: cursor-agent, codex, or a shell alias

The README states that api.anthropic.com/v1/messages is routed by default, and that agents reaching their provider over some other base URL need a rule. A rule that names a port matches only that port:

bash
pxpipe warp --route '127.0.0.1:9090/v1/*=http://127.0.0.1:47821' -- codex

There is also a Docker path. The compose.yml publishes the proxy on 127.0.0.1 with PXPIPE_PORT defaulting to 47821, runs the container read-only with cap_drop: ALL and no-new-privileges, and persists /data as a named volume. The Dockerfile sets PXPIPE_CONFIG=/data/config.json and PXPIPE_LOG=/data/events.jsonl inside that volume.

If you only want the rendering and not the proxy, the README documents an offline export:

bash
npx pxpipe-proxy export src/
cat prompt.txt | npx pxpipe-proxy export --stdin
npx pxpipe-proxy export --git

Each export run writes a fresh pxpipe-export-XXXXXX/ folder containing page-*.png, factsheet.txt, manifest.json and prompt.txt. The README describes the intended use as uploading the PNG pages and pasting the prompt into image-upload clients such as Cursor.

The lossy part, and why the factsheet is not a full answer

This is the limitation that decides the tool. The README states plainly that pxpipe is lossy. On exact 12-character hex strings inside dense imaged content it reports 13/15 on Fable 5 and 0/15 on Sol, and it describes the misses as silent confabulations rather than errors. A silent confabulation is worse than a parse failure: nothing in the output tells you the value is wrong.

The mitigation is a factsheet that selectively preserves up to 96 recognized precision-critical tokens. Not every identifier. The README notes that a dedicated verbatim-risk guard is not built yet, which means the boundary between preserved and unpreserved identifiers is a recognition heuristic you cannot fully audit from the outside.

The documented escape hatch is subagents on non-allowlisted models, which pass through as text. The README suggests routing byte-exact work there with CLAUDE_CODE_SUBAGENT_MODEL=claude-sonnet-4-6, or model: sonnet in agent frontmatter. That works, but it also means the token savings and the correctness guarantee live in different parts of the session, and someone has to remember which is which.

Model choice changes the risk profile too. The README reports claude-opus-5 at verbatim 2/15 against Fable 5's 13/15, while still scoring 100/100 on arithmetic and 0/16 on never-stated items. Suggested effort for Opus 5 is medium. So the same pipeline is materially less reliable at recall depending on which model reads the image.

Where the savings come from, and where they do not

The headline figure in the README is a roughly 59 to 70 percent lower end-to-end bill at then-current Fable list prices, with the caveat that prices move and workloads differ. The durable number it points to instead is the token cut itself, measured per request against a free count_tokens counterfactual in ~/.pxpipe/events.jsonl. That is the number to trust, because it is measured on your traffic rather than on a demo.

Savings also depend on the client. The README states that savings track uncached bulk the client still re-sends as text, that Claude Code re-sends system plus tools plus history on /anthropic/messages, and that this typically lands around 60 to 70 percent. The measured splits are in docs/CACHING_AND_SAVINGS.md. If your client caches more aggressively, or sends less bulk, the arithmetic changes.

The evaluation evidence is small and the README does not hide it. A SWE-bench Lite pilot came out 10/10 on both arms at 65 percent lower request size. SWE-bench Pro came out 14/19 with pxpipe on versus 15/19 with it off at 60 percent lower request size, verdicts agreeing on 18 of 19, and the single split re-resolved 3/3 on replication, which the README attributes to run-to-run variance rather than compression. Small n, receipts in eval/. One caveat is visible in the demo clip: the pxpipe arm needed a nudge to match the requested one-line output format.

How pxpipe differs from text-compression tools

The obvious alternative is a tool that trims or summarizes context as text. The difference is the mechanism, and it changes the failure mode. Text trimming is deterministic: you can read the trimmed prompt and know exactly what was removed. pxpipe keeps the content but changes the channel, so the model reads a rendering. The README's own framing is that the reader is the same vision channel used for screenshots, which means accuracy now depends on the model's OCR-like recall rather than on string handling.

That trade buys capacity that text trimming cannot. The README's chart claims roughly 19.0M characters for Fable 5 (4.8x) and roughly 21.3M characters for Gemini 3.6 Flash (5.3x) in the same 1M-token windows when read through pxpipe images, against text lines that top out near 4M characters. The chart is generated by npx tsx scripts/gen-context-chart.ts from a live render, not hand-typed, per the README.

A second alternative is simply choosing a model with a larger text window. That avoids the lossy step entirely, but the README's chart places Grok 4.5 as a text-window point at 500K, so the comparison is really about whether you can get more effective context from a fixed window by changing representation rather than by buying a bigger one. Neither approach is free: one costs money per token, the other costs verbatim fidelity.

Maintenance, licence and what upgrading costs you

The repository is not archived. The last push was on 2026-08-21, the same day as the v0.13.2 release, following v0.13.1 on 2026-08-11 and v0.13.0 on 2026-08-09. Three releases in under two weeks is a fast-moving surface, and the package version in package.json is 0.13.2, so the project is still pre-1.0. Expect configuration keys and model allowlists to move; the README already documents PXPIPE_MODELS, PXPIPE_GPT_HISTORY_MAX_IMAGES, PXPIPE_DISABLE, PXPIPE_UPSTREAM and several provider-specific variables, and the Dockerfile pins PXPIPE_CONFIG and PXPIPE_LOG into /data.

The licence is MIT, and the Dockerfile labels the image with org.opencontainers.image.licenses="MIT". The package ships assets/*LICENSE.txt into ./licenses/ in the container. That is a permissive licence, but it says nothing about the terms of the model provider you route through, which is a separate question and not one the README addresses. Nothing here is legal advice; read the LICENSE file and your provider terms.

Operationally, the upgrade cost is mostly the re-validation loop. The README's own measurement path is ~/.pxpipe/events.jsonl, and the eval/ directory holds the receipts. After a version bump, the thing worth re-checking is whether the profitability gate and the per-model render profiles still behave the way your last measurement showed, because those are the parts that decide whether imaging helps or hurts.

Editorial conclusion

Adopt pxpipe if your Claude Code sessions are dominated by bulky, token-dense context (system prompt, tool docs, long file dumps) and you can keep byte-exact identifiers in text turns. Do not adopt it if your workload is sparse prose, or if you need verbatim recall of hashes, IDs and secrets from imaged content: the README reports 13/15 on Fable 5 and 0/15 on Sol for exact 12-char hex strings, with misses described as silent confabulations. Before adopting, run the proxy with PXPIPE_MODELS=off to confirm your client still works unchanged, then read ~/.pxpipe/events.jsonl after a real session to see whether the per-request token cut actually applies to your traffic, and check docs/CACHING_AND_SAVINGS.md for the caching split.

Frequently asked questions

How do I install pxpipe and start the proxy?

The README's quick start runs npx pxpipe-proxy to start the proxy on 127.0.0.1:47821, then sets ANTHROPIC_BASE_URL=http://127.0.0.1:47821 before launching claude. The dashboard is at http://127.0.0.1:47821/. There is also a Docker path via compose.yml, which publishes the proxy on 127.0.0.1 with PXPIPE_PORT defaulting to 47821.

Is pxpipe lossy, and which values should stay as text?

The README states that pxpipe is lossy: exact 12-character hex strings in dense imaged content scored 13/15 on Fable 5 and 0/15 on Sol, and the misses are described as silent confabulations. Byte-exact values such as IDs, hashes and secrets must stay text, and recent turns do. A factsheet preserves up to 96 recognized precision-critical tokens, not every identifier.

Does pxpipe work with models other than the default ones?

The default PXPIPE_MODELS value is claude-fable-5,gemini-3.6-flash,gemini-3.7-flash. Opus 5, Sol, GPT 5.5 and Grok are opt-in only, enabled through the dashboard chips or the PXPIPE_MODELS environment variable, and everything else passes through byte-identical. PXPIPE_MODELS=off disables imaging.

Official sources

  1. Official README
  2. Project repository
  3. Release notes
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/teamchong-pxpipe.svg)](https://hysenlabs.com/projects/teamchong-pxpipe)