# Paritok: a non-destructive compression gateway for coding agents

> Paritok sits between your agent and the LLM API, filters tool schemas, compresses tool results and summarizes stale history, then forwards the request billed on the compressed tokens. The README claims 25% savings on turn one rising past 85% in long sessions, and the whole thing installs as a Python package.

**Paritok-official/paritok-4b-v1** — Non-destructive compression gateway for AI coding agents. Cuts token bills 25% on turn 1 to past 85% in long or saturated sessions, and fits ~3× more turns in the same context window. Powered by our open-source code-native 4B model. Drop-in for Claude Code, Cursor, Codex, OpenHands, and any BASE_URL agent.

- Repository: https://github.com/Paritok-official/paritok-4b-v1
- Website: https://www.paritok.com/
- Stars: 1,453 · Forks: 140
- Language: Python
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/paritok-official-paritok-4b-v1

## The bill Paritok is trying to cut

A coding agent re-sends the same material every turn. Dozens of tool schemas go out in full JSON, the message history grows, and file reads and tool outputs pile up inside it. The README states that once MCP servers are added, an agent can expose 70 or more tools on every request, and that on a typical Claude Code turn the tool block alone is around 29K tokens. That is the target. Paritok is a proxy that rewrites the outgoing request before it reaches Anthropic or OpenAI, so the compressed payload is what gets billed. The audience is narrow and specific: people running Claude Code, Cursor, Codex or OpenHands against a metered API, who care about input-token cost and context-window pressure more than they care about keeping the request byte-identical. If you use a flat-rate subscription or a local model, the savings have no line item to land on.

## Three levers, and only one of them pays off on turn one

The README describes three independent mechanisms that stack. The first is a semantic tool-schema filter. It keeps the handful of tools relevant to the user's intent in full schema and stubs the rest, using the local embedding model BAAI/bge-small-en-v1.5 through fastembed and onnxruntime on CPU. The README calls this the largest single-turn lever, dropping the tool block from roughly 29K to roughly 8K tokens on a typical Claude Code turn, and notes the selection is frozen per conversation so the tools array stays byte-stable and does not invalidate the provider's KV cache. A filtered tool is not gone: the model calls gateway_search_tools to get the full schema back. The second lever compresses each tool result, file read and stale history turn down to about 26% of its original size, tagged [REF:id], with the agent able to pull the exact original back through read_original. The README is candid that this is the smaller lever on a single turn, since most of a turn's cost is the fixed prefix, and that its value appears across a session. The third summarizes turns beyond the recent window once the context fills, so a long session stays inside the window instead of overflowing. The design point worth naming: the wire format is lossy, the store is not.

## Installing the Paritok gateway and pointing an agent at it

The package is published as paritok and the build backend is hatchling. The proxy dependencies are an optional extra, and the pyproject.toml comment states that the quantized bge-small ONNX weights are bundled in the wheel under paritok/data/bge-small-onnx/, so nothing downloads at runtime. The extra is declared as proxy = ["uvicorn>=0.30.0", "starlette>=0.38.0", "fastembed>=0.7", "numpy>=1.24"].

```toml
proxy = ["uvicorn>=0.30.0", "starlette>=0.38.0", "fastembed>=0.7", "numpy>=1.24"]
```

That pulls uvicorn, starlette, fastembed and numpy alongside the base dependencies tiktoken, click, pyyaml and httpx. If you only want the embedding-based tool selection without the proxy, the toolselect extra installs fastembed and numpy alone. Persistent shadow storage is a separate extra: the redis extra adds the redis client, and the pyproject comment says entries survive restarts when shadow_storage is set to redis.

```toml
redis = ["redis>=5.0"]
```

The README's architecture diagram shows the agent pointing at Paritok instead of Anthropic or OpenAI, with the response flowing back unchanged. The repository carries a deploy.sh at the top level and an examples/inference directory, but the README excerpt does not spell out the exact serve command or the default port, so read deploy.sh and the examples before assuming a flag. The one configuration key the README does name is tool_discovery.strategy, set to embedding for the semantic filter.

## The 4B model is the part with real operational weight

Content compression is not a heuristic. It runs a 4B model, Paritok-4B-v1, built on a Qwen3-4B backbone and trained on 45K teacher-distilled samples drawn from real agent trajectories. The README says the model learned to protect identifiers, paths and error strings while dropping noise, which is the behaviour that separates it from a generic summarizer. That model has to run somewhere. The pyproject description offers two paths: self-host the 4B model, or use the Paritok GPU server. The README's own session numbers come from the GPU path, which it labels explicitly. The training requirements file pins transformers below 4.50, torch is deliberately left out of requirements.txt with a note to install it separately, and the header warns that Python 3.12 and above has issues with some flash-attn builds. None of that affects a proxy-only install, but it does describe the cost of the self-hosted route. The embedding filter, by contrast, is CPU-only and ships its weights in the wheel. Two very different resource profiles inside one package.

## What the README does not settle

The savings figures are the weakest part of the documentation. The README states the 25% turn-one figure and the past-85% long-session figure, and then, in the section on compounding savings, begins a sentence about an A/B run over five consecutive turns on a read-only find-the-bug task with the phrase important: in this A/B the too. The excerpt ends there. Whatever caveat was coming, it is not in the documentation available here. Treat the headline numbers as vendor claims with an unfinished qualifier attached, not as an independent measurement. Several operational questions go unanswered. The README does not document rollback if the gateway misbehaves mid-session. It does not say how a read_original expansion is billed, whether the restored bytes are re-sent upstream at full price or served locally. It does not list which upstream providers have been exercised beyond the named agents. And it does not describe behaviour when the 4B model is unreachable, which matters because the compression step sits in the request path. A gateway that fails open passes uncompressed traffic; one that fails closed blocks the agent. The README does not say which this is.

## Where a plain proxy or client-side compaction fits better

The obvious alternative is a generic LLM proxy such as LiteLLM, which also sits between agent and provider and rewrites the base URL. The difference is the mechanism. A generic proxy routes, logs, retries and normalizes provider APIs, but it does not decide which tool schemas are relevant to the current intent, and it does not run a model to compress tool output. If your problem is key management or provider fallback, a routing proxy is the right tool and Paritok is not. The second alternative is what the README calls aggressive client-side compaction: when the context window fills, the client drops or summarizes history itself. That is destructive by construction. Once the client discards a file read, the exact bytes are gone and the agent has to re-read the file, which costs tokens again. Paritok's claim is that it keeps the original addressable through [REF:id] and read_original. Whether that round trip is cheaper than re-reading depends on the pricing question above. The third case is simpler: if your sessions are short, the fixed prefix dominates and the compounding lever never engages. A one-turn task gets the tool filter and little else.

## Licence, maintenance and the upgrade surface

The project is Apache-2.0, declared in pyproject.toml and in the LICENSE file at the repository root. That is a permissive licence with an explicit patent grant, and it is the same licence as the bge-small embedding model the README cites, which is also MIT. The bundled ONNX weights ship inside the wheel, so redistributing the package redistributes those weights; Apache-2.0 and MIT both permit that, but anyone vendoring the wheel into a commercial product should confirm the notice requirements themselves rather than take this as legal advice. On maintenance, the last push was on 2026-09-08, and the repository is not archived. The README's news section lists v1.3.8 on 2026-08-17 adding a visual /stats dashboard and a VS Code extension, v1.3.0 on 2026-07-31 as a stability release with edit-recovery and the read_original rename from expand_context, and v1.2.0 on 2026-07-19 shipping the embedding tool filter. The installed package version in pyproject.toml is 1.3.12, which is ahead of anything in the news list, so the changelog lags the release. The rename is the upgrade cost to watch: if you wrote against expand_context, v1.3.0 moved it. The README also carries a CITATION.cff and an arXiv identifier, so the project expects to be cited as research as well as used as a tool.

## Conclusion

Adopt Paritok if you run long coding-agent sessions against a token-billed API and want the tool-schema block cut before anything else. Do not adopt it if you cannot run a local embedding model on CPU or if your workflow depends on byte-stable prompts you never want rewritten. Before trusting it, verify what the README does not state: which upstream providers it has been exercised against, how a restored read_original segment is billed, and what the gateway does when the 4B model is unavailable.

## FAQ

### Does Paritok work with Claude Code, Cursor and Codex without changing my agent?

The README states it is a drop-in proxy for Claude Code, Cursor, Codex, OpenHands and any agent that honors BASE_URL, and that you do not change a line of your agent. The agent points at Paritok instead of Anthropic or OpenAI, and the response flows back unchanged.

### What happens to the content Paritok compresses? Is it lost?

The README describes it as non-destructive: compressed segments are tagged [REF:id] and the agent calls read_original (renamed from expand_context in v1.3.0) to pull back the exact untouched bytes locally. Filtered tool schemas come back through gateway_search_tools.

### Do I need a GPU to run the Paritok gateway?

The embedding tool filter runs entirely on CPU via fastembed and onnxruntime with no torch and no CUDA, and its quantized ONNX weights are bundled in the wheel. The 4B compression model is the part that needs a GPU if you self-host it, and the pyproject description offers the Paritok GPU server as the alternative.

### Which token savings should I expect from Paritok?

The README claims roughly 25% on turn one rising past 85% in long, context-saturated sessions, plus about 3x more turns in the same context window. It also says the tool-schema filter is the largest single-turn lever, taking a typical Claude Code tool block from around 29K to around 8K tokens.

### What licence is Paritok released under?

Apache-2.0, declared in pyproject.toml and in the LICENSE file at the repository root. The bge-small-en-v1.5 embedding model it uses is MIT per the README.

## Sources

- [Issues](https://github.com/Paritok-official/paritok-4b-v1/issues)
- [License: Apache-2.0](https://github.com/Paritok-official/paritok-4b-v1/blob/main/LICENSE)
- [Paritok-official/paritok-4b-v1 on GitHub](https://github.com/Paritok-official/paritok-4b-v1)
- [Project website](https://www.paritok.com/)
- [README](https://github.com/Paritok-official/paritok-4b-v1/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/paritok-official-paritok-4b-v1
