fast-jev-compaction: verbatim context compaction for Claude Code
Claude Code plugin that replaces the compaction summary with Jev decisions: every tool call and result is scored in one fast request, stale ones are dropped or truncated, everything kept stays verbatim.
At a glance
- What is it?
- A Claude Code plugin and npm library that drops or truncates stale tool calls instead of summarizing them, keeping every retained message word for word. It depends on TypeSafe's Jev model and throws when the history will not fit.
- Who is it for?
- Adopt it if you run long Claude Code sessions where exact file paths, error strings and command output must survive compaction, and you are willing to send the whole conversation to TypeSafe's Jev endpoint with a TYPESAFE_API_KEY. Do not adopt it if your history is mostly prose, if you cannot send transcripts to a third-party API, or if you need compaction that never throws.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 13 days ago.
- What is it written in?
- Mainly TypeScript, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The lossy summary problem fast-jev-compaction targets
Most context compaction hands old turns to an LLM and asks for a summary. The README states the objection plainly: a summary is lossy, and a file path, an exact error, a constraint or a command can disappear even when it matters later. That is the failure this project is built around. An agent that was told never to edit src/generated may keep working for another twenty turns; if the constraint is paraphrased away during compaction, nothing in the transcript flags the loss.
The audience is narrow and specific. It is for people running Claude Code sessions long enough to hit compaction, and for developers building their own agent loop who already have a Message-shaped transcript and want a compaction step they can call directly. The package is TypeScript, ships as ESM only ("type": "module"), and requires Node 18 or newer. It is also a Claude Code plugin: the repository root holds hooks/ and .claude-plugin/, and the README describes the root as a function-hook plugin whose adapter feeds session.compact transcripts through src/. If you are not using Claude Code and not writing TypeScript, the library form is the only entry point.
How Jev decisions replace the summary
The mechanism is deletion, not rewriting. Every tool_use is paired with its tool_result by tool_use_id. Calls in the first message, and calls in the newest preserveRecentMessages messages (6 by default), are pinned and never touched. Everything else becomes a candidate.
The state sent to Jev is the whole conversation, oldest first, with each tool result replaced by a short note such as "ok, 4213 chars (omitted)". Tool inputs and text are included as they are. Nothing is summarized before the model sees it, which is the design bet: Jev judges the real conversation, not a digest of it.
For each non-pinned call, Jev receives two questions. Should the call stay, knowing it was made with that input still matters? Should the result stay verbatim, meaning its contents are still needed and re-running the tool would not reproduce them? The answers are compared against keepThreshold, 0.5 by default. A keepResult at or above the threshold keeps call and result. Otherwise a keepCall at or above it keeps the call and truncates the result to its first truncateHeadChars characters (300 by default) plus a one-line note. Otherwise both go.
Rebuilding is conservative in one direction: a message that loses all its content is removed, untouched messages are returned as the same objects, and no result is ever left without its call. The trade-off is the opposite of a summary. A summary compresses everything and keeps the shape of the conversation; this keeps selected turns byte-for-byte and removes others entirely.
Fitting the state under maxStateTokens
Jev has a 32k request limit, and the full state is resent with every request. That makes fitting the state the most mechanical part of the project, and the README documents it as a staged sequence where each stage applies only if the previous one did not bring the estimate under maxStateTokens (25,000 by default).
The stages degrade in this order: tool inputs truncated to 1000, then 200, then 60 characters; long texts abridged to head plus tail, oldest non-pinned messages first; old non-pinned messages collapsed to a "[… N chars omitted …]" note; old tool calls reduced to one line each, in the README's example format "t12 Read file_path=src/a.ts → ok 480ch"; old call-less messages left out; runs of old call-only messages folded into one entry. If the history still does not fit, compaction throws rather than proceeding with a state it cannot send.
Token counts are estimates, not tokenizer output. The README describes the heuristic: a word per six letters, half a token per digit, roughly one per other symbol, calibrated to land a little above the counts Jev reports. That calibration direction is deliberate, since overshooting the estimate is safer than undershooting the 32k limit, but it means result.stats reports estimated tokens, not measured ones. Questions are then split into as many requests as needed to keep state plus questions under maxRequestTokens (30,000 by default), the same full state is resent each time, and the requests run concurrently with their answers merged. A history near the state ceiling therefore costs one request per handful of questions, which the README lists as a limitation rather than hiding.
Installing fast-jev-compaction and running it on a transcript
The package installs from npm and reads the API key from the environment. The README warns against committing the key or putting it in a source file.
npm install fast-jev-compaction
export TYPESAFE_API_KEY=...The library entry point is compactMessages, which takes a Message array and options. The README's example builds a small transcript with one user turn, one assistant turn carrying a Read tool_use, and a user turn carrying the matching tool_result, then compacts it while preserving the newest 4 messages.
import { compactMessages, reductionRatio, type Message } from 'fast-jev-compaction';
const result = await compactMessages(transcript, { preserveRecentMessages: 4 });
console.log(result.messages, result.decisions, result.stats);
if (reductionRatio(result) < 0.25) {
// not worth it: keep the original transcript, or summarize instead
}After the call you get three things: the rebuilt message list, the per-call decisions, and stats covering message and character counts before and after, per-reason decision counts, the state size in estimated tokens, which fitting stage was needed, and the number of requests. Message is a subset of Claude Code's SessionMessage, so a session transcript can be passed in as is. If you already have a transport, implement JevAsker with its single ask(state, questions) method and call compact(messages, asker, options); buildJevRequest and parseJevResponse produce the HTTP body and validate the response. apiKey defaults to process.env.TYPESAFE_API_KEY, model defaults to jev-latest, and baseUrl defaults to https://api.typesafe.ai/v1/systemone.
For Claude Code itself, function hooks are an early-access feature (2.1.274 or newer), so the opt-in flag has to be set wherever Claude Code runs. The README gives ~/.claude/settings.json as an example location.
{ "env": { "CLAUDE_CODE_ENABLE_FUNCTION_HOOKS": "1", "TYPESAFE_API_KEY": "<your key>" } }The README then says to add the repository as a plugin marketplace and install the plugin, either from the shell or as slash commands inside a session, and shows the beginning of the shell route. The README as provided is truncated at that point, so the exact install command is not documented here. Once installed, hooks/fast-jev.ts feeds session.compact transcripts through src/ and falls back to Claude Code's built-in summary on errors or insufficient reduction.
Where fast-jev-compaction fails or is the wrong tool
Failures throw. Jev failures, malformed answers, a missing key, or a history that cannot be fitted all raise, and the README says the caller or the Claude Code hook decides what to fall back to. In the plugin path that fallback is Claude Code's built-in summary. In the library path you own it, and the README's example suggests keeping the original transcript or summarizing instead.
The scope limit is the sharper one. Only tool calls and results are candidates; text messages are never removed or shortened in the output, only abridged in the state Jev sees. A session dominated by long prose, planning discussion, or pasted documents gets almost nothing back. The reductionRatio check in the README exists for exactly that case.
Second, a probability is not a proof. The README concedes that calibration is at the request level and that a keep probability does not establish a result is safe to delete. The stated mitigation is that the assistant can always re-run the tool, which holds for a Read or a Grep and fails for anything with side effects or a result that cannot be reproduced, such as output from a command that mutated state. The truncation path has the same character: keeping the first 300 characters of a result plus a note is a real loss for stack traces where the last line carries the error.
Third, the full state goes to TypeSafe's System One endpoint on every request. If the transcript cannot leave your network, this design is unusable, and no local option is documented.
How it differs from summarization-based compaction
The obvious alternative is the compaction Claude Code already ships: ask a model to summarize old turns. The difference in approach is not tuning, it is what survives. A summary produces a new, shorter representation of everything; fast-jev-compaction produces a subset of the original objects, with no rewriting of what stays. If your work depends on exact strings (a file path, a compiler error, a flag), the subset model preserves them by construction and the summary model preserves them only if the summarizer chose to.
The cost side runs the other way. A summary makes one model call over the history and returns text you can inspect; this project sends the full state with every batch of questions, so a long history costs several requests, each carrying the whole conversation. The README states the full state is repeated with every request and lists the resulting request count as a limitation. A summarizer also degrades gracefully when the history is huge, because it compresses rather than refuses; fast-jev-compaction throws when the staged fitting cannot get under maxStateTokens. And a summarizer needs no extra vendor: this project requires a TYPESAFE_API_KEY and talks to https://api.typesafe.ai/v1/systemone.
If you want to keep the summarizer but avoid the loss, the middle path is the one the README's own example implies: run compactMessages, look at reductionRatio, and fall back to the summary when the reduction is small. That gives up the verbatim guarantee only in cases where the verbatim path was not buying much.
Licence, maintenance and the cost of upgrading
The licence is MIT, declared in package.json and shipped as LICENSE in the published files. MIT permits commercial use and modification; it also means the authors extend no warranty, which matters here because the failure mode is silent context loss if you misjudge keepThreshold. That is a description of the licence terms, not legal advice.
On maintenance, the repository is not archived, and the last push was on 2026-09-18, four days before this writing. There are no retrieved releases, so version tracking runs through package.json, which is at 0.2.0, and through npm. The 0.x version number is the honest signal: the options table includes defaults such as keepThreshold and truncateHeadChars that you will want to tune per project, and a minor bump can move them.
Upgrade cost concentrates in three places. The plugin path depends on a Claude Code feature marked early-access at 2.1.274 or newer, and hooks/README.md carries a type reference for that version, so a Claude Code release can invalidate the adapter. The library path depends on the shape of Message, documented as a subset of SessionMessage. And the fitting stages are tuned against Jev's 32k request limit, so a change to that limit or to the model name jev-latest changes the arithmetic. The package has no runtime dependencies listed in package.json, only devDependencies (typescript, tsx, vitest, @types/node), which keeps the upgrade surface small.
Editorial conclusion
Adopt it if you run long Claude Code sessions where exact file paths, error strings and command output must survive compaction, and you are willing to send the whole conversation to TypeSafe's Jev endpoint with a TYPESAFE_API_KEY. Do not adopt it if your history is mostly prose, if you cannot send transcripts to a third-party API, or if you need compaction that never throws. Before wiring it into a session, run the library form on a real transcript and check reductionRatio(result); if it comes back below 0.25 the README's own example says the reduction was not worth it.
Frequently asked questions
What is compaction in AI agents, and how does fast-jev-compaction handle it?
Compaction is the step that shortens a long conversation before it exceeds the model's context window. Most implementations ask an LLM to summarize old turns; fast-jev-compaction instead scores each tool call and result with Jev and deletes or truncates the stale ones, leaving everything it keeps verbatim.
Does fast-jev-compaction work for text-only conversations?
Not usefully. The README states that only tool calls and results are candidates, and text messages are never removed or shortened in the output. A session made mostly of prose will see little reduction, which is why the README's example checks reductionRatio and falls back when it is below 0.25.
What happens when fast-jev-compaction cannot fit the history?
Compaction throws. The README lists Jev failures, malformed answers, a missing key, and a history that cannot be fitted as throwing conditions, and says the caller or the Claude Code hook decides what to fall back to; the plugin falls back to Claude Code's built-in summary.
Which API key and endpoint does fast-jev-compaction use?
apiKey defaults to process.env.TYPESAFE_API_KEY, and baseUrl defaults to https://api.typesafe.ai/v1/systemone, with model defaulting to jev-latest. The README says never to commit the key or put it in a source file.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/tamaratran-fast-jev-compaction)