Model or dataset
fendouai/CodexSaver avatar
fendouai/CodexSaver

CodexSaver's manifest declares zero dependencies and no licence

Make Codex cheaper without making it dumber with DeepSeek.

645 stars50 forksPythonLicense varies

At a glance

What is it?
CodexSaver is an MCP tool that routes low-risk work away from Codex to cheaper workers, keeping judgment and final review in place. Its manifest declares an empty dependency list and no licence field, and its headline cost figures come from five self-run tasks.
Who is it for?
CodexSaver suits someone already paying per token for Codex whose work is mostly explanation, scanning, docs, and bounded tests, and who wants those delegated without giving up review. It does not suit anyone whose codebase work is dominated by auth, payments, migrations, or architecture calls, since the tool is explicit that it loses there.
Can I use it commercially?
Not without permission. GitHub finds no licence file in the repository, and without a licence all rights are reserved by default: you may read the code but not reuse it. Check the README, or ask the authors, before using it.
Is it still maintained?
Yes. The repository last received commits 127 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 5, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The description still leads with DeepSeek while the default worker changed

The repository description reads Make Codex cheaper without making it dumber with DeepSeek. The readme's own headline drops the last four words, and the body explains why. Pi Agent is the default v3 worker, with automatic discovery for other local Agent Card workers, while DeepSeek and other providers remain available for legacy v1 and v2 provider-backed delegation. So the sentence most people see first describes an architecture the project has already moved past, and the thing that actually runs by default is a locally installed agent rather than a hosted API. The local provider configuration is a one-time file edit, which also means the tool has no credentials of its own to leak and nothing to configure in every workspace beyond the global Codex install it relies on.

The manifest declares zero dependencies and no licence

The packaging metadata is unusually bare. The manifest lists a name, a version of 0.3.6 matching the only release, a one-line description, a Python floor of 3.10, and an empty dependency list. That last one is the striking entry: a tool whose whole purpose is calling model providers and serving an MCP interface declares nothing to install. It also has no build-system table, so the backend is whatever a tool assumes by default rather than something pinned, and no licence field at all, which is why the hosting site shows no licence for the repository. The distribution is two top-level modules alongside the package: one is the console entry point named codexsaver, the other is the MCP server module. Setup instructions are deferred to a separate developer readme.

The evidence is five tasks, run by the project itself

The headline numbers come from a table labelled as repo-local evidence, and they are specific. On five bounded tasks the v2 benchmark is reported at 45% to 100% estimated savings. The five-task run completed successful tasks between 0.03 and 14.95 seconds, and a v3 readonly swarm succeeded in 6.45 seconds. Quality is given as bounded work packets passing verifier gates, plus a v3 swarm that produced 10 findings at a 0.75 quality score. Those are honest, narrow claims, and the readme does not pretend otherwise: it says the clearest proven v3 advantage today is readonly specialist orchestration, and that the next advantage, higher patch success through verified repair, is still to be proven. Savings are estimated, on a handful of tasks, by the project that wrote the router.

A satisfied task returns before a worker call is paid for

The v2 lane hands the worker a bounded work packet instead of an open instruction, and the packet has six fields: an exact goal, allowed files or globs, forbidden paths, acceptance criteria, allowlisted commands, and caps on iterations and diff size. The worker may propose a patch, but the patch is applied only inside a temporary sandbox and accepted only if it stays within policy and the allowlisted checks pass. One branch of that flow deserves attention. If the task is already satisfied when the packet is checked, v2 returns a preflight_satisfied result without spending a worker model call at all. For a tool whose argument is cost, having a zero-cost path for work that needs no work is the difference between a router and a wrapper, and it is documented as a normal return value rather than an optimisation.

Three routing states, and one of them costs nothing either

Tool responses carry an interaction block instead of a silent blob of JSON, which is what makes a routing decision inspectable after the fact. It names the tool, the mode, a headline sentence, a route label that encodes the route, the task type, and a risk level, and a next step telling the reviewer what to do with the result. Three states matter. Preview is a routing preview with no external model call, delegated_execution means a delegated run completed, and codex_takeover means the task stayed with Codex because the risk was too high or the request was ambiguous. When worker-output compression is switched on, the same block reports the active compression level, so a terse delegated reply explains itself instead of looking like a truncated answer. The MCP entry point for the packet lane is a single tool name.

The command line has two flags that look identical

The v2 lane is also reachable from a shell, and the flag surface is where the rough edge sits. The example invocation passes a goal, a workspace, an acceptance criterion, an allowlisted command, and two separate file flags:

code
--files README.md
--allowed-file docs/v2-smoke.md

Their difference is not explained anywhere in the visible documentation, and the example compounds the confusion by passing one filename to the plural flag while the acceptance criterion and the allowlisted command both concern a different file. That may well be deliberate, since a work packet plausibly needs both a read scope and a write scope, but a reader has to infer it. The other four fields are self-describing, and the command the allowlist example uses is a single Python one-liner that asserts the file exists inside the sandbox.

v3 aggregation falls back to Codex when two patches collide

The orchestration lane keeps the earlier guarantees rather than replacing them. Codex still owns judgment and final review, CodexSaver plans a work graph, readonly specialists run in parallel, and bounded patch specialists reuse the same sandbox and verifier path v2 established. The part that decides whether parallelism is safe is the aggregation rule, which the readme describes as conservative and falls back to Codex on overlap. That is the load-bearing sentence in the whole design: when two specialists propose changes to the same place, the tool does not try to merge them, it hands the conflict back to the model that was supposed to keep the judgment. The stated speed advantage follows from the same shape, since total latency becomes the slowest specialist plus orchestration overhead instead of the sum of everything.

What it refuses is listed as precisely as what it claims

The strongest part of the readme is the pair of lists. The winning lanes are code explanation, repository scanning, performance hinting, docstrings and readme maintenance, bounded test generation, and small bounded refactors with an explicit file scope. The list of what it is not strongest at is auth, security, payment and permissions; destructive migrations; ambiguous architecture decisions; multi-file behavioural change with weak verification; and anything that still needs model-level judgment at every step. The shared criterion is verifiability rather than difficulty, which is a defensible line: the tool wins where a verifier can decide whether the answer is right without asking the expensive model. The thesis block reduces the same idea to three sentences:

code
CodexSaver wins first on readonly specialist orchestration.
CodexSaver wins second on bounded, verifiable patch work.
Codex remains the judge for everything risky or unclear.

Editorial conclusion

CodexSaver suits someone already paying per token for Codex whose work is mostly explanation, scanning, docs, and bounded tests, and who wants those delegated without giving up review. It does not suit anyone whose codebase work is dominated by auth, payments, migrations, or architecture calls, since the tool is explicit that it loses there. Before adopting it, set the worker in the local config file, check that the protected-path list covers the directories you actually care about, run a work packet with a deliberately tight acceptance criterion to see the preflight path fire, and remember the benchmark is five tasks run by the project itself rather than an independent measurement.

Frequently asked questions

How can I save tokens in Codex?

Route the work that does not need judgment to a cheaper worker. CodexSaver does that through Codex as an MCP tool: readonly specialists such as explanation, scanning, and review first, bounded and verifiable patch work second, while Codex keeps architecture, security, protected domains, and final review.

What is the default worker for CodexSaver?

Pi Agent is the default v3 worker, with automatic discovery for other local Agent Card workers. DeepSeek and other providers remain available for legacy v1 and v2 provider-backed delegation, and the local provider setup happens once in a config file under the user's home directory.

How does CodexSaver decide whether to delegate a task?

A router judges whether the task is safe enough. Tool responses then report one of three states: preview for a routing decision with no model call, delegated_execution for a completed delegated run, or codex_takeover when the risk was too high or the request was ambiguous.

What can a CodexSaver work packet restrict?

Six things: an exact goal, allowed files or globs, forbidden paths, acceptance criteria, allowlisted commands, and caps on iterations and diff size. The worker proposes a patch, it is applied only inside a temporary sandbox, and it is accepted only when it stays within policy and the allowlisted checks pass.

What results does CodexSaver report for cost and speed?

As repo-local evidence on five bounded tasks: 45% to 100% estimated savings, successful tasks completing between 0.03 and 14.95 seconds, and a v3 readonly swarm succeeding in 6.45 seconds with 10 findings at a 0.75 quality score. These are the project's own measurements.

Official sources

  1. fendouai/CodexSaver on GitHub
  2. Issues
  3. README
  4. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/fendouai-codexsaver.svg)](https://hysenlabs.com/projects/fendouai-codexsaver)