# UltraCode-Shim is an effort envelope wrapped around a backend you already pay for

> UltraCode-Shim is a local proxy that gives Claude Code's high-effort mode to any model you already have a subscription for, splits the single model slot into an orchestrator and a worker tier, and can route each task by quality score. Its most interesting engineering is not the routing but the four ways it stops a long agent run from stalling.

**OnlyTerp/UltraCode-Shim** — Give Claude Code's ultracode mode to ANY model you already pay for. A tiny local proxy + one config.json. Point your AI at AGENTS.md and it sets itself up.

- Repository: https://github.com/OnlyTerp/UltraCode-Shim
- Stars: 434 · Forks: 45
- Language: Python
- License: MIT
- Published: 2026-09-20 · Updated: 2026-09-20 · Language: en
- Canonical page: https://hysenlabs.com/projects/onlyterp-ultracode-shim

## UltraCode is an envelope, not a model

The premise is stated bluntly, in the form of a question and its answer.

At the API level, the high-effort mode is an effort setting at its highest level, plus adaptive thinking, plus a large output token allowance, plus one system reminder. There is no secret model behind it. That is a falsifiable claim, and the project backs it by pointing at a documentation file containing the breakdown together with what it calls reverse-engineering evidence.

If that is right, the implementation is small in principle: take every request that would go to the vendor, add the envelope, and forward it to a backend that has no idea what UltraCode is. The proxy exists to do exactly that, plus the parts a plain request rewriter would not handle.

Two deployment details make it usable rather than a curiosity. Your existing Claude Code install is left untouched, so nothing about the client you already use changes. And the setup path is an agent-readable file: you point your coding assistant at the repository's AGENTS file and it installs the proxy and writes the configuration, which is a different onboarding model from reading a manual.

The shipped example configuration contains ready-to-use entries for several hosted models, a few aggregators, and local models, with the instruction to keep the ones you have a plan for and delete the rest.

## One model slot becomes two tiers

The design responds to a specific behaviour of the host, and naming it is what makes the feature make sense.

The model menu in Claude Code has a single slot. Separately, the host's dynamic-workflow engine issues most of its background and sub-agent traffic using the stock model regardless of what you selected. So on a large fan-out, the dozens of parallel workers doing the bulk of the work follow your selection for the interactive part and quietly bill a different model for the rest.

The proxy's answer is to make the single slot into two. A launcher presents a two-column selector before the client starts: an orchestrator on the left for the main interactive loop, and a worker on the right for every workflow and task sub-agent. The same choices appear later inside the model menu, where the proxy adds a worker entry for every model you have configured.

The routing decision is structural rather than clever. The main loop carries interactive-only tools such as an ask-the-user tool, and sub-agents never do, so that difference classifies a request without inspecting its content. Pick one model for both and everything runs on it; pick two and the strong model plans while the cheap one fans out. An environment variable turns the split off.

Workers run fully parallel, with a threaded proxy and no artificial concurrency cap, which matters when the point of the split is to fan out.

## A classifier that cannot see price

The automatic router is the part with a real design decision in it, and the decision is what the classifier is not allowed to know.

You nominate a small, cheap model to act as a classifier. For each candidate backend it scores from zero to one how likely that backend is to handle the current task, and it does that by reading a short capability card that you write for each backend yourself. The proxy then routes to the cheapest candidate that clears a quality bar, with the bar defaulting to seven tenths.

The classifier never sees cost. That is the whole point of the cost column existing separately in the configuration: the model doing the scoring cannot learn that an expensive backend is preferable, so it cannot quietly become a bias toward spending more. Cost is applied afterwards, by the proxy, as a tie-break among candidates that already passed the quality bar.

Three more properties make it safe to leave on. Decisions are cached per task, so a repeated step does not re-pay for classification. Any failure falls back to a sensible default rather than breaking the request. And the whole feature is off until you turn it on, though the shipped example file already contains a block you can enable.

The example configuration makes the shape concrete: a cheap candidate at three tenths of a unit of cost with a card about single-file edits and simple refactors, and a frontier candidate at five units with a card about large refactors, hard debugging, and images.

## Four ways the proxy avoids a stalled run

A router that spans many backends inherits every backend's failure modes, and on a long autonomous run one unhandled hiccup can end a forty-minute session. The project lists the three failure modes it hit, plus one more, and each answer is specific.

An empty turn is retried. A backend returning a turn with no text and no tool call is transparently re-issued, which covers both a transient blip and a reasoning turn that ran out of budget at high effort. The retry buffers only until the first real token arrives, so a normal turn adds no measurable latency, and it never retries once real output exists, so output cannot be duplicated.

A silent stream becomes a quick retry rather than a hang. If a stream opens and then goes quiet mid-turn, a bounded idle timeout fires, so one stuck sub-agent no longer stalls an entire multi-agent run for minutes.

A declined tool call is repaired. Some strict backends return an error when the client declines or skips a tool, which the proxy handles by fixing up the tool-call sequence and synthesizing a stub reply for anything unanswered, including half-finished parallel calls. That one is tracked as a numbered issue, which suggests it was reported by a user rather than anticipated.

And reasoning models get a keepalive, because a model that thinks for seconds before its first token would otherwise look frozen to a client watching the connection.

## One proxy file, two installers, and a demo that runs offline

The repository layout is flat enough to read in one screen, and it explains what this project considers its core.

A single proxy module sits at the top level, with a directory of provider adapters beside it, a launcher directory, a checked-in example configuration, a test file for the proxy itself, and the documentation directory holding the guides the README keeps linking. There is also a role pipeline workflow script and an examples directory.

That last part is the one to try first. The auto router ships an offline demonstration that needs no keys:

```bash
python3 examples/auto_router_demo.py
```

Run it and it prints a table of tasks, the classifier scores each candidate produced, which backend was chosen, and what that cost. Five rows show the escalation working, including an image task that only a vision-capable candidate can take and therefore lands on the expensive one, and a repeated task served from cache.

Two installers and a directory named for Windows make the intended platform coverage explicit. A shell installer and a PowerShell installer sitting side by side is a stronger signal than a claim in prose, and it is the sort of detail that saves you from discovering a gap after installing.

The licence is MIT and there are no tagged releases, so nothing here is versioned in the way you would pin a dependency. The last push is dated 2026-08-22.

## Ten backends in the example, and one that cannot work

The example configuration is worth reading as a compatibility list, because its omissions are as informative as its entries.

It ships ready-to-use entries for a hosted frontier model reached through a different login flow, several other hosted models including a mixture of general and fast tiers, two model families offered in a larger and a faster variant, a cloud-hosted local-model service, a subscription-based coding tool, two aggregators, and local models. The instruction accompanying it is to keep the ones you have a plan for and delete the rest.

That last clause is the honest bit. A configuration file with ten provider entries is a convenience and a liability, and telling users to prune it is better than shipping dead credentials.

The omission is a specific one: one coding assistant's composer mode is excluded, with the reason given as needing its own command line client and not being reachable over HTTP. Since the proxy's entire mechanism is HTTP request rewriting, that is a structural exclusion rather than a missing integration, and it is documented in a separate guide for adding a model.

The same guide is where you would look to add something the example does not cover, which makes the unsupported case the best-documented one in the repository.

## Conclusion

UltraCode-Shim suits someone with a subscription to a model Claude Code cannot talk to, who wants the high-effort harness and is willing to debug a proxy in the middle. It does not suit anyone treating model choice as invisible, because the routing decisions are the proxy's and the documentation is candid that some backends cannot decline a tool call natively. Before pointing it at a real workload, read how credentials are stored, since one config file ends up holding keys for every backend you keep, and try the offline router demo before spending anything.

## FAQ

### What is Ultracode in Claude Code and how is it used?

The project states that at the API level UltraCode is an effort setting at its highest level plus adaptive thinking, a large output token allowance, and one system reminder, with no secret model behind it. The proxy adds that envelope to every request so any backend gets the same treatment.

### How does UltraCode-Shim support two different models at once?

It turns the single-slot model menu into two tiers: an orchestrator for the main interactive loop and a worker for every workflow or task sub-agent. Requests are classified by a structural signal, since the main loop carries interactive-only tools that sub-agents never use.

### How does the UltraCode-Shim auto router choose a model?

A cheap classifier model you nominate scores each candidate from zero to one against a short capability card you write, and the proxy routes to the cheapest candidate clearing a quality bar that defaults to 0.7. The classifier never sees price, decisions are cached per task, and any failure falls back rather than breaking the request.

### What happens when a backend fails during a long UltraCode-Shim run?

Empty turns are re-issued with buffering only until the first real token, a stalled stream hits a bounded idle timeout instead of hanging, and a declined tool call is repaired with a synthesized stub reply for anything unanswered, including partial parallel calls.

### Which backends ship in the UltraCode-Shim example config?

Entries for several hosted models, aggregators, a cloud-hosted local-model service, and local models, with the instruction to keep the ones you have a plan for and delete the rest. One coding assistant's composer mode is excluded because it needs its own command line client and is not HTTP based.

## Sources

- [Issues](https://github.com/OnlyTerp/UltraCode-Shim/issues)
- [License: MIT](https://github.com/OnlyTerp/UltraCode-Shim/blob/main/LICENSE)
- [OnlyTerp/UltraCode-Shim on GitHub](https://github.com/OnlyTerp/UltraCode-Shim)
- [README](https://github.com/OnlyTerp/UltraCode-Shim/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/onlyterp-ultracode-shim
