# autocontext ships a default provider whose value is the word deterministic

> A recursive harness that runs a task against an evaluation, keeps the lessons that helped, discards the rest, and leaves the traces on disk. The configuration template it ships sets the agent provider to deterministic and the executor to local, while the page's own thirty-second run sets a provider by hand. Four install surfaces carry three different version numbers.

**greyhaven-ai/autocontext** — a recursive self-improving harness designed to help your agents (and future iterations of those agents) succeed on any task

- Repository: https://github.com/greyhaven-ai/autocontext
- Stars: 1,303 · Forks: 114
- Language: Python
- License: Apache-2.0
- Published: 2026-09-15 · Updated: 2026-09-15 · Language: en
- Canonical page: https://hysenlabs.com/projects/greyhaven-ai-autocontext

## Four install surfaces carrying three version numbers

The install table has four rows and three distinct versions. The Python CLI and the Python library both pin 0.19.1, the TypeScript and Node CLI pins 0.19.0, and the Pi extension pins 0.12.0. So a project that presents itself as one harness is published as three release lines, and someone who installs the Pi extension next to the Python CLI ends up seven minor versions apart. The table also has to warn about a name collision, because the PyPI package is autocontext, the command is autoctx, and the npm package is autoctx and not the unrelated npm package that happens to carry the same name as the Python one. Two language runtimes, one command, and a third-party package to steer around. The Node surface also carries its own floor, requiring version 22.19.0 or newer, with contributors pointed at a separate pin held in ts/.nvmrc.

```bash
uv tool install autocontext==0.19.1
```

## The shipped configuration names a stub as the provider

Everything is configured through environment variables, and the template the project ships sets the agent provider to the literal value deterministic and the executor to local. That default is a stub rather than a model. Meanwhile the thirty-second run near the top of the page sets the provider by hand, choosing pi, which the page calls the lowest-friction option because it reuses your local agent auth.

```bash
AUTOCONTEXT_AGENT_PROVIDER=pi \
AUTOCONTEXT_PI_COMMAND=pi \
autoctx solve "improve customer-support replies for billing disputes" --iterations 3
```

So the file a new user copies and the example a new user runs describe different systems, and the gap between them is the gap between a loop that consults a model and one that does not. Other providers are named as values of the same variable, including anthropic, openai-compatible, openrouter, claude-cli, codex and pi-rpc. Set it before you draw a conclusion from a first run.

## Six model roles, one variable each, one precedence rule

The template carries six model variables, one per role: competitor, analyst, coach, architect, translator and curator. All six sit commented out, and the sample values split across two models, with coach and architect shown on claude-opus-5 and the other four on claude-sonnet-5, which is a sensible division between the roles that propose changes and the roles that judge them. The rule that resolves conflicts between them is stated in comments rather than in prose. Leave the variables unset and they inherit the selected provider's defaults. On a local endpoint, AUTOCONTEXT_LOCAL_MODEL fills every role and tier slot that is otherwise unset. Explicit per-role or per-tier values take precedence over that. Three levels of granularity, and the precedence rule lives inside a template file that a user has to open and read line by line.

## Two credential names per provider, and one base URL the SDK cannot carry

The credentials follow the same doubling. There is ANTHROPIC_API_KEY and an AUTOCONTEXT_ANTHROPIC_API_KEY kept as a compatibility alias, then OPENROUTER_API_KEY and an AUTOCONTEXT_OPENROUTER_API_KEY, and the template says plainly to prefer the provider-native names. A third provider adds an API base URL, marked compatibility-only, with a comment giving the reason: the installed SDK cannot safely receive a per-client custom base URL. That is a real constraint stated without embarrassment, and it is the sort of thing a user discovers at three in the morning if nobody wrote it down. The other switch worth knowing is AUTOCONTEXT_CONSTRAINED_OUTPUT, true by default because OpenAI-compatible role generation uses output schemas by default, and set to false it sends ordinary unconstrained requests and parses the returned Markdown instead. One boolean decides whether six roles emit objects or prose.

## Three records for one run, two of them inside the repository

A run leaves two directory trees and a database. Under runs/<run_id> there is a trace as JSON lines, a generations folder where attempt n holds a strategy in JSON, an analysis in Markdown and a score in JSON, a report, and an artifacts folder. Under knowledge/<scenario> there is a playbook, hints, a tools folder, and a context_bundles folder holding bundles, candidates, promotions and an active file. Alongside both, the configuration points a database at runs/autocontext.sqlite3 and an event stream at runs/events.ndjson. Calling this filesystem-first is fair in the useful sense: every artefact is a file you can read, diff and replay. It also means three record formats for one run, each with its own configurable path, and neither the template nor the page says which one wins when they disagree. Both trees are committed at the repository root, so the accumulated knowledge and the raw traces travel together.

## A context change cannot be served until a matched trial says so

Context changes do not take effect when they are written. Changes from the coach and architect roles are stored as immutable candidates and are not served until matched candidate and incumbent trials confirm them. Serving can additionally require a cancellable independent audit, plus a durable budget for false promotions spread across the whole campaign, and causal credit is granted only for verified single-component manifest additions, so a change touching more than one thing is not credited at all. Attribution runs through controlled component trials so that prompt selection can demote low-value context without presenting a correlation with edit size as a cause. The schema-migration study compares a baseline, textual context, executable reuse and cheaper-model controls under frozen paired evaluation, and the page takes the trouble to add that fixture runs are infrastructure evidence only. That is a stricter experimental standard than most agent tooling publishes, stated in the same place as the feature it guards.

## One loop pointed at three different research problems

The same machinery is aimed at three problems. The first is agent task improvement, entered through solve for a plain-language goal and run for a saved scenario. The second is context bundle promotion, entered through immutable candidates and matched trials. The third is GPU kernel evolution, where Python composes bounded studies across variable-shape matmul, fused elementwise and reduction, and causal-attention families. Each kernel family keeps independent primary and confirmation evidence plus per-case floors, and cross-shape, cross-hardware and cross-family trials sort results into portable, partially transferring, specialist and plateau, specifically so that no aggregate score can hide a workload that failed. It is worth sitting with the fact that one promotion rule is being asked to answer both should this prompt help and is this kernel faster at this shape. The documentation does not claim the two transfer cleanly between each other.

## Paid accelerator work is claimed before it is returned

Remote execution on paid providers carries the most careful wording on the page, and that wording is the documentation. Accelerator requests are opt-in and validated for type and count, immutable image, region and telemetry capability, and the page states that a run fails before provider creation when the configured pool cannot satisfy the request, and never downgrades accelerator work to CPU. On the billing side, the shipped Prime generation and campaign paths persist a durable pre-dispatch claim plus the complete result and ledger projection before returning paid work, and a restart never treats an unresolved or already committed request as permission to provision another sandbox. Read as a whole, that is a set of invariants about money rather than a list of features: refuse the request rather than substitute hardware, write the claim before the work, and refuse to re-provision on a restart. Kernel campaigns extend the same contract with exact provider-generation receipts, bounded paid-call accounting, content-addressed lineage and safe stop, status and resume.

## Conclusion

The design here is more careful than the packaging, and both halves are worth reading before you commit. The parts to admire are the ones that refuse to be sloppy: a change is not served until a matched trial confirms it, causal credit is withheld unless exactly one component changed, paid accelerator work is claimed durably before it is returned and never silently downgraded, and a failed workload cannot hide behind an average. The parts to watch are the seams. Set the provider explicitly before trusting a first run, because the shipped template will not do it for you. Pin your Python, npm and Pi versions independently, since three release lines are in flight. And decide up front whether you are using this to improve prompts, to improve GPU kernels, or both, because the same promotion rule is being asked to answer both questions and the documentation does not claim it transfers cleanly between them.

## FAQ

### How do you install autocontext?

Four surfaces are documented and they are not at the same version: the Python CLI and library pin 0.19.1, the TypeScript and Node CLI pins 0.19.0, and the Pi extension pins 0.12.0. The command is autoctx on both the Python and npm sides, and the npm package autoctx is not the unrelated package called autocontext.

### Which model provider does autocontext use by default?

The shipped environment template sets AUTOCONTEXT_AGENT_PROVIDER to deterministic, a stub, with the executor mode set to local. The quickstart on the page instead sets the provider to pi explicitly, and anthropic, openai-compatible, openrouter, claude-cli, codex and pi-rpc are also available as values.

### What does an autocontext run leave behind?

Two trees and a database. A run directory holds trace.jsonl, per-generation strategy.json, analysis.md and score.json, report.md and an artifacts folder, while a knowledge directory per scenario holds playbook.md, hints.md, a tools folder and context bundles. The configuration also points at runs/autocontext.sqlite3 and an events.ndjson stream.

### How does autocontext decide to serve a context change?

Coach and architect changes are stored as immutable candidates and are not served until matched candidate and incumbent trials confirm them. Serving can also require a cancellable independent audit and a campaign-wide false-promotion budget, and causal credit is only granted for verified single-component manifest additions.

### Can autocontext run models on your own hardware?

Yes, on vLLM, Ollama or any OpenAI-compatible endpoint. A local endpoint can set AUTOCONTEXT_LOCAL_MODEL to fill every otherwise-unset role and tier slot, and can declare AUTOCONTEXT_PROVIDER_HOSTING=local with a capability of fast, mid_tier or frontier. Paid remote execution refuses an accelerator request it cannot satisfy rather than downgrading it to CPU.

## Sources

- [greyhaven-ai/autocontext on GitHub](https://github.com/greyhaven-ai/autocontext)
- [Issues](https://github.com/greyhaven-ai/autocontext/issues)
- [License: Apache-2.0](https://github.com/greyhaven-ai/autocontext/blob/main/LICENSE)
- [README](https://github.com/greyhaven-ai/autocontext/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/greyhaven-ai-autocontext
