Model or dataset
OnlyTerp/openclaw-optimization-guide avatar
OnlyTerp/openclaw-optimization-guide

OpenClaw optimization guide: the harness thesis and a public retraction

Make your OpenClaw AI agent faster, smarter, and cheaper. Speed optimization, memory architecture, context management, model selection, and one-shot development guide.

388 stars44 forksJavaScriptMIT

At a glance

What is it?
A documentation set rather than code, organized as numbered part files, whose central claim is that most agent capability comes from the harness rather than the model weights. What makes it worth reading is the top-of-file correction, where the author retracts a release number and several configuration keys that an earlier revision of the same guide had invented.
Who is it for?
Read this if you run an OpenClaw agent and want a specific, dated set of opinions rather than general advice, because every claim is tied to a named release and the author flags which of his own earlier claims were wrong. Three caveats to carry with you.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 98 days ago.
What is it written in?
Mainly JavaScript, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 7, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The guide retracts its own earlier claims at the top

The first substantive thing in this repository is a correction, and it is worth more than most of the advice that follows. A July 2026 sweep sets the stable baseline at a specific OpenClaw release, version 2026.6.11, released on 2026-06-30, and covers the trains between two earlier point releases. The same paragraph states that this revision corrects the previous one, which had referenced a release that does not appear to have shipped and several configuration keys that were never real. A second retraction appears further down, in parentheses, where the per-agent cost item notes there is no per-agent budget cap setting and that an earlier revision of the guide claimed one in error. That is a specific kind of documentation failure, the kind that only happens when someone writes plausible-sounding settings from memory rather than from the source, and the fact that both retractions are visible without scrolling to a changelog is unusual. It also tells you how to read the rest: check the claim against the named release before you paste it into a config file.

Per-agent cost, and the advice to cron it and watch for jumps

The cost guidance has become a per-agent measurement rather than a global figure, and the recommendation is to automate the reading of it:

bash
openclaw gateway usage-cost --agent <id>

The same command reports all configured agents at once. The stated practice is to schedule it daily and treat a jump in spend as a signal that a context regression has landed, which is a genuinely useful reframing: a cost spike in an agent system is usually a prompt that grew, a tool that started returning more, or a loop that stopped terminating, and none of those are visible in a monthly invoice. What is not there is any cap. The guide is explicit that no per-agent budget configuration key exists, which means the mechanism is detection after the fact rather than prevention, and that is a real gap for anyone running an agent against a metered API. The related subscription advice in the same section is blunt for a different reason, covered in its own section below.

Four files with hard size caps, injected on every message

The file hierarchy is the most transferable idea in the guide, and the size caps are the mechanism. Three curated files are injected into every single message, and each has a budget: the identity file holding invariants and non-negotiables is capped below one kilobyte and written by a human, the operational rules file covering decision trees and tool routing is capped below two kilobytes and written by a human and the agent in an auditable way, and the durable memory file is capped below three kilobytes and written by the agent through a promotion command. A fourth file, a reflection diary, holds the latest few entries rather than a fixed cap. Read together, the caps are a context budget expressed in the unit an operator can actually enforce, and the read column is what makes the design work: these four are always loaded, so anything that belongs in them has to be worth the tokens every single time. Everything else is either loaded on activation or not loaded at all.

The raw vault and the daily notes stay out of the injected context

The opposite half of the hierarchy is where the volume goes, and separating it is what makes the caps above survivable. Raw source notes, transcripts and links live in a vault directory with no size cap at all, written by automatic capture and by humans, and a separate dated file per day holds a rolling short-term rollup. Neither is read on every message. Both are read when a search is performed, which means the retrieval step is a search over material that did not have to fit in the prompt, and the promotion path is how something moves from a daily rollup into the capped durable memory file. Generated output sits in a third category again, as one-shot artifacts such as pull requests, commits, reports and transcripts, which are written once and read by people rather than by the agent. The three tiers are described as mapping one to one onto a published three-tier note-taking pattern for language models, and the guide does credit that source rather than claiming the structure as its own invention.

The harness thesis, with the percentage explicitly disclaimed

The argument the whole guide rests on is stated as a single line: most agent capability comes from the harness rather than the weights. What is more credible than the claim is the disclaimer attached to it, where the author says the exact percentage is rhetoric and that the point is the operator lesson. Model swaps help, but the larger wins come from seven areas an operator controls, and the diagram that follows enumerates them as layers: the instruction files, context engineering through budgets and progressive disclosure, tools and approvals organised by semantic category, guardrails built from hooks and redaction, the memory layer, and orchestration, which the guide says has five coordination patterns. The closing line of that section is the one to remember when evaluating any of this, that you usually cannot change the weights but you can change everything else. What you cannot change is also what gets benchmarked and marketed, which is the imbalance the guide is correcting for.

Failover only works with at least two configured lanes

The release notes claim that failover now genuinely fires, and the mechanism is more interesting than the claim. Provider overloads are classified correctly rather than being lumped together, and usage-limit responses are routed to fallbacks rather than failing the turn. The useful part is that the fallback list is per job, so an individual scheduled task can carry its own list of alternatives, or can be run strict with an explicitly empty list so that it fails rather than degrading somewhere you did not choose. That is a better design than a global fallback chain, because a nightly report job and an interactive session have very different tolerances for silently using a different model. The caveat the guide attaches is the one to hold onto: none of it helps without at least two configured lanes. A failover mechanism with a single lane is a configuration that fails silently in the only way that matters, which is not failing at all.

Fast mode is automatic, and long sessions got cheaper

Two changes in this batch point the same direction, toward fewer manual interventions. Fast mode became automatic, with a command that runs short conversational turns in the provider's faster lane and returns to normal behaviour for longer work, and the status reporting is described as correct through retries and fallback switches, which is the part that matters since a fast-lane indicator that lies during a fallback is worse than none. The second change is about cost over time: tool-heavy sessions now retain prompt-cache savings as results accumulate. The practical advice that follows is a reversal of the common instinct, to keep sessions long-lived and keep compaction healthy rather than resetting them constantly, because a session that is thrown away throws away whatever caching had built up. Both are version-stamped, one to the earlier point release and one to the baseline, so they are the kind of thing that changes again and needs re-checking.

Stuck channels self-recover, and the file numbering has gaps

The last operational change is the one that retires operational toil: chat channel queues that wedge after a crash or a long-running task now resume on their own, and the guide's conclusion is that restarting the gateway should leave your runbooks, which is a measurable outcome rather than an aspiration. After that, the repository itself has some rough edges worth knowing about. The part files are numbered, and the numbers are not contiguous, with several values absent from the top level. More confusingly, the most recent field guide is in a file named for late April while the text refers to it as the June 2026 field guide, so a link following the name lands you on a document that disagrees with the link text. The infrastructure around the prose is otherwise careful, with a link checker configuration, a markdown lint configuration, a documentation build configuration, and two setup scripts, one per platform. None of that changes whether the advice is good, but it does tell you what the review process does and does not catch.

Editorial conclusion

Read this if you run an OpenClaw agent and want a specific, dated set of opinions rather than general advice, because every claim is tied to a named release and the author flags which of his own earlier claims were wrong. Three caveats to carry with you. It is prose, not software, so nothing here is enforced by the tool itself and the size caps on the context files are conventions you maintain by hand. The version numbers are dated to a June 2026 baseline, and a field guide in a 2026 world ages fast, so treat the release-specific sections as a snapshot. And the author is candid that the headline framing is rhetoric, which is a good sign about the rest. The last push was on 2026-07-01 and there are no tagged releases.

Frequently asked questions

What is the OpenClaw optimization guide?

It is a documentation set rather than a library, MIT licensed and organized as numbered part files covering speed, context engineering, memory architecture, tool permissions, orchestration and production concerns. Its stated stable baseline is OpenClaw 2026.6.11, released 2026-06-30, and it ships a scorecard, templates, benchmark methodology and a symptom-indexed gotchas page.

What does the harness thesis mean in the OpenClaw optimization guide?

It is the claim that most agent capability comes from the harness rather than the model weights, which the author frames as rhetoric rather than a measurable percentage. The operator-controlled layers it names are instructions, context engineering through budgets and progressive disclosure, tools and approvals, guardrails, the memory layer, and orchestration.

How does the OpenClaw optimization guide describe provider failover?

Provider overloads are classified and usage-limit responses are routed to fallbacks, with each scheduled job able to carry its own list of fallbacks or to run strict with an empty list. The guide is explicit that none of it helps without at least two configured lanes, so a single-lane setup fails silently.

What are the SOUL.md, AGENTS.md and MEMORY.md files for in OpenClaw?

They are the curated files injected into every message: identity and invariants under 1 KB, operational rules and tool routing under 2 KB, and durable facts promoted from short-term memory under 3 KB. Raw vault notes and dated daily rollups sit outside that context and are read on a search, and generated artifacts such as reports are one-shot output.

Why does the OpenClaw optimization guide say Claude subscription advice is dead?

Because an Anthropic policy change on April 4 broke the assumption that a Pro or Max subscription covered the agent's usage. The guide's position is to treat Claude as paid API, Bedrock or provider-routed usage unless your own install proves otherwise, and it pairs that with per-agent cost reporting rather than any budget cap, since no such configuration key exists.

Official sources

  1. Issues
  2. License: MIT
  3. OnlyTerp/openclaw-optimization-guide on GitHub
  4. Project website
  5. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/onlyterp-openclaw-optimization-guide.svg)](https://hysenlabs.com/projects/onlyterp-openclaw-optimization-guide)