LongHorizon-Harness: a durable execution loop for Claude Code, Codex and OpenCode
The long-horizon computer-use harness. Run AI agents across desktop apps and the CLI for extended periods while preserving task state and making reliable progress on complex workflows. Features fresh-context execution, durable verified state, independent auditing, recoverable progress, and native Claude Code / Codex / OpenClaw integration.
At a glance
- What is it?
- LongHorizon-Harness wraps an existing computer-use agent in a plan, act, verify and checkpoint loop so a task can survive dozens of hours. Here is how the loop works, how to install it, and where it still falls short.
- Who is it for?
- Adopt LongHorizon-Harness if you already drive a CLI agent such as Claude Code, Codex or OpenCode and your tasks outlive a single context window; the loop gives you checkpointing and an independent audit you would otherwise build yourself. Do not adopt it if you need one-shot answers, a hosted service, or GUI computer use on the DeepSeek backend, which the release notes say is still a later phase.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 42 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The problem LongHorizon-Harness targets: tasks that outlast one context window
A single agent round is bounded by the model's context and by how much work fits in one prompt. Long tasks break that bound in a predictable way: the agent loses the thread, re-derives state it already established, or reports success on work it never checked. The README frames the split directly, saying the model determines what an agent can do in one round while the harness engineers the loop around it.
The project is aimed at people already running a computer-use agent from the terminal or a desktop app. It does not train a model and it does not replace the agent. It sits above Claude Code, Codex, OpenCode and DeepSeek Harness, and turns them into long-running systems by deciding what to do next, verifying the result in the real environment, and preserving accepted progress. If your work fits in one prompt, this layer adds overhead for nothing.
Plan, act, verify, checkpoint: the loop and its three roles
The README's flowchart is the clearest description of the mechanism. The original goal plus verified state feeds a planner that produces the next bounded step. That step is executed on a desktop app or the CLI with fresh context. Verification then inspects files, UI, logs and tests in the real environment. A pass checkpoints verified progress; a failure records evidence and returns to the goal-plus-state node. The loop repeats until the task is complete, and the terminal output is a verified result.
The three roles are described as implementation boundaries inside one loop rather than three agents each growing their own version of the task. That distinction matters in practice: a multi-agent design where each worker keeps its own plan tends to drift, while a single ledger with role-scoped workers keeps one source of truth. Fresh context per step is the other half of the design, and it is what makes the loop survivable across a long run.
Version 0.1.7 changed the end of a run. According to the release notes, a finished run is now a conversation: you read the reply, type a follow-up, and the run continues on its own round ledger instead of replanning from scratch. A message sent mid-round is claimed by the next round, so stopping and continuing does not drop it. The same release added --reasoning-effort for every role, with per-role overrides such as --manager-reasoning-effort forwarded to whichever backend exposes it.
Installing LongHorizon-Harness with pip and starting a run
The package name on PyPI is lh-harness and the console entry point is lh-harness, both defined in pyproject.toml. Python 3.10 or newer is required. The project's one-command install claim is the starting point; the commands below follow the entry point and the Web workbench described in the README.
pip install lh-harness
lh-harness doctorThe doctor command is the diagnostic surface the release notes keep extending, and it is the fastest way to confirm that a backend is reachable before you spend hours on a task. Run it against the agent you intend to use.
The browser workbench is the recommended path in the README. It is a React and FastAPI application, launched by the same binary, and it lets you start a task, choose a backend and model per role, answer approvals, send an instruction mid-run, and stop or restart a run.
lh-harness webThe wheel is built from src/lh_harness, and the Web bundle is shipped as a build artifact rather than source. The pyproject.toml comment notes that artifacts skips a missing path silently, so a checkout without a Node toolchain still installs, but a source tree installed without a prebuilt bundle will not carry the Web UI.
For a terminal-only workflow, the README documents running a task from the command line and selecting a backend with --agent. OpenCode support arrived in v0.1.6 as --agent opencode, running the opencode run prompt form with role-scoped read and write permissions and normalized JSON results. DeepSeek Harness arrived in v0.1.5 as --agent deepseek_harness, running dsh --profile headless with an isolated DSH_HOME.
Where the harness stops: backend gaps, permissions and rollback
The most concrete limitation is stated in the release notes rather than hidden. DeepSeek Harness support is phase one: GUI computer-use and MCP support are listed as following in a later phase. If your task depends on driving desktop applications through DeepSeek Harness, the harness does not cover it yet.
Role isolation is a design constraint, not a free feature. The v0.1.2 notes describe stronger auditor read-only checks and role isolation, which implies the safety of the loop depends on those checks holding. The auditor is meant to inspect rather than mutate, and any configuration that widens an auditor's write scope weakens the verification step the whole design rests on.
The README does not document rollback of a checkpoint, and it does not describe what happens to a partially written file when a worker is force-stopped. A graceful stop escalating to a force stop is mentioned in v0.1.7, but the recovery semantics for a killed mid-write process are not spelled out. Treat checkpointing as forward progress preservation rather than transactional undo until you have verified the behaviour yourself.
Finally, the benchmark claims in the README cover WeaveBench, OSWorld 2.0 and Terminal-Bench 2.1. Those are the project's own evaluations; the harness is not a substitute for measuring your own task mix.
How LongHorizon-Harness differs from running Codex or Claude Code directly
Running Codex or Claude Code directly means the agent owns the whole task: it plans, acts and decides when it is done, inside one context. LongHorizon-Harness keeps the same backends but moves planning, execution and verification into separate role boundaries with their own permissions, and inserts an explicit verification step against the real environment before progress is accepted.
The practical difference shows up in two places. First, context: the harness re-executes each bounded step with fresh context instead of carrying a growing transcript, which is what allows a run to continue for dozens of hours. Second, verification: a direct agent run reports completion from its own reasoning, while the harness checks files, UI, logs and tests before checkpointing. If you only need a fast edit or a single command, direct invocation is simpler and cheaper. The harness earns its cost when a task has many steps and a wrong completion claim is expensive.
Licence, packaging and the cost of keeping up
The project is MIT licensed, with the licence text at LICENSE and declared in pyproject.toml as license = "MIT" with license-files = ["LICENSE"]. MIT is permissive, so redistribution and commercial use are not restricted by the licence itself. That is not legal advice, and anyone shipping the harness inside a product should read the licence text and the licences of the dependencies rather than rely on this summary.
The dependency list is short: packaging, tomli on Python below 3.11, fastapi, uvicorn with the standard extra, and websockets. The test extra pins httpx2, with a comment explaining that starlette imports httpx2 in its TestClient and only falls back to the deprecated httpx. A short dependency list lowers the cost of upgrades, but it also means the Web workbench's FastAPI and websockets versions move with upstream releases.
Upgrade cadence is the real maintenance question. Releases v0.1.5, v0.1.6 and v0.1.7 landed within a week of each other in August 2026, and the last push to the repository was on 2026-08-20. The README itself says the project is iterating rapidly. Plan for frequent upgrades and pin the version you deploy.
What to check before you commit to the harness
Start with lh-harness doctor against the backend you actually intend to use, because backend coverage is uneven: OpenCode and DeepSeek Harness are recent additions and DeepSeek is explicitly phase one. Then run one short task in the Web workbench and watch whether the verification step catches a deliberate mistake, since the auditor is the component the whole loop depends on. If the auditor passes work that is wrong, the checkpoint is worthless.
Also confirm how the harness behaves when a worker is stopped mid-write, because the documentation does not describe rollback. That single behaviour determines whether a long run is recoverable or merely restartable.
Editorial conclusion
Adopt LongHorizon-Harness if you already drive a CLI agent such as Claude Code, Codex or OpenCode and your tasks outlive a single context window; the loop gives you checkpointing and an independent audit you would otherwise build yourself. Do not adopt it if you need one-shot answers, a hosted service, or GUI computer use on the DeepSeek backend, which the release notes say is still a later phase. Before committing, run lh-harness doctor against your chosen backend, confirm which model each role is bound to in the Web workbench, and check whether the MIT licence and the dependency set in pyproject.toml fit your deployment.
Frequently asked questions
What is LongHorizon-Harness and how does it work?
It is a loop-engineering harness that wraps an existing computer-use agent such as Claude Code, Codex, OpenCode or DeepSeek Harness. It repeatedly turns the remaining work into a bounded step, executes it with fresh context, verifies the result in the real environment, and then checkpoints accepted progress or records failure evidence for the next round.
Why is LongHorizon-Harness called a harness?
The README describes it as a Loop Engineering system that provides the durable execution loop around an existing agent rather than training a model or replacing one. The harness owns planning, verification, checkpointing and recovery, while the underlying agent still performs each step.
How do I install LongHorizon-Harness?
Install the Python package with pip install lh-harness on Python 3.10 or newer, then run lh-harness doctor to check the backend. The README recommends launching the browser workbench with lh-harness web.
Which agent backends does LongHorizon-Harness support?
The README lists Claude Code, Codex, OpenCode and DeepSeek Harness. OpenCode support arrived in v0.1.6 as --agent opencode, and DeepSeek Harness arrived in v0.1.5 as --agent deepseek_harness, with GUI computer-use and MCP support noted as coming in a later phase.
What does code harness mean in this project?
In LongHorizon-Harness the term refers to the execution loop around a coding or computer-use agent: plan the next bounded step, act with fresh context, verify against files, UI, logs and tests, then checkpoint or recover. The harness is the surrounding machinery, not the model itself.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/amap-ml-longhorizon-harness)