Harbor: a Python harness for evaluating agents in sandboxed environments
Framework for evaluating and improving agents
At a glance
- What is it?
- Harbor is the official harness for Terminal-Bench-2.0 and a general runner for agent and model evaluations. It installs as a CLI, runs agents such as Claude Code against datasets, and can spread jobs across cloud sandbox providers.
- Who is it for?
- Adopt Harbor if you already run agent evaluations and want one CLI that covers local Docker and hosted sandbox providers, especially if you need Terminal-Bench-2.0 results. Do not adopt it if your evaluation needs are a single script over a fixed fixture set, because the dataset, agent and provider abstractions are overhead you will not use.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository received new commits within the last day.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What Harbor is for, and who ends up using it
Harbor is a framework for evaluating and optimizing agents and language models, published by the creators of Terminal-Bench. The README lists four things you can do with it: evaluate arbitrary agents such as Claude Code, OpenHands and Codex CLI; build and share your own benchmarks and environments; run experiments in many environments in parallel through providers like Daytona, Modal, LangSmith, Blaxel, Novita Sandbox, Tensorlake and Runta; and generate rollouts for RL optimization.
The audience is narrower than that list suggests. If you are comparing two models on a chat task, Harbor is the wrong shape: it is built around agents that act inside a container, receive a task, and produce a result the harness can score. The people who get value are evaluation engineers maintaining a benchmark suite, and research teams that need many parallel rollouts rather than a handful of scored prompts. The RL rollout use case points the same way: you want volume and reproducibility, not a leaderboard screenshot.
The project is not archived, and its last push was on 2026-09-22. Releases are frequent: v0.21.0 on 2026-08-10, v0.22.0 on 2026-08-22, v0.23.0 on 2026-09-12. That cadence is a real cost as well as a signal, because a CLI whose flags and dataset references move this quickly will occasionally break the command you saved three weeks ago.
The mechanism: datasets, agents, models and sandbox providers
Harbor separates four things that are usually tangled together in evaluation scripts: the dataset, the agent, the model, and the environment the run happens in. The CLI takes all four as arguments. A dataset is named with an at-sign version, for example [email protected]. An agent is named separately, for example claude-code. A model is named with a provider prefix, for example anthropic/claude-opus-4-1. The environment defaults to local Docker and can be switched with a flag.
That separation is the whole design. The same dataset can be run against a different agent without editing task definitions, and the same agent and dataset can move from a laptop to a hosted provider by changing one argument. The repository layout supports the reading: adapters/ holds agent integrations, packages/ holds workspace members such as harbor-rewardkit and harbor-langsmith, and examples/ is split into agents, configs, exec, jobs, metrics, prompts and tasks. Those directories suggest the intended extension points are agent adapters and task definitions rather than a single monolithic runner.
Concurrency is a first-class argument, not something you bolt on. The README's local example uses --n-concurrent 4 and the Daytona example uses --n-concurrent 100. The provider list in pyproject.toml is long, and each provider is an optional extra with its own dependencies, which tells you the sandbox layer is pluggable rather than hardcoded to one vendor.
Installing Harbor and running Terminal-Bench-2.0 locally
The README gives two installation paths. With uv, the command installs Harbor as a tool. The equivalent pip command is also documented. Python 3.12 or newer is required according to pyproject.toml.
uv tool install harborAfter installation, the package exposes three console scripts: harbor, hr and hb, all pointing at the same Typer application. The README uses harbor throughout.
The first real use is the official Terminal-Bench-2.0 harness. Export your Anthropic key, then run the benchmark locally. This launches Docker containers on your machine.
export ANTHROPIC_API_KEY=<YOUR-KEY>
harbor run --dataset [email protected] \
--agent claude-code \
--model anthropic/claude-opus-4-1 \
--n-concurrent 4You should see Harbor start four concurrent runs and report results as they finish. If you are unsure which agents are supported, the README points to harbor run --help, and harbor datasets list enumerates third party benchmarks such as SWE-Bench and Aider Polyglot. A shorter form exists for those datasets.
harbor run -d "<dataset@version>" -m "<model>" -a "<agent>"The same run moves to a cloud provider by adding an environment flag and the provider's key. The README's example uses Daytona with 100 concurrent runs.
export DAYTONA_API_KEY=<YOUR-KEY>
harbor run --dataset [email protected] \
--agent claude-code \
--model anthropic/claude-opus-4-1 \
--n-concurrent 100 \
--env daytonaWhere Harbor gets in the way
The provider extras are the first friction point. pyproject.toml defines daytona, modal, e2b, islo, runloop, tensorlake, gke, ec2, novita, cua, langsmith, adapter and huggingface as optional dependency groups. A plain pip install harbor does not pull any of them. The README's Daytona example shows the environment flag but does not show the extra you must install first, so a reader following only the README will hit an import or connection failure and have to find the extra name in pyproject.toml. That is a documentation gap, not a design flaw, but it costs time.
Local Docker is the default and it is also the ceiling. Running Terminal-Bench-2.0 at --n-concurrent 4 on a laptop is realistic; the README's own step up to 100 uses a hosted provider. If your evaluation depends on a specific container image, your local machine must be able to build and run it, and the harness does not abstract that away.
Version churn is the second cost. Three minor releases landed in roughly five weeks, and the project is pre-1.0. Dataset references are versioned ([email protected]), which helps, but nothing in the README documents how a pinned dataset version behaves after a harness upgrade. There is no rollback procedure in the README either. If you need a frozen evaluation environment for a paper or a compliance record, you will be managing that yourself.
Finally, the name is a genuine search problem. Harbor is a common English word, and the related searches around it are dominated by hardware stores, ferry terminals and a video game map. Finding the framework's own discussions takes deliberate querying.
How Harbor differs from writing your own eval loop
The obvious alternative is a homegrown script: a Python file that loops over tasks, shells out to an agent, and writes scores to JSON. That approach is smaller, has no dependencies beyond your agent's SDK, and you understand every line. For a fixed set of twenty tasks on one model, it is the better choice, and Harbor's abstractions would be overhead.
The difference shows up when any of the four axes moves. Harbor's dataset argument is resolved by name and version, so sharing a benchmark with another team means sharing a string rather than a directory of fixtures. Its agent argument is resolved through adapters, so swapping Claude Code for OpenHands is a flag change. Its environment argument is resolved through provider extras, so the same command that ran four local containers can run a hundred remote ones. A hand-written loop can do all of this, but you would be rebuilding the same four-way separation, and that is precisely the part Harbor already ships.
The other comparison worth naming is Terminal-Bench itself. Harbor is described as the official harness for Terminal-Bench-2.0, so if your work is Terminal-Bench, Harbor is not an alternative to it but the supported way to run it. If your work is a different benchmark, Harbor's dataset listing is what tells you whether it is already covered or whether you need to author tasks yourself.
Editorial conclusion
Adopt Harbor if you already run agent evaluations and want one CLI that covers local Docker and hosted sandbox providers, especially if you need Terminal-Bench-2.0 results. Do not adopt it if your evaluation needs are a single script over a fixed fixture set, because the dataset, agent and provider abstractions are overhead you will not use. Before committing, verify three things: which optional extra matches your provider, whether your agent needs an adapter package, and whether the version you pin still resolves the dataset you intend to run.
Frequently asked questions
How do I install the Harbor framework?
The README gives two options: uv tool install harbor, or pip install harbor. Python 3.12 or newer is required, and the installation exposes the harbor, hr and hb console scripts.
Which agents and models can Harbor evaluate?
The README names Claude Code, OpenHands and Codex CLI as examples of arbitrary agents, and the CLI takes an agent flag alongside a model flag. To see the full list of supported agents and other options, the README points to harbor run --help.
How do I run a benchmark on a cloud provider instead of locally?
Pass the --env flag with a provider name, as in the README's Daytona example, and export that provider's API key first. The provider is an optional dependency group in pyproject.toml, so the corresponding extra has to be installed.
What datasets does Harbor support besides Terminal-Bench?
The README states that running harbor datasets list shows all supported third party benchmarks, naming SWE-Bench and Aider Polyglot as examples. A dataset is referenced in the form dataset@version.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/harbor-framework-harbor)