# Harbor: a Python harness for evaluating agents in sandboxed environments

> Harbor is the official harness for Terminal-Bench-2.0 and a general runner for agent and model evaluations. It installs as a CLI, runs agents such as Claude Code against datasets, and can spread jobs across cloud sandbox providers.

**harbor-framework/harbor** — Framework for evaluating and improving agents 

- Repository: https://github.com/harbor-framework/harbor
- Website: https://harborframework.com/
- Stars: 5,729 · Forks: 1,892
- Language: Python
- License: Apache-2.0
- Published: 2026-09-22 · Updated: 2026-09-22 · Language: en
- Canonical page: https://hysenlabs.com/projects/harbor-framework-harbor

## What Harbor is for, and who ends up using it

Harbor is a framework for evaluating and optimizing agents and language models, published by the creators of Terminal-Bench. The README lists four things you can do with it: evaluate arbitrary agents such as Claude Code, OpenHands and Codex CLI; build and share your own benchmarks and environments; run experiments in many environments in parallel through providers like Daytona, Modal, LangSmith, Blaxel, Novita Sandbox, Tensorlake and Runta; and generate rollouts for RL optimization.

The audience is narrower than that list suggests. If you are comparing two models on a chat task, Harbor is the wrong shape: it is built around agents that act inside a container, receive a task, and produce a result the harness can score. The people who get value are evaluation engineers maintaining a benchmark suite, and research teams that need many parallel rollouts rather than a handful of scored prompts. The RL rollout use case points the same way: you want volume and reproducibility, not a leaderboard screenshot.

The project is not archived, and its last push was on 2026-09-22. Releases are frequent: v0.21.0 on 2026-08-10, v0.22.0 on 2026-08-22, v0.23.0 on 2026-09-12. That cadence is a real cost as well as a signal, because a CLI whose flags and dataset references move this quickly will occasionally break the command you saved three weeks ago.

## The mechanism: datasets, agents, models and sandbox providers

Harbor separates four things that are usually tangled together in evaluation scripts: the dataset, the agent, the model, and the environment the run happens in. The CLI takes all four as arguments. A dataset is named with an at-sign version, for example terminal-bench@2.0. An agent is named separately, for example claude-code. A model is named with a provider prefix, for example anthropic/claude-opus-4-1. The environment defaults to local Docker and can be switched with a flag.

That separation is the whole design. The same dataset can be run against a different agent without editing task definitions, and the same agent and dataset can move from a laptop to a hosted provider by changing one argument. The repository layout supports the reading: adapters/ holds agent integrations, packages/ holds workspace members such as harbor-rewardkit and harbor-langsmith, and examples/ is split into agents, configs, exec, jobs, metrics, prompts and tasks. Those directories suggest the intended extension points are agent adapters and task definitions rather than a single monolithic runner.

Concurrency is a first-class argument, not something you bolt on. The README's local example uses --n-concurrent 4 and the Daytona example uses --n-concurrent 100. The provider list in pyproject.toml is long, and each provider is an optional extra with its own dependencies, which tells you the sandbox layer is pluggable rather than hardcoded to one vendor.

## Installing Harbor and running Terminal-Bench-2.0 locally

The README gives two installation paths. With uv, the command installs Harbor as a tool. The equivalent pip command is also documented. Python 3.12 or newer is required according to pyproject.toml.

```bash
uv tool install harbor
```

After installation, the package exposes three console scripts: harbor, hr and hb, all pointing at the same Typer application. The README uses harbor throughout.

The first real use is the official Terminal-Bench-2.0 harness. Export your Anthropic key, then run the benchmark locally. This launches Docker containers on your machine.

```bash
export ANTHROPIC_API_KEY=<YOUR-KEY>
harbor run --dataset terminal-bench@2.0 \
   --agent claude-code \
   --model anthropic/claude-opus-4-1 \
   --n-concurrent 4
```

You should see Harbor start four concurrent runs and report results as they finish. If you are unsure which agents are supported, the README points to harbor run --help, and harbor datasets list enumerates third party benchmarks such as SWE-Bench and Aider Polyglot. A shorter form exists for those datasets.

```bash
harbor run -d "<dataset@version>" -m "<model>" -a "<agent>"
```

The same run moves to a cloud provider by adding an environment flag and the provider's key. The README's example uses Daytona with 100 concurrent runs.

```bash
export DAYTONA_API_KEY=<YOUR-KEY>
harbor run --dataset terminal-bench@2.0 \
   --agent claude-code \
   --model anthropic/claude-opus-4-1 \
   --n-concurrent 100 \
   --env daytona
```

## Where Harbor gets in the way

The provider extras are the first friction point. pyproject.toml defines daytona, modal, e2b, islo, runloop, tensorlake, gke, ec2, novita, cua, langsmith, adapter and huggingface as optional dependency groups. A plain pip install harbor does not pull any of them. The README's Daytona example shows the environment flag but does not show the extra you must install first, so a reader following only the README will hit an import or connection failure and have to find the extra name in pyproject.toml. That is a documentation gap, not a design flaw, but it costs time.

Local Docker is the default and it is also the ceiling. Running Terminal-Bench-2.0 at --n-concurrent 4 on a laptop is realistic; the README's own step up to 100 uses a hosted provider. If your evaluation depends on a specific container image, your local machine must be able to build and run it, and the harness does not abstract that away.

Version churn is the second cost. Three minor releases landed in roughly five weeks, and the project is pre-1.0. Dataset references are versioned (terminal-bench@2.0), which helps, but nothing in the README documents how a pinned dataset version behaves after a harness upgrade. There is no rollback procedure in the README either. If you need a frozen evaluation environment for a paper or a compliance record, you will be managing that yourself.

Finally, the name is a genuine search problem. Harbor is a common English word, and the related searches around it are dominated by hardware stores, ferry terminals and a video game map. Finding the framework's own discussions takes deliberate querying.

## How Harbor differs from writing your own eval loop

The obvious alternative is a homegrown script: a Python file that loops over tasks, shells out to an agent, and writes scores to JSON. That approach is smaller, has no dependencies beyond your agent's SDK, and you understand every line. For a fixed set of twenty tasks on one model, it is the better choice, and Harbor's abstractions would be overhead.

The difference shows up when any of the four axes moves. Harbor's dataset argument is resolved by name and version, so sharing a benchmark with another team means sharing a string rather than a directory of fixtures. Its agent argument is resolved through adapters, so swapping Claude Code for OpenHands is a flag change. Its environment argument is resolved through provider extras, so the same command that ran four local containers can run a hundred remote ones. A hand-written loop can do all of this, but you would be rebuilding the same four-way separation, and that is precisely the part Harbor already ships.

The other comparison worth naming is Terminal-Bench itself. Harbor is described as the official harness for Terminal-Bench-2.0, so if your work is Terminal-Bench, Harbor is not an alternative to it but the supported way to run it. If your work is a different benchmark, Harbor's dataset listing is what tells you whether it is already covered or whether you need to author tasks yourself.

## Conclusion

Adopt Harbor if you already run agent evaluations and want one CLI that covers local Docker and hosted sandbox providers, especially if you need Terminal-Bench-2.0 results. Do not adopt it if your evaluation needs are a single script over a fixed fixture set, because the dataset, agent and provider abstractions are overhead you will not use. Before committing, verify three things: which optional extra matches your provider, whether your agent needs an adapter package, and whether the version you pin still resolves the dataset you intend to run.

## FAQ

### How do I install the Harbor framework?

The README gives two options: uv tool install harbor, or pip install harbor. Python 3.12 or newer is required, and the installation exposes the harbor, hr and hb console scripts.

### Which agents and models can Harbor evaluate?

The README names Claude Code, OpenHands and Codex CLI as examples of arbitrary agents, and the CLI takes an agent flag alongside a model flag. To see the full list of supported agents and other options, the README points to harbor run --help.

### How do I run a benchmark on a cloud provider instead of locally?

Pass the --env flag with a provider name, as in the README's Daytona example, and export that provider's API key first. The provider is an optional dependency group in pyproject.toml, so the corresponding extra has to be installed.

### What datasets does Harbor support besides Terminal-Bench?

The README states that running harbor datasets list shows all supported third party benchmarks, naming SWE-Bench and Aider Polyglot as examples. A dataset is referenced in the form dataset@version.

## Sources

- [harbor-framework/harbor on GitHub](https://github.com/harbor-framework/harbor)
- [License: Apache-2.0](https://github.com/harbor-framework/harbor/blob/main/LICENSE)
- [Project website](https://harborframework.com/)
- [README](https://github.com/harbor-framework/harbor/blob/main/README.md)
- [Releases](https://github.com/harbor-framework/harbor/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/harbor-framework-harbor
