Model or dataset
harbor-framework/terminal-bench-1 avatar
harbor-framework/terminal-bench-1

Terminal-Bench: Evaluating AI Agents on Real-World Terminal Tasks

A benchmark for LLMs on complicated tasks in the terminal

2,597 stars571 forksPythonApache-2.0

At a glance

What is it?
Terminal-Bench is a benchmark and execution harness for measuring how well AI agents handle complex end-to-end tasks in real terminal environments: compiling code, training models, setting up servers, and similar work that requires reasoning about system state over multiple steps. The current beta has approximately 100 tasks and a leaderboard built around a versioned task dataset called Terminal-Bench-Core.
Who is it for?
Terminal-Bench is worth using for teams building or evaluating AI agents that need to operate in terminal environments rather than code editors or chat interfaces. The pip-installable `tb` CLI makes setup straightforward, but Docker and `uv` are required dependencies.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 81 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 28, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What Terminal-Bench Measures and Who Needs It

Most AI agent benchmarks evaluate agents on tasks that resemble software engineering work: fixing bugs in a codebase, answering questions about code, or making pull requests. Terminal-Bench targets a different category: tasks that require a sequence of commands in a real terminal to reach a verifiable end state. Compiling a specific binary, training a model to a target accuracy, configuring a server, and setting up a development environment are the kinds of tasks the README describes.

The primary users are teams building LLM-based agents that interact with real systems, benchmarking frameworks that need a task suite with well-defined verification, and researchers stress-testing system-level reasoning. The README describes Terminal-Bench as a benchmark for both building and evaluating agents, and it is structured to serve both purposes: the task dataset provides evaluation targets, and the execution harness connects any agent to those tasks.

Task Structure: Instruction, Test Script, and Reference Solution

Each task in Terminal-Bench has three components. The instruction describes what the agent should do, in English. The test script verifies whether the agent completed the task correctly; it runs after the agent's actions and returns pass or fail. The reference solution, called the oracle, demonstrates one correct way to complete the task and is used to validate that each task is solvable.

Tasks are stored in the `tasks/` directory of the repository. The README points to a Task Gallery at tbench.ai/tasks for browsing current tasks, and a Task Ideas page at tbench.ai/docs/task-ideas for community-sourced suggestions. The verification approach means task pass rates are objective and reproducible: the same test script runs in the same Docker-based sandbox regardless of which agent is being evaluated. The README notes that each task includes a test script to verify correct completion, which distinguishes Terminal-Bench from benchmarks that rely on model-graded evaluation.

Installing and Running the Terminal-Bench Harness

Terminal-Bench is distributed as a pip package with a `tb` CLI:

bash
uv tool install terminal-bench

or

bash
pip install terminal-bench

Once installed, view the harness options with:

bash
tb run --help

Running an evaluation requires Docker and `uv` in addition to the Python package. The harness connects a language model to a sandboxed terminal environment managed by Docker. The README gives this example for submitting to the leaderboard:

bash
tb run \
    --agent terminus \
    --model anthropic/claude-3-7-latest \
    --dataset-name terminal-bench-core \
    --dataset-version 0.1.1 \
    --n-concurrent 8

The `--n-concurrent 8` flag runs eight tasks simultaneously, which affects total evaluation time significantly.

The Leaderboard and Terminal-Bench-Core Dataset

The leaderboard is built around a versioned dataset called Terminal-Bench-Core. The current beta leaderboard uses Terminal-Bench-Core v0.1.1. To submit results to the leaderboard, the README instructs passing `--dataset-name terminal-bench-core` and `--dataset-version 0.1.1` to the harness run command. A separate leaderboard submission guide at tbench.ai/docs/submitting-to-leaderboard covers the complete submission process.

The versioning of the dataset matters because tasks can change between versions. Comparing results across submissions requires that both runs used the same dataset version. The README describes a registry system at tbench.ai/docs/registry for understanding how datasets and versions are managed. The `registry.json` file in the repository root is likely the machine-readable source of that registry. The leaderboard covers agents on Terminal-Bench-Core v0.1.1 as the current reference point, but the dataset is expected to expand as the project moves out of beta.

Terminal-Bench vs SWEBench: Different Task Categories

SWEBench evaluates agents on real GitHub issues, asking the agent to write code that passes the issue's test suite. The tasks are software engineering tasks drawn from open-source repositories, with success defined by test passage. Terminal-Bench evaluates agents on terminal operations that go beyond code editing: compiling binaries, running training jobs, configuring system services, and setting up environments from scratch.

The difference is in the action space. SWEBench agents primarily read and write files in a codebase. Terminal-Bench agents issue shell commands, manage processes, interpret error output, and adjust configuration in response to runtime behavior. An agent that scores well on SWEBench may not generalize to Terminal-Bench tasks if it lacks the ability to reason about process state, file permissions, package installation, and system configuration. Both benchmarks are complementary rather than redundant.

The README describes Terminal-Bench tasks as covering real-world, end-to-end work that must be completed autonomously. The word autonomously is significant: the harness does not prompt the agent with hints or partial solutions between steps. The agent must determine the correct sequence of commands from the initial instruction alone.

Maintenance, the Move to Harbor, and License

The last push to this repository was on 2026-07-11. The README includes an announcement at the top: new users should check out harbor, described as a new framework that can be used to run Terminal-Bench 2.0. The harbor framework is at github.com/laude-institute/harbor. This repository covers Terminal-Bench 1 (the current beta), while the next generation is being built separately.

The current version in pyproject.toml is 0.2.18. The project has no GitHub releases; versioning is tracked through pyproject.toml changes. The Apache 2.0 license permits unrestricted use, modification, and redistribution. The adapters directory, which the README describes as the mechanism for connecting different agents and benchmarks, and the tasks directory are both excluded from the ruff linter configuration, suggesting they are treated as external content rather than library code.

The repository includes a `dashboard/` directory and documentation at tbench.ai/docs/dashboard, indicating a web interface is available for browsing task results. A `discord-bot/` directory suggests there is an automated Discord integration for community interaction with the benchmark results.

Editorial conclusion

Terminal-Bench is worth using for teams building or evaluating AI agents that need to operate in terminal environments rather than code editors or chat interfaces. The pip-installable `tb` CLI makes setup straightforward, but Docker and `uv` are required dependencies. Verify whether your agent has an existing adapter or whether you need to write one before committing to the evaluation. The last push to the repository was on 2026-07-11, and the README announces Terminal-Bench 2.0 via a new framework called harbor. The Apache 2.0 license permits unrestricted use.

Frequently asked questions

What does terminal-bench mean?

Terminal-Bench is a benchmark for testing AI agents on tasks completed in a real terminal environment, such as compiling code, training models, and configuring servers. The name reflects the terminal as the environment where agents are evaluated, not a GUI or code editor.

What is terminal-bench vs SWEBench?

The README does not compare Terminal-Bench directly to SWEBench. Terminal-Bench focuses on open-ended terminal tasks including system administration, compilation, and server configuration in a Docker-based sandbox. The evaluation is based on a test script that verifies whether the agent reached the correct end state.

What is terminal-bench hard?

The README does not describe a dataset specifically named terminal-bench hard. The current leaderboard dataset is Terminal-Bench-Core v0.1.1. The README describes the benchmark as covering complicated tasks and notes that the project will expand into a comprehensive testbed for AI agents in text-based environments.

How many tasks are in terminal-bench?

The README states Terminal-Bench is currently in beta with approximately 100 tasks. The project intends to expand the task count over the coming months.

Official sources

  1. harbor-framework/terminal-bench-1 on GitHub
  2. Issues
  3. License: Apache-2.0
  4. Project website
  5. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/harbor-framework-terminal-bench-1.svg)](https://hysenlabs.com/projects/harbor-framework-terminal-bench-1)