Terminal-Bench-Science: a 70-task benchmark for AI agents on real research workflows
Terminal-Bench-Science: Evaluating AI agents on research workflows across scientific domains
At a glance
- What is it?
- Terminal-Bench-Science collects expert-curated scientific tasks and runs them through Harbor, scoring agents on whether they can finish work a domain researcher would recognise. The design is sound; the cost of running it is the part worth checking before you commit.
- Who is it for?
- Adopt Terminal-Bench-Science if you are evaluating a terminal-capable agent against scientific work and you have a sandbox provider plus the budget for repeated trials, because the oracle-first workflow in the README is designed to be run before any agent is scored. Do not adopt it if your interest is software engineering tickets, since the task corpus in this release is scientific workflows only, or if you cannot supply the sandboxing environment the benchmark expects.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 1 day ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What problem Terminal-Bench-Science is built to solve
Most agent evaluations measure work that software engineers already do: reading a repository, patching a function, closing a ticket. Terminal-Bench-Science starts from a different premise. Its tasks are drawn from scientific research workflows across the life, physical, earth, mathematical, and engineering sciences, and the README describes them as authored and reviewed by domain experts. The intended audience is therefore narrow and specific: researchers who want to know whether an agent can support the work they already do, and agent developers who need a signal that is not another coding benchmark.
The README frames the project as a community effort rather than a fixed test set. It calls Terminal-Bench-Science a continuous benchmark that evolves alongside frontier AI, with a feedback loop between scientific needs and AI development. That is a meaningful design commitment. A static benchmark decays as models improve; a continuously contributed one has to solve governance instead, which is why the repository ships a proposal form, a contribution guide, and a review process rather than just a task directory.
How the task corpus and review pipeline fit together
The repository layout shows the mechanism more clearly than the README prose does. Tasks live under tasks/, with supporting material in rubrics/, scripts/, tools/, ci_checks/, and lite/. The contribution flow is stated as Propose, Build, Review, and the README gives the three steps plainly: propose the idea through a task proposal form to get feedback and approval, build the task and open a pull request following CONTRIBUTING.md, then go through automated checks, parallel domain and technical review, and final approval before merge.
The commented-out material in the README is where the automation detail sits. It lists static checks covering path validation, Dockerfile sanity, canary, metadata and test references; an implementation rubric review run through a command called harbor check against a 39-criteria rubric stored in rubrics/task-implementation.toml; TF-IDF duplicate detection for similarity; a Docker build step that validates both an oracle solution and a no-op solution, where the solution must pass and the no-op must fail; multi-agent trials; and adversarial cheat trials aimed at reward-hack detection. That last pair is the part worth noting. Requiring a no-op to fail is what separates a task that measures capability from a task an agent can pass by doing nothing plausible.
The README also states that the benchmark currently contains 70 expert-curated tasks and is growing toward 100+. Coverage can be tracked on a separate task dashboard, and releases are tagged on Harbor Hub.
Installing Harbor and running the oracle solutions
Terminal-Bench-Science does not ship its own runner. The README directs you to install Harbor and then run the oracle solutions five times to confirm all tasks work as expected in your sandboxing environment. Modal or Daytona are the recommended environments. The install line uses uv with both extras:
uv tool install "harbor[modal,daytona]"
# or pip install "harbor[modal,daytona]"The README gives the pip alternative on the following line, so either package manager is acceptable. Once Harbor is installed, the oracle run is the first command to execute:
harbor run -d terminal-bench-science/terminal-bench-science@latest \
-k 5 \
--agent oracle \
--n-concurrent 32 \
--env modalThe -k 5 flag repeats each task five times, which is why the README recommends this before anything else: it is a check on your own sandbox rather than on any model. If the oracle flakes on your setup, the README asks you to open an issue on the repository rather than assume the task is broken. To score an actual agent, you swap the oracle for an agent and model:
harbor run -d terminal-bench-science/terminal-bench-science@latest \
--agent claude-code \
--model anthropic/claude-opus-5 \
--ak reasoning_effort=max \
--n-concurrent 32 \
--env modalThe --ak flag passes agent-specific keyword arguments; reasoning_effort=max is the example the README uses. The dataset reference terminal-bench-science/terminal-bench-science@latest is what pulls the current task set, and @latest is deliberate for a benchmark that keeps moving.
Where the benchmark gets expensive and awkward
The --n-concurrent 32 setting in the README's own examples is the honest signal about cost. Each task runs in its own containerised environment, and the oracle pass alone means five repetitions of every task across the whole corpus. The README does not state a price, a runtime, or a per-task cost estimate, and it does not document what happens when a run fails partway through or how to resume it. Anyone budgeting for a full evaluation is estimating from the concurrency number and their provider's rates.
There is a second constraint that matters more than money. The benchmark depends on external sandboxing infrastructure, Modal or Daytona in the documented path, and the README treats environment flakiness as a real possibility by asking you to validate the oracle first. An agent that fails a task on a machine with an unstable container runtime produces a number that looks like a capability result and is not one. The README's instruction to open an issue when the oracle flakes is the right response, but it also means your evaluation is only as reproducible as the provider underneath it.
The third limitation is scope. If the workflow you care about is software engineering, this is the wrong tool, because the task corpus is drawn from scientific domains. And the README is explicit that tasks must produce outcomes objectively verifiable in a terminal environment, which rules out research work whose success is a judgement call rather than a checkable artefact.
How it differs from SWE-bench-style evaluation
The natural comparison is SWE-bench and its descendants. Those benchmarks draw tasks from real software repositories, typically as issues paired with the pull request that resolved them, and they score an agent on producing a patch that passes the project's tests. The unit of work is a code change, and the ground truth is a test suite that already exists in the upstream project.
Terminal-Bench-Science inverts both assumptions. Tasks are authored rather than harvested, by domain experts rather than extracted from repository history, and the ground truth is constructed by the task contributor as part of the submission. That is more work per task and it is why the project needs a proposal form, a 39-criteria implementation rubric, duplicate detection, and adversarial cheat trials. It also buys something SWE-bench cannot offer: coverage of domains where the correct output is a data analysis, a simulation result, or a computed quantity rather than a merged patch. The trade-off is that task quality rests on the review pipeline in a way a harvested benchmark does not have to worry about, and the corpus grows at the speed of expert contribution rather than at the speed of repository activity.
Licence, releases and the cost of keeping up
The repository is Apache-2.0, which permits commercial and academic use with the usual attribution and notice obligations; the repository carries a CITATION.cff for academic citation and a DOI, so published work built on it has a defined reference. Nothing in the README suggests a separate licence for the task content, but the README does not discuss licensing at all, so anyone redistributing the tasks should read the LICENSE file rather than infer terms from the repository badge.
Upgrade cost is shaped by the @latest reference and the release cadence. The most recent release noted is v0.1.0, from 2026-08-26, and the last push to the repository was on 2026-09-09. There is a RELEASING.md and a release-notes-v0.1.0.md in the repository root, which suggests releases are a documented process rather than an afterthought. The practical consequence is that a score recorded today is tied to a task set that is explicitly designed to change; the README describes the benchmark as continuous and growing toward 100+ tasks. Reproducing a published number later means pinning a specific dataset reference instead of @latest, and the README's own examples do not show a pinned form, so that is something to work out from Harbor's documentation.
Editorial conclusion
Adopt Terminal-Bench-Science if you are evaluating a terminal-capable agent against scientific work and you have a sandbox provider plus the budget for repeated trials, because the oracle-first workflow in the README is designed to be run before any agent is scored. Do not adopt it if your interest is software engineering tickets, since the task corpus in this release is scientific workflows only, or if you cannot supply the sandboxing environment the benchmark expects. Verify two things first: that the oracle passes repeatedly on your own infrastructure rather than assuming it will, and that the task set covers the domains you actually care about, since the README states 70 tasks today against a stated goal of 100+.
Frequently asked questions
What does Terminal-Bench-Science actually measure?
It measures AI agent capability on expert-curated scientific research workflows across the life, physical, earth, mathematical, and engineering sciences. Tasks are contributed by domain experts and must produce outcomes that can be objectively verified in a terminal environment.
How much does it cost to run Terminal-Bench-Science?
The README does not give a price or a runtime estimate. It recommends Modal or Daytona as sandbox environments and uses --n-concurrent 32 in its example commands, and it asks you to run the oracle solutions five times before scoring any agent, so the cost scales with the size of the task set and your provider's rates.
What is the difference between Terminal-Bench-Science and SWE-bench?
SWE-bench-style benchmarks score an agent on producing a patch that passes an existing project's tests. Terminal-Bench-Science uses tasks authored and reviewed by scientific domain experts rather than harvested from repository history, and validates them through a review pipeline that includes an oracle pass and a required no-op failure.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/harbor-framework-terminal-bench-science)