Model or dataset
Ayanami0730/deep_research_bench avatar
Ayanami0730/deep_research_bench

DeepResearch Bench changed its judge twice, and the leaderboard is still split in two

DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents

838 stars86 forksPythonApache-2.0

At a glance

What is it?
DeepResearch Bench scores deep research agents with one frontier model as judge, a smaller one for citation checking, and a third for cleaning the articles first. The judge changed from one Google model to another in May 2026 for a stated reason, the page carries a migration deadline that has already passed, and the repository is still version 0.1 with four dependencies.
Who is it for?
Use this benchmark to compare agents, not to trust an absolute number, because the evaluator changed twice in a year and the two leaderboards are only comparable inside themselves. Before quoting a score, check which judge produced it and whether the entry has been migrated.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 6 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 4, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The judge changed for a dated reason, and the leaderboards were split to keep scores comparable

The news entries read like a changelog for a benchmark rather than for a library, and the largest one is the evaluator. On 11 May 2026 the official evaluator switched to GPT-5.5, and the trigger is stated: Google's announced deprecation of Gemini-2.5-Pro on 17 June 2026. A benchmark whose score is the output of a model cannot keep a judge the model vendor is retiring, so the same month the page published a migration plan. Results evaluated under the old judge and the new one are both accepted through 31 May 2026 and displayed on two separate leaderboards, so a ranking stays internally consistent instead of mixing judges. From 1 June the paper's own results were to be re-evaluated and migrated automatically, community entries had to be resubmitted, and the legacy code was preserved on a branch named after the old model. The page also said the new leaderboard was still under construction, expected within a week.

The candidate table picks a winner on three of five columns

The replacement judge was chosen by measurement rather than by reputation, and the numbers are published. Three frontier models were run on the human annotated subset, described as 50 tasks across 4 target agents, so 200 articles, and each candidate was scored for alignment with human judgements against a stated inter annotator agreement baseline of 68.78 percent. All three cleared it, by between 1.3 and 3 points. The table has five columns and the adoption is not a clean sweep: GPT-5.5 wins overall at 71.82, PAR at 73.00, and FAS at 59.23, while Gemini-3.1-Pro takes OPC at 90.14 against 89.70, and Claude-Opus-4-7 takes FAP at 66.70 against 65.35. So the official judge gave up a point on two sub metrics to win the composite and one other, which is a defensible choice and a specific one.

Three models do three jobs, and only the cheapest one moved in September

Reading the pipeline as three separate roles explains most of the news entries. One model cleans the articles, one is the official RACE judge, and a smaller one handles the FACT pipeline. The most recent change, dated 22 Sep 2026, moved only the cleaning step, from GPT-5.5 to GPT-5.6 Luna, and the entry says why in its own framing: it reduces the cost of the cleaning stage while preserving the official scoring setup. That distinction matters for anyone reproducing results, because a cleaning model swap is invisible in the score while a judge swap invalidates every earlier number. The earlier configuration, from the 15 July 2025 update, was the same shape with different models: Gemini-2.5-Pro for RACE and Gemini-2.5-Flash for FACT, which the page marks as superseded. The cleaning logic itself was also reworked in the May 2026 pipeline release, moving to a chunk based strategy for very long articles.

One hundred tasks in twenty two fields, served by one root script

The benchmark is described as filling a gap: a comprehensive benchmark for systematically evaluating deep research agents, built from 100 PhD level research tasks crafted by domain experts across 22 distinct fields. The fields listed include science and technology, covering physics, chemistry, biology, environmental science, and engineering, and finance and business, covering investments, personal finance, marketing, and human resources, with software among the categories the page begins to list before it ends. The repository is shaped like a research script rather than a package. At the root there is a single evaluation script named for the RACE pipeline, a shell runner, and directories for data, prompts, results, utilities, and images. Nothing at the root is named for the FACT pipeline, so a reader looking for the citation scoring code is left with the docs to go on.

The package says version 0.1 and the requirements are four lines

The packaging is the least maintained corner of an actively dated changelog. setup.py declares a package name, a version of 0.1, and a call to find packages, with nothing else: no metadata, no dependencies, no extras, no Python version floor. The dependency file is four lines, a progress bar, a dataframe library, an array library, and an HTTP client. What that list does not contain is any model client library, which means every call to a frontier model in this pipeline goes out through a hand written request rather than a vendor SDK. For a benchmark whose entire cost is model inference, that is a deliberate trade: the numbers stay reproducible and the bill stays visible, but the caller owns retries, rate limits, and token accounting. The same applies to the results, which the page says are published as raw articles and scores on a hosted leaderboard rather than generated by the repository.

The homepage is a paper PDF and the leaderboards live on two accounts

The project link fields are an accurate map of where the artefacts actually are. The recorded homepage is a link to a PDF of the paper on arXiv, not a documentation site, though a project site also exists on GitHub pages. The dataset sits on a model hosting platform under one account, and the leaderboards are split across two spaces on that platform under two different accounts, which is consistent with the dual evaluator arrangement rather than an accident. The same evaluation interface is mirrored on a third party evaluation platform, an arrangement the page dates to a partnership announced on 18 July 2025. Submissions are handled by email, and both addresses given sit on the same Chinese university domain, so the people who run the benchmark and the people who answer submissions are the same lab.

The successor benchmark lives in another repository, and this one says it continues

A benchmark that outlives its own judge usually outlives its own dataset too, and this project says so directly. On 6 Feb 2026 the authors released DeepResearch Bench II in a separate repository, with its own homepage, its own paper, and a different evaluation focus: 9,430 fine grained binary rubrics covering information recall, analysis, and presentation, derived from expert written articles. The announcement is careful to add that DRB, as the predecessor, will continue to be maintained and updated. The same entry lists the rest of the lab's output, two more benchmarks that use Wikipedia's good articles as expert references and that benchmark graph retrieval on long heterogeneous documents with 1,100 questions, plus two agent frameworks. So the naming is confusing on purpose: a reader searching for the current version of deep research bench has three candidates and no way to tell from the names which one is which.

Editorial conclusion

Use this benchmark to compare agents, not to trust an absolute number, because the evaluator changed twice in a year and the two leaderboards are only comparable inside themselves. Before quoting a score, check which judge produced it and whether the entry has been migrated. And if you plan to run the pipeline yourself, budget for model calls by hand: the repository ships four dependencies and no client library, so the request layer is yours to write.

Frequently asked questions

What is DeepResearch Bench?

It is a benchmark for evaluating deep research agents, built from 100 PhD level research tasks crafted by domain experts across 22 fields. Scoring runs through a RACE pipeline for overall quality and a separate FACT pipeline for citations.

Which model does DeepResearch Bench use as its judge?

GPT-5.5 is the official RACE judge and GPT-5.4-mini runs the FACT pipeline. Since 22 Sep 2026 the article cleaning step uses GPT-5.6 Luna. The judge itself moved to GPT-5.5 on 11 May 2026, replacing Gemini-2.5-Pro ahead of that model's deprecation.

How were the DeepResearch Bench judge candidates compared?

Three frontier models were scored on the human annotated subset of 50 tasks and 4 target agents, 200 articles, against a human inter annotator agreement baseline of 68.78 percent. All three cleared it by 1.3 to 3 points, and GPT-5.5 was adopted after winning on overall, PAR, and FAS.

Is there a second version of DeepResearch Bench?

Yes. DeepResearch Bench II was released on 6 Feb 2026 in a different repository, with 9,430 fine grained binary rubrics and a different evaluation focus. The page states that the original benchmark continues to be maintained and updated.

What is inside the deep_research_bench repository?

A single root evaluation script named for the RACE pipeline, a shell runner, and directories for data, prompts, results, utilities, and images. setup.py declares version 0.1, and the dependency file lists four packages with no model client library.

Official sources

  1. Ayanami0730/deep_research_bench on GitHub
  2. Issues
  3. License: Apache-2.0
  4. Project website
  5. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/ayanami0730-deep-research-bench.svg)](https://hysenlabs.com/projects/ayanami0730-deep-research-bench)