OpenResearcher: A 30B-A3B Model and an Open Pipeline for Long-Horizon Deep Research Trajectories
OpenResearcher: A Fully Open Pipeline for Long-Horizon Deep Research Trajectory Synthesis
At a glance
- What is it?
- OpenResearcher is a fully open recipe for deep research: a 96K trajectory dataset, a 30B-A3B model trained on it, and a lightweight evaluation framework that runs against a self-built retriever over an 11B-token corpus. The pipeline is the interesting part, and the hardware cost of the local search stack is the catch.
- Who is it for?
- Adopt OpenResearcher if you need an open, reproducible deep research trajectory pipeline and can pay the storage and indexing cost of the local retriever, or if you only want the 30B-A3B model and the 96K dataset for your own training. Do not adopt it as a hosted research assistant: there is no service, the demo is a Hugging Face Space, and the GAIA path still needs a Serper API key.
- Can I use it commercially?
- Not without permission. GitHub finds no licence file in the repository, and without a licence all rights are reserved by default: you may read the code but not reuse it. Check the README, or ask the authors, before using it.
- Is it still maintained?
- Yes. The repository last received commits 112 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The gap OpenResearcher fills: nobody publishes the trajectory, only the score
Deep research agents are usually reported as a number on a benchmark. The training data, the retriever, the tool-call format and the filtering rules stay inside the lab. OpenResearcher's stated goal is to publish that middle layer: the README describes a "fully open-source recipe" covering a 96K DeepResearch trajectory dataset with 100+ turns generated by GPT-OSS-120B with native browser tools, the 30B-A3B model trained on it, a distillation recipe, and a lightweight evaluation framework. The intended audience is not an end user looking for answers. It is a team that wants to train or evaluate a long-horizon research agent and needs trajectories that show what a hundred turns of search, read and synthesis actually look like, plus a harness that scores them.
The project also positions the model as a release in its own right. The README reports 54.8% accuracy on BrowseComp-Plus and claims it surpasses GPT-4.1, Claude-Opus-4, Gemini-2.5-Pro, DeepSeek-R1 and Tongyi-DeepResearch, with the benchmark table rendered as an image rather than text. Those are the project's own numbers, evaluated on the project's own harness, and the README does not describe how the baselines were run. Treat the comparison as a claim to reproduce, not a settled result.
How the pipeline works: a self-built retriever replaces the search API
The design decision that shapes everything else is the retriever. Instead of calling a commercial search API during trajectory generation, OpenResearcher builds its own retriever over a dedicated ~11B-token corpus, described in the README as eliminating the need for external Search APIs and significantly reducing training costs. Trajectories were generated at scale against that local index using GPT-OSS-120B with native browser tools, which is why the dataset can be 96K trajectories with 100+ turns each: no per-query billing, no rate limits, no vendor downtime in the middle of a rollout.
The repository layout reflects a two-part system. The evaluation side is Python: eval.py runs a benchmark, browser.py and backend.py expose the browsing and retrieval layer, deploy_agent.py stands up the agent, and run_agent.sh drives a run. The training side is packaged separately, with utils/ and a tevatron package declared in pyproject.toml, which is the usual home for retrieval training code. The dependency list tells you what the retriever is built from: pyserini, faiss-cpu and duckdb for indexing and querying, datasets for corpus handling, vllm==0.13.0 for generation, and fastapi plus httpx for the serving and tool-call layer.
The trade-off is explicit in that list. You are not calling someone else's search endpoint; you are hosting a corpus and its indexes. That is what makes the recipe reproducible and what makes it expensive to stand up.
Installing OpenResearcher and running the BrowseComp-Plus path locally
The README points to a setup.sh at the repository root and a .env.template for configuration. Python 3.12 or newer is required by pyproject.toml, and the pinned vllm==0.13.0 dependency means the environment is not flexible: install into a dedicated virtual environment rather than into an existing one.
git clone https://github.com/TIGER-AI-Lab/OpenResearcher
cd OpenResearcher
cp .env.template .env
bash setup.shThe .env file is where the API keys and endpoint settings go; the README's Configuration section covers it but the cleaned README does not enumerate the keys, so read .env.template before editing. After setup, the documented first real use is the BrowseComp-Plus benchmark with the local search engine, which needs the corpus and indexes in place. The README's Quick Start section shows the agent being launched through run_agent.sh, and eval.py is the entry point for scoring the resulting trajectories.
If you do not want to host the corpus, the README gives a second path: GAIA with the Serper API, described as needing no local search. That swaps the self-built retriever for an external search provider, which is the right way to try the agent without committing to the storage and indexing work. The README also links a Gradio demo on Hugging Face Spaces for readers who only want to see the model behave.
The local retriever is the real adoption cost
The feature list sells the self-built retriever as a cost reduction, and for training runs at scale it is. For a single team trying the project out, it is the largest line item. An ~11B-token corpus has to be downloaded, stored and indexed with pyserini and faiss-cpu, and the dependencies are CPU builds, which means the index lives in RAM or on disk rather than on a GPU. The README does not state the corpus size in gigabytes, the index build time, or the memory needed to serve queries, so you cannot size the machine from the documentation alone. Check the corpus dataset page before you start.
Generation is a separate cost. The model is 30B-A3B, and the evaluation harness depends on vllm==0.13.0, so you need a GPU host that can serve it. The README does not publish a minimum VRAM figure, and the pinned vllm version means you cannot simply upgrade to whatever your cluster already has installed.
The GAIA path exists precisely because of this. It substitutes Serper for the local retriever and removes the corpus and index from the equation, at the cost of an external API key and per-query spend. If your goal is to evaluate the model rather than to reproduce the data-generation pipeline, that is the path to start with, and the README presents it that way.
Where OpenResearcher is the wrong tool
It is not a research assistant you point at a question. There is no hosted service, no CLI that answers a query and returns a cited report, and no documented deployment path for production traffic. The Gradio Space is a demo, not an API contract. A team that wants a deep research product should look at an agent framework with a stable interface and bring their own model.
It is also the wrong choice if you cannot host a 30B-A3B model. The evaluation framework is lightweight, but the thing being evaluated is not, and the pinned vllm==0.13.0 dependency narrows the environments where it will run at all. Teams on managed inference endpoints will find that the harness assumes local serving.
A third boundary is licensing. The repository does not state a licence, and the README does not discuss terms for the code, the 30B-A3B model weights, the 96K trajectory dataset or the ~11B-token corpus. Those are four separate artifacts hosted in different places, and each may carry different terms. That is a question for whoever handles compliance on your side, not something the README answers, and it is worth resolving before the pipeline becomes load-bearing.
How it differs from calling a search API inside an agent loop
The obvious alternative is the common pattern: an agent framework that calls a hosted search API or a commercial deep research endpoint per step. The difference is not quality, it is where the retrieval happens and what you can inspect afterward. A hosted search API gives you fresh web results and no infrastructure, but the trajectories you collect depend on a vendor's ranking, and you cannot replay them later against the same index. OpenResearcher's self-built retriever makes the retrieval step a local artifact: the corpus is fixed, the index is yours, and a trajectory generated today can be re-run against the same index next year. For data synthesis, that reproducibility is the point.
The cost of that choice is freshness. A fixed ~11B-token corpus does not contain this week's pages, which is fine for generating training trajectories and fine for benchmarks with frozen evidence, but wrong for questions whose answers change. The README's GAIA example acknowledges this by offering the Serper path, which is the pragmatic middle: keep the agent and the harness, borrow someone else's index when you need the live web.
A second difference is scale of effort. A search-API agent is an afternoon. OpenResearcher is a corpus download, an index build, a GPU host and a pinned inference stack. Choose it when you intend to generate or study trajectories, not when you need one answer.
Maintenance, upgrade cost and what the repository does not say
The last push to the default branch was on 2026-06-10. The repository is not archived, and the README's news section shows a steady cadence through the first half of 2026: the training code in February, the paper in March, and the Nemotron 3 Ultra adoption note in June. There are no retrieved releases, so versioning is by commit rather than by tag, and the pyproject.toml version is 0.1.0. Upgrading means tracking main and re-resolving dependencies.
The upgrade surface is narrow but sharp. vllm is pinned to exactly 0.13.0, so a cluster-wide vllm upgrade will break the harness until the pin moves. transformers, peft and datasets are given lower bounds rather than pins, which means a fresh install can pull newer versions than the authors tested against. pyserini and faiss-cpu are the retriever's foundation; changing either implies rebuilding indexes. If you are running this in CI or on shared infrastructure, freeze the environment after a working install rather than reinstalling later.
On licensing, the repository states no licence, which is the single most important unresolved item for anyone planning to ship something built on it. The code, the model, the dataset and the corpus are separate artifacts, and the README does not describe their terms. That is a question for your own review, not a conclusion this article can reach.
Editorial conclusion
Adopt OpenResearcher if you need an open, reproducible deep research trajectory pipeline and can pay the storage and indexing cost of the local retriever, or if you only want the 30B-A3B model and the 96K dataset for your own training. Do not adopt it as a hosted research assistant: there is no service, the demo is a Hugging Face Space, and the GAIA path still needs a Serper API key. Verify first whether your hardware can hold the ~11B-token corpus and the pyserini and faiss-cpu indexes, whether you are willing to pin vllm==0.13.0, and what licence actually covers the code, since the repository does not state one.
Frequently asked questions
What is OpenResearcher?
It is a fully open pipeline for long-horizon deep research, consisting of a 96K trajectory dataset with 100+ turns generated by GPT-OSS-120B with native browser tools, a 30B-A3B model trained on that data, a distillation recipe, and a lightweight evaluation framework. The README describes it as an agentic large language model for deep research scenarios.
Is ChatGPT a good research tool compared with OpenResearcher?
The README does not compare OpenResearcher with ChatGPT as a research tool. It reports 54.8% accuracy on BrowseComp-Plus and claims OpenResearcher surpasses GPT-4.1, Claude-Opus-4, Gemini-2.5-Pro, DeepSeek-R1 and Tongyi-DeepResearch, but it does not describe how those baselines were run.
How do I install OpenResearcher?
Clone the repository, copy .env.template to .env, and run setup.sh, which the README lists as the setup entry point. Python 3.12 or newer is required by pyproject.toml, and vllm is pinned to version 0.13.0.
Do I need a search API to run OpenResearcher?
Not necessarily. The README documents two paths: BrowseComp-Plus with a local search engine built over a dedicated ~11B-token corpus, and GAIA with the Serper API, which it describes as needing no local search.
What hardware does OpenResearcher need?
The README does not publish a minimum VRAM figure for the 30B-A3B model or a size and memory figure for the ~11B-token corpus and its pyserini and faiss-cpu indexes. Both need to be sized from the model and corpus dataset pages before you commit to the local-search path.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/tiger-ai-lab-openresearcher)