ToolOrchestra: NVIDIA's RL Framework for Training Tool-Orchestrating Models
ToolOrchestra is an end-to-end RL training framework for orchestrating tools and agentic workflows.
At a glance
- What is it?
- ToolOrchestra is NVIDIA's end-to-end reinforcement learning framework for training small orchestrator models that coordinate tools and larger language models to solve complex tasks. The released Orchestrator-8B outperforms GPT-5 on three benchmarks while using roughly 30 percent of the cost.
- Who is it for?
- ToolOrchestra is the right starting point for ML researchers who want to study or reproduce the orchestration training pipeline described in the accompanying arxiv paper on model and tool orchestration. It requires NVIDIA HPC infrastructure with multiple H100 GPUs, a Tavily search API key, and familiarity with Slurm environments.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Activity is slowing. The repository last received commits 6 months ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What ToolOrchestra Solves and Who Uses It
Large language models struggle with tasks that require calling multiple specialized tools in sequence, because a single model must both reason about strategy and execute each step. ToolOrchestra addresses this by separating the orchestration problem: a small 8B orchestrator model decides which tool to call and when, while the actual computation happens in a separate, more capable system. The orchestrator can route to web search, a code interpreter, domain-specific coding or math models, or generalist models such as GPT-5, Llama-Nemotron-Ultra-253B, or Claude Opus 4.1.
The primary audience is ML researchers who want to reproduce or extend the results published in the accompanying paper, and teams at institutions with NVIDIA HPC access who want to train orchestration models using the ToolScale dataset and the provided RL pipeline. The project is not structured as a general-purpose library for deploying orchestration in production systems.
How the Orchestrator Architecture Works
The Orchestrator model runs a multi-turn loop: it receives a task, emits a reasoning step, calls a tool, receives the tool output, and repeats until it produces a final answer. This loop is what the README calls alternating between reasoning and tool calling in multiple turns.
During training, the Orchestrator is jointly optimized by three signals. An outcome reward measures whether the final answer is correct. An efficiency reward penalizes unnecessary tool calls. A preference reward distinguishes between tool call sequences that reach the same answer at different costs. The three rewards are combined via end-to-end reinforcement learning rather than supervised fine-tuning alone.
To supply enough training signal for RL, the project includes an automatic pipeline that synthesizes both environment tasks and tool-call tasks at scale. The resulting dataset, called ToolScale, was released on Hugging Face. Training data generation and the RL loop together require the vLLM inference backend and flash-attention, which limits the pipeline to CUDA-capable hardware.
The tool set the Orchestrator interacts with during training spans basic utility tools (web search via Tavily and a code interpreter), specialized smaller models focused on coding or mathematics, and generalist large models. The Orchestrator does not run the tools itself: it emits structured tool-call requests, which the evaluation harness routes to the correct service. This separation means the Orchestrator's weights encode call strategy, not domain knowledge, which is why a model of only 8 billion parameters can outperform much larger generalist models on benchmarks that require multi-step reasoning.
Setting Up the Training and Retrieval Environments
ToolOrchestra uses three separate Conda environments: one for training, one for retrieval, and one for running vLLM models during evaluation. Setting them up requires cloning two additional Hugging Face repositories for the index files and model checkpoints, then setting four environment variables before any scripts run.
The training environment setup starts with:
conda create -n toolorchestra python=3.12 -y
conda activate toolorchestra
pip install -r requirements.txt
pip install flash-attn --no-build-isolation
pip install flashinfer-python -i https://flashinfer.ai/whl/cu124/torch2.6/
pip install -e training/rolloutThe retrieval environment adds FAISS with GPU support:
conda create -n retriever python=3.12 -y
conda activate retriever
conda install pytorch==2.4.0 torchvision==0.19.0 torchaudio==2.4.0 pytorch-cuda=12.4 -c pytorch -c nvidia
pip install transformers datasets pyserini psutil
conda install -c pytorch -c nvidia faiss-gpu
pip install uvicorn fastapiA third environment (`vllm1`) handles vLLM-based model serving during evaluation. The README also requires a Tavily API key for web search, set via the `TAVILY_KEY` environment variable before running any evaluation script.
One important note: the Tau2-Bench evaluation requires rebuilding a Docker image with an Enroot environment running on a Slurm cluster, using GPU partitions with 8 H100 GPUs. This is not a setup that works on a standard workstation.
Running Training and Evaluating on Benchmarks
Once the environments are configured and the checkpoint and index directories are set, training starts with a single command from the `training` directory:
cd training
python resume_h100.pyThe filename `resume_h100.py` reflects that the script is designed to run on H100 nodes and supports resuming interrupted runs. The README explains that parallel experiments can be run by modifying three variables, `{EXPERIMENT_NAME1}`, `{EXPERIMENT_NAME2}`, and `{EXPERIMENT_NAME3}`, in `training/resume_h100.py`, with each variable corresponding to a file name in the experiment directory.
Evaluation covers three benchmarks:
cd evaluation
python run_hle.py
python run_frames.pyThe HLE evaluation requires both the `vllm1` and `retriever` environments running simultaneously. The README notes that connection errors to host models during HLE evaluation can be avoided by commenting out one line in `run_hle.py` and then running the generation and evaluation scripts as separate processes.
The published results show Orchestrator-8B at 37.1 percent on HLE versus GPT-5 at 35.1 percent, with 2.5x lower compute cost. On tau2-Bench and FRAMES, Orchestrator-8B surpasses GPT-5 while using approximately 30 percent of the cost.
Customizing Tools, Prompts, and LLM Backends
The README documents four customization points. First, LLM backends: modifying the `get_llm_response` function in `LLM_CALL.py` lets callers replace vLLM or OpenAI with any other inference service. Second, prompts: lines 455 to 458 in `eval_hle.py` and 506 to 509 in `eval_frames.py` contain the system and user prompts that the Orchestrator sees.
Third, tool configuration: the `tool_config` variable in line 27 of both `eval_frames.py` and `eval_hle.py` controls which tools are available to the Orchestrator during evaluation. Fourth, individual tools and models are defined in `tools.json` together with the `call_tool` function in `eval_hle.py`.
These four entry points give researchers a way to swap in a different tool set or a different LLM provider without rewriting the evaluation harness from scratch. The customization surface is narrow, though: there is no plugin API or configuration file format, and changes require direct code edits at specific hard-coded line numbers, which makes tracking upstream changes difficult when the repository is updated.
Maintenance Status, Limitations, and Comparison with Other Frameworks
The last push to the repository was on 2026-03-25. There are no formal GitHub releases. The repository is not archived, but it was created to accompany a research paper rather than as a sustained engineering product.
The project's main limitation is infrastructure dependency. The training pipeline assumes Slurm-managed clusters with H100 GPUs and Enroot container management. Researchers without access to this environment cannot run training and may face difficulties reproducing even the evaluation numbers exactly, since vLLM performance and model serving behavior vary across hardware configurations. There is no documented path for running training on a single machine or in a cloud VM.
The requirements.txt file spans hundreds of pinned packages, which creates a significant dependency-management burden when upgrading or porting the code. The file includes RAPIDS packages (cuDF, cuGraph, cuML) that are commented out but listed, indicating the environment was developed on specialized NVIDIA infrastructure.
For teams that want to build agentic pipelines without training a new model, LangGraph from LangChain is the nearest alternative. LangGraph provides a graph-based agent loop with support for multiple tool calls and conditional edges, but it is not an RL training framework: it is a runtime orchestration library. ToolOrchestra's contribution is the training methodology that produced Orchestrator-8B, not a reusable runtime API.
The Apache-2.0 license permits commercial use and redistribution without requiring derivative works to be open-sourced.
Editorial conclusion
ToolOrchestra is the right starting point for ML researchers who want to study or reproduce the orchestration training pipeline described in the accompanying arxiv paper on model and tool orchestration. It requires NVIDIA HPC infrastructure with multiple H100 GPUs, a Tavily search API key, and familiarity with Slurm environments. Teams who only want to deploy a capable orchestrator model without training one can download Orchestrator-8B from Hugging Face directly without running the training code. The codebase does not have formal releases on GitHub, and the last push to the repository was on 2026-03-25, so check the repository for any upstream fixes before building on it.
Frequently asked questions
What hardware does ToolOrchestra require for training?
The training script is named resume_h100.py and is designed to run on NVIDIA H100 GPUs within a Slurm cluster environment. The Tau2-Bench evaluation additionally requires an Enroot container with 8 GPUs. The README does not document a path for training on consumer-grade hardware.
Can I use Orchestrator-8B without running the ToolOrchestra training pipeline?
Yes. The Orchestrator-8B model weights are published on Hugging Face at nvidia/Orchestrator-8B and can be downloaded and used independently of the training code in this repository.
What is the ToolScale dataset used for?
ToolScale is a synthetic dataset of tool-call tasks generated by the automatic data synthesis pipeline included in this repository. It is hosted on Hugging Face under nvidia/ToolScale and provides the training signal for the reinforcement learning loop that produces the Orchestrator model.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/nvlabs-toolorchestra)