ThunderAgent and the one field it adds to a request
A simple, fast and robust program-aware agentic inference system.
At a glance
- What is it?
- An agentic inference scheduler that sits between your agent client and a vLLM or SGLang deployment, grouping work into programs so the KV cache is reused and memory stays balanced across nodes. Integration is a single extra field in an otherwise OpenAI-shaped request.
- Who is it for?
- It fits a team running agent rollouts at volume on multi node GPU infrastructure, where cache reuse and memory balance across nodes are the bottleneck and the same request library already speaks the OpenAI shape. It does not fit a single GPU experiment or anyone who would have to fork their agent client to add a field.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 97 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 9, 2026, and from our analysis. They are not legal advice.
Editorial analysis
It sits between your agent client and the GPUs
The placement is the whole idea. ThunderAgent is described as sitting between agent clients and the infrastructure layer, working as an agentic workflow scheduler rather than as a model server of its own. From one side it improves inference throughput for vLLM or SGLang across multiple GPU nodes, and from the other it offers a unified tool management interface for resources such as Docker containers and remote APIs. That second job is the one that distinguishes it from a routing layer: an agentic workflow holds on to things that are not tokens, a container it started, a remote endpoint it called, and those need reclaiming when the run ends. The stated target use is agentic inference and rollout, which is why the framing throughout is a library for producing many agent trajectories rather than for serving one chat. The supporting links match that audience: a wiki, a blog on the project's own domain which doubles as the repository homepage, and a paper. There is also a demo, delivered as a hosted video attachment rather than as anything you can run, which is the right trade for a scheduler whose interesting behaviour only appears under load.
The program is the scheduling unit, and cache hits are the point
Program-aware scheduling is the mechanism behind the speed claim. A program is the unit the scheduler reasons about, and the stated effect of that grouping is a higher KV cache hit rate plus less memory imbalance across nodes. Both matter for rollouts: a cache miss means recomputing prefix tokens for a branch that differs from its neighbours only slightly, and imbalance means one node holds the long-context requests while another idles. The reported result is a throughput increase of 1.5 to 3.6 times across agentic workloads, with SWE-Agent, OpenHands and ToolOrchestra named as the ones measured. That the abstraction is treated as schedulable rather than advisory is the part worth noticing, since the news section records integration into NVIDIA Dynamo as part of Dynamo 2.0, where the proposed program abstraction operates as a first-class scheduling unit. A concept that upstream accepts as a scheduling primitive is a stronger signal than a benchmark table. One caveat belongs next to the numbers: the README states the 1.5 to 3.6 times range twice, once as throughput across agentic workflows and once as vLLM throughput across SWE-Agent, OpenHands and ToolOrchestra, and it does not describe the benchmark method, the request mix or the node counts behind either figure. Treat the range as the project's own evaluation of its own system.
Integration is one field in the request body
The API surface is the smallest part of this project and the most useful thing about it. Rather than asking you to adopt a client, the interface is an OpenAI-compatible passthrough where the only change is adding one field. The embedding example in the documentation shows the same chat completions call before and after, with the original passing model and messages, and the ThunderAgent version building an extra_body dictionary that carries a program_id value, plus an optional docker_ids list when the workflow runs in containers:
# original openai sender
openai.client.chat.completions.create(
model=self.config.model_name,
messages=messages,
)
# ThunderAgent openai sender
extra_body = {}
extra_body["program_id"] = "unique_id"
# if you use docker for your agentic workflow
# extra_body["docker_ids"] = ["docker_id1", "docker_id2", ...]
openai.client.chat.completions.create(
model=self.config.model_name,
messages=messages,
extra_body = extra_body
)The program_id is what tells the scheduler which trajectories belong together, so it is the one value you have to get right in your agent loop.
Two ports: 8000 for the model, 9000 for the scheduler
The getting started path is short and makes the topology obvious. Installation is from source:
git clone https://github.com/ThunderAgent-org/ThunderAgent.git
cd ThunderAgent
pip install -e .Then you pick one backend, and the example uses vLLM, which means serving a model yourself before the scheduler has anything to sit in front of:
uv pip install vllm --torch-backend=auto # install vllm
vllm serve Qwen/Qwen3-32B --port 8000 # serve a model
thunderagent --backend-type vllm --backends http://localhost:8000 --port 9000 --metrics --profile # launch ThunderAgent, make sure to send request through 9000.The last line is the one to copy carefully. ThunderAgent listens on 9000 and forwards to the model server on 8000, and the trailing comment in the command says in as many words that requests must be sent through 9000. Get that backwards and you will be measuring your model server directly, with no scheduling and no metrics. Installation is an editable install from a clone, and that is the only route given: no wheel, no package index name and no version pin appear anywhere in the getting started section, so you are expected to work against the checkout you just made.
The dependency list explains what kind of thing this is
The package metadata is the fastest way to understand the shape of the project, and it is unusually small. Three runtime dependencies are declared: fastapi, httpx and uvicorn. That is an HTTP service that forwards requests, not a machine learning framework, and nothing in the list pulls in a tensor library, because the heavy inference engine is expected to be running separately as vLLM or SGLang. The declared support is Python 3.10 or newer, the package name is ThunderAgent, and the console script maps the thunderagent command to the package's main entry point. The version is 0.1.0 and the project describes itself in the metadata as an agentic scheduler with program state tracking, which is a narrower description than the marketing copy and a more accurate one about what the code holds. Package discovery is scoped to the ThunderAgent prefix, which matches a single package directory at the root of the repository, and a wiki directory sits alongside it mirroring the hosted wiki link.
Rollouts are the primary use, not an afterthought
Several details only make sense if your workload is many short agent runs rather than one long conversation. Tool-call lifecycle management with automatic resource reclaim is listed as a core capability, described as the thing that makes long running rollouts more stable, and it pairs naturally with the docker_ids field: a rollout that starts containers has to give them back, or a few thousand iterations later the host is out of resources. The training examples make the same assumption. The documentation lists agentic RL training setups built on a Search-R1 agent with slime, and mini-swe-agent with SkyRL, and the news section records the SkyRL integration as including an example training recipe that accelerates SWE Agent rollout by 3x on 40 H100 GPUs. Observability is oriented the same way, with real-time visualization of trajectory metrics including total tokens, tool-use time and per-program profiling, which are exactly the two flags in the launch command.
Scale adoption runs through an email address
The project publishes a paper, accepts to a venue and names authors, so this is research code with an intent to be used beyond a paper. The news section records acceptance to ICML 2026 as a Spotlight, described as the top 2.2 percent, alongside the Dynamo and SkyRL integrations. The citation block in the README lists ten authors, an arXiv identifier of 2602.13692 and a primary class of cs.OS, so the systems framing is explicit. Alongside that, the contact section invites enterprises interested in adopting or deploying ThunderAgent at scale, including technical consulting, sponsorship and partnership inquiries, at a single academic address. That combination, a paper, a venue acceptance, two upstream integrations and a consulting contact, is the profile of a project whose next step is deployment rather than further benchmarking. The repository itself is MIT licensed through a LICENSE.md file at the root, and its layout keeps the package in a ThunderAgent directory with examples split into datagen, inference and rl_training, plus assets and a wiki directory. There are no GitHub releases, and the last push landed on 2026-07-05 with the project not archived.
Editorial conclusion
It fits a team running agent rollouts at volume on multi node GPU infrastructure, where cache reuse and memory balance across nodes are the bottleneck and the same request library already speaks the OpenAI shape. It does not fit a single GPU experiment or anyone who would have to fork their agent client to add a field. Before you adopt it, check the two things the documentation leaves open: the reported 1.5 to 3.6 times throughput range comes from the project's own evaluation across named workloads rather than from your workload, and the package is still version 0.1.0 despite the integrations listed in the news section.
Frequently asked questions
What does ThunderAgent add to an OpenAI-style request?
One field. The API is an OpenAI-compatible passthrough, and you add a program_id inside extra_body, plus an optional docker_ids list when the agentic workflow runs in containers.
Which inference backends does ThunderAgent support?
vLLM and SGLang. You choose one when launching, and the getting started example serves a model with vllm serve before pointing ThunderAgent at that address.
Which port should my agent send requests to?
Port 9000, where ThunderAgent listens. The example serves the model on 8000 and the launch command's own comment says to send requests through 9000.
Which agentic RL training stacks does ThunderAgent show examples for?
A Search-R1 agent with slime, and mini-swe-agent with SkyRL. The SkyRL integration is described as including a recipe that accelerates SWE Agent rollout by 3x on 40 H100 GPUs.
Is ThunderAgent open source and what does it require?
It is MIT licensed through a LICENSE.md file at the root. The package metadata declares version 0.1.0, requires Python 3.10 or newer, and depends only on fastapi, httpx and uvicorn.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/thunderagent-org-thunderagent)