RUC-NLPIR/DeepAgent: A Research Agent That Folds Its Own Memory
[WWW‘26 Oral🔥] DeepAgent: A General Reasoning Agent with Scalable Toolsets
At a glance
- What is it?
- DeepAgent is an MIT-licensed Python research codebase for a reasoning agent that discovers tools inside a single stream of thought and compresses its history with Autonomous Memory Folding. It is a paper implementation, not a product, and the README ships installation commands but no run instructions.
- Who is it for?
- Adopt DeepAgent if you are evaluating the paper's ideas, want the ToolPO training recipe, or need a codebase that scales tool discovery from tens to more than ten thousand APIs. Do not adopt it if you need a supported product with a documented CLI, a UI, or a stable release tag, because none of those exist in the repository.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 170 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The problem DeepAgent targets: tool selection at ten-thousand-tool scale
Most agent frameworks assume you already know which tools the model will call. You register a handful of functions, write a prompt, and let the model pick. DeepAgent starts from the opposite assumption. The README describes an agent that searches for and uses appropriate tools from more than 16,000 RapidAPIs, an end-to-end agentic reasoning process. The intended user is a researcher or an engineer building an agent against a toolset too large to enumerate in a prompt, or against a domain where the right action is not known in advance, such as ALFWorld navigation or web browsing. The paper behind it is arXiv 2510.21618, accepted at WWW 2026, and the repository is a Python implementation of that work rather than a hosted service.
How the reasoning loop differs from the ReAct cycle
The README frames the design as a departure from predefined workflows, naming ReAct's Reason-Act-Observe cycle as the thing being replaced. In DeepAgent, thinking, tool discovery and action execution happen inside one coherent stream, so the model keeps a global view of the task instead of rebuilding context at each step. The second mechanism is Autonomous Memory Folding. When a run risks getting stuck in a wrong exploration path, the agent can trigger a fold that consolidates the interaction history into a structured schema with three parts: episodic memory for key events and sub-task completions, working memory for the current sub-goal and near-term plans, and tool memory for tool-related interactions. The third piece is ToolPO, an RL method that pairs an LLM-based tool simulator with tool-call advantage attribution, which assigns credit to the specific tokens of a correct tool invocation. That last detail matters: it is a training-time mechanism, so you only benefit from it if you intend to train, not just to run inference.
Installing DeepAgent and getting to a first run
The README's Installation section begins with a conda environment pinned to Python 3.10, and the badge above it advertises Python 3.9+. Treat 3.10 as the tested path.
# Create conda environment
conda create -n deepagent python=3.10After creating the environment, the repository carries a requirements.txt at the top level. Installing from it pulls in the runtime surface, including vllm, transformers, openai, fastapi, aiohttp, and the domain packages alfworld, spotipy and exa-py. Expect a heavy install: vllm and transformers are the largest dependencies, and alfworld is only needed for the embodied tasks.
conda activate deepagent
pip install -r requirements.txtThe README does not document a run command, a CLI, or an entry script. The top-level layout shows src/, config/ and data/ alongside docs/, and the README points to QwQ and Qwen3 collections as the reasoning models to deploy with, so the practical first step is to read src/ and config/ to find the entry point your task expects. The README also links a DeepAgent-Datasets page on Hugging Face for data. If you want a first real use, pick one of the documented task families, general tool use, ALFWorld, or deep research, and locate the matching config rather than guessing at flags.
Where DeepAgent will disappoint you
The repository has no releases, so there is no version to pin and no changelog to read before upgrading. The README documents installation and nothing about operation: no CLI reference, no example invocation, no expected output, no rollback procedure if a run goes wrong. The demo section is candid about a related weakness. The 16,000-API demonstration uses LLM-simulated API responses because some APIs in ToolBench are unavailable, which means the flagship demo shows system behaviour rather than live API behaviour. If your evaluation depends on real third-party API latency, quota errors or schema drift, that demo will not tell you how the agent behaves. And if you need a graphical interface, a hosted endpoint, or an npm package, this is the wrong project: it is a Python research codebase with a config directory and a paper, and the README offers no UI.
DeepAgent versus LangGraph and LangChain
The search data around this project keeps pairing it with LangGraph and LangChain, and the comparison is real but asymmetric. LangGraph gives you an explicit graph: you define nodes, edges and state transitions, and the control flow is something you can read, test and constrain. DeepAgent removes that structure on purpose. Its README argues that predefined workflows prevent the model from holding a global perspective, and it replaces the graph with a single reasoning stream plus autonomous memory folding. The trade-off is inspectability. With a graph you can point at the node that failed. With DeepAgent, the decision to fold memory is made by the model, and the README does not describe a way to force or forbid a fold. If your requirement is auditable, replayable control flow, LangGraph is the more direct fit. If your requirement is an agent that copes with a toolset too large to enumerate, DeepAgent is built for that case.
Maintenance, licence and the cost of upgrading
The repository is not archived and the last push was on 2026-04-13. That is roughly five months before today, so the codebase is not stale, but the absence of releases means you track the main branch or pin a commit yourself. There is no upgrade path to follow, and requirements.txt pins nothing: vllm, transformers and openai are listed as bare names, so a fresh install months apart can resolve to different versions. For a research codebase that is normal, and it is also the main operational risk. The licence is MIT, which permits commercial use and modification provided the copyright notice and permission notice are retained; the repository's LICENSE file is the authoritative text. This is a description of the licence, not legal advice, and the dependencies carry their own terms, including vllm and the model weights you choose to serve.
Editorial conclusion
Adopt DeepAgent if you are evaluating the paper's ideas, want the ToolPO training recipe, or need a codebase that scales tool discovery from tens to more than ten thousand APIs. Do not adopt it if you need a supported product with a documented CLI, a UI, or a stable release tag, because none of those exist in the repository. Before committing, verify that src/ contains the entry point you need, that requirements.txt resolves against your Python version, and that the DeepAgent-Datasets page on Hugging Face has the split you intend to train on.
Frequently asked questions
What is RUC-NLPIR/DeepAgent?
It is an MIT-licensed Python research codebase implementing DeepAgent, described in the README as an end-to-end deep reasoning agent that performs autonomous thinking, tool discovery and action execution within a single reasoning process. The accompanying paper is arXiv 2510.21618 and was accepted at WWW 2026.
How do I install and use DeepAgent?
The README creates a conda environment with Python 3.10 and then installs the top-level requirements.txt. It documents no run command or CLI, so you have to read src/ and config/ to find the entry point for the task you want, and the README points to QwQ and Qwen3 as the reasoning models to deploy with.
Is DeepAgent free?
The repository is licensed under MIT, which permits use and modification with the copyright and permission notice retained. That covers the code only; the model weights you serve and the third-party APIs you call have their own terms.
Is DeepAgent good?
The README reports that DeepAgent achieves superior performance across the evaluated scenarios, covering ToolBench, API-Bank, TMDB, Spotify, ToolHop, ALFWorld, WebShop, GAIA and HLE. Those are the project's own reported results, and no independent reproduction is documented in the repository.
What is DeepAgent AI?
In this repository the name refers to a general reasoning agent with scalable toolsets, built by RUC and Xiaohongshu Inc. and presented at WWW 2026. It is code and a paper, not a hosted AI service.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/ruc-nlpir-deepagent)