Model or dataset
TencentCloudADP/youtu-agent avatar
TencentCloudADP/youtu-agent

Youtu-Agent: Tencent's agent framework that ships benchmark numbers with it

A simple yet powerful agent framework that delivers with open-source models

4,622 stars485 forksPythonNOASSERTION

At a glance

What is it?
Youtu-Agent is Tencent's MIT-licensed Python framework for building, running and evaluating autonomous agents, built on openai-agents and delivering data analysis, file processing and deep research with open-weight models like DeepSeek-V3. It reports state-of-the-art results on WebWalkerQA and GAIA, generates agent configurations and tools automatically, and adds experience-based learning through Training-Free GRPO plus end-to-end RL scaling to 128 GPUs.
Who is it for?
Use Youtu-Agent when agent work should run on open-weight models, DeepSeek and similar, and the evaluation side matters as much as the running side, since benchmarks, tracing and RL training pipelines are built in rather than bolted on. Choose a lighter agent loop when the research apparatus is overhead.
Can I use it commercially?
Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
Is it still maintained?
Activity is slowing. The repository last received commits 6 months ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The benchmark claims, stated with models named

The framework's performance section leads its feature list, and the claims are specific enough to check. On WebWalkerQA, it achieved 60.71 percent accuracy with DeepSeek-V3-0324, improving to 71.47 percent with the newer DeepSeek-V3.1, described as setting a new SOTA. On GAIA, it achieved 72.8 percent pass@1 on the text-only validation subset, both runs using purely open-weight models and establishing what the project calls a strong open-source baseline. The consistent pattern, naming the exact model version beside each score, is the honest form of benchmark reporting, since an agent framework's results are a product of the model underneath as much as the harness around it, and the collection of RL checkpoints on Hugging Face extends the verification to trained artifacts. The practical use cases named beside the benchmarks, data analysis, literature review, personal file organization, retrieval-augmented generation and PPT generation, show where the framework expects its agents to land, and the PPT generation and RAG examples ship in the repository as working starting points.

Two generation paradigms: Workflow and Meta-Agent

Automated agent generation is the framework's second pillar, introducing two paradigms, a Workflow mode for standard tasks and a Meta-Agent mode for complex requirements. The framework supports automated generation of tool code, prompts and configurations, achieving over 81 percent tool synthesis success rate, the figure describing how often the described capability becomes working tool code. A September 2025 update added companion tooling, describe a capability once and let Youtu-Agent build the tool for you. The design intent is to move agent construction from hand-writing configuration files toward specification, with the two modes covering the range from predictable workflows to open-ended requirements the meta-agent decomposes itself.

Training-Free GRPO: RL at eight dollars

The most distinctive research contribution is Training-Free Group Relative Policy Optimization, and its announcement reads like a challenge, RL for DeepSeek-V3.2 at 8 dollars, yes, it's possible. The method keeps the model frozen, learns a token prior from roughly 100 samples, and delivers the gains through in-context optimization without parameter updates, with verified math and web search improvements and the Agent Practice module reporting a plus 5.4 percent gain on AIME 2025. The technique landed in the main branch in November 2025, with the practice documentation covering math reasoning and web search tasks, and it reframes what agent improvement costs, no GPUs, no fine-tuning infrastructure, just inference budgets and accumulated experience.

Agent RL scaling to 128 GPUs

For full parameter training, the Agent RL module provides a complete pipeline for end-to-end reinforcement learning, integrating with distributed frameworks to address stability and scalability challenges, achieving 40 percent training speedup and scaling to 128 GPUs. The December 2025 collaboration with Microsoft's Agent-Lightning team implemented efficient model training in various scenarios with multi-node deployment on 128 GPUs, and the work lives on the rl/agl branch. The two learning paths sit at opposite ends of the cost spectrum, Training-Free GRPO at roughly 8 dollars per run and full RL on a GPU cluster, and the framework's positioning is that an agent team can start with the former and graduate to the latter without changing frameworks.

Configuration through env, YAML and extras

The basic environment file shows what a deployment needs, the LLM settings carrying UTU_LLM_TYPE set to chat.completions, UTU_LLM_MODEL defaulting to deepseek-chat, the DeepSeek base URL and an API key, plus tool keys, SERPER_API_KEY for web search and JINA_API_KEY for content extraction. Optional settings cover tracing through a Phoenix OTEL endpoint with project name and cloud key, a database URL for tracing and evaluation data defaulting to SQLite, and a log level. Agent configurations live as YAML files in the configs tree, the RAG example being a named config, and the package's optional dependencies partition capabilities, litellm for additional model providers, search through crawl4ai, scholar through the arxiv client, wiki, e2b for code execution sandboxes, and local-python for the in-process interpreter.

Built on openai-agents, with tracing built in

The architecture is stated as built on openai-agents, OpenAI's Python agent framework, with extensible support for diverse model APIs from DeepSeek to gpt-oss, tool integrations and framework implementations, the pin at exactly 0.10.4 in the dependency list showing how tightly the base is held. Observability is a dependency-level feature rather than an integration, with openinference instrumentation for both the OpenAI clients and the agents library, plus the OpenTelemetry OTLP exporter, feeding the Phoenix tracing setup from the environment file. Storage and display round out the stack, sqlmodel with psycopg2 for the evaluation database, rich and prompt-toolkit for the terminal experience, hydra for configuration composition, and jinja2 for the prompt templates the generation system fills. The evaluation database storing tracing and assessment runs is what connects the running side to the learning side, since the experience the practice module distills has to come from recorded agent behavior somewhere.

Agent Skills, and the Youtu-Tip desktop sibling

The January 2026 news added Agent Skills, extending agents with modular, domain-specific knowledge and workflows inspired by Anthropic's Claude Code skills, with documentation on the project site. A separate announcement introduced Youtu-Tip, an extension running on macOS powered by offline models through Ollama, automating file reading and web browsing, with the promise that agents built with Youtu-Agent will run through it more easily, and Youtu-LLM, a model released inside the youtu-tip repository. The family structure, a Python framework, a desktop wrapper and a model, mirrors how the Tencent Cloud ADP organization positions the project, the open-source research core beside the commercial Agent Development Platform the announcements reference for enterprise solutions.

A research project with a frontend and docs site

The repository carries more than a Python package, a frontend directory with a web UI built through npm and shipped as a wheel, a docker directory, mkdocs documentation deployed to the GitHub Pages site, demo scripts, and a Makefile orchestrating sync, format, lint, docs and the UI build through uv. The development setup is uv-native, make sync pulling all extras and packages plus the dev group, ruff handling format and lint. Releases are tagged sparsely, v0.1.2 and v0.1.3 in October 2025 around the frontend's v0.3.0, with the last push on 2026-03-21, and the MIT license sits beside an arXiv paper and a DeepWiki entry, the documentation triangle of a project published for both engineers and researchers.

Editorial conclusion

Use Youtu-Agent when agent work should run on open-weight models, DeepSeek and similar, and the evaluation side matters as much as the running side, since benchmarks, tracing and RL training pipelines are built in rather than bolted on. Choose a lighter agent loop when the research apparatus is overhead. Before adopting, note the package is version 0.1.3 and marked Beta with the last push in March 2026, configure the LLM and search keys in .env before running, and pick extras deliberately, the litellm, search, scholar, wiki, e2b and local-python sets each carry their own dependencies.

Frequently asked questions

What is Youtu-Agent?

Youtu-Agent is Tencent's MIT-licensed Python framework for building, running and evaluating autonomous agents, built on openai-agents. It delivers data analysis, file processing and deep research with open-weight models like DeepSeek-V3, reports 71.47 percent on WebWalkerQA and 72.8 percent on GAIA, and supports automated agent generation plus experience-based learning through Training-Free GRPO.

How do you configure Youtu-Agent?

Set the LLM settings in .env, UTU_LLM_TYPE, UTU_LLM_MODEL, UTU_LLM_BASE_URL and UTU_LLM_API_KEY, defaulting to deepseek-chat at the DeepSeek API, plus tool keys like SERPER_API_KEY for web search and JINA_API_KEY for extraction. Optional Phoenix tracing, a database URL for evaluation data, and YAML agent configs in the configs directory complete the setup.

What is Training-Free GRPO in Youtu-Agent?

Training-Free GRPO is an experience-based learning method that keeps the model frozen and learns a token prior from about 100 samples, improving agent performance through in-context optimization without parameter updates, at roughly 8 dollars per run. The Agent Practice module reports gains like plus 5.4 percent on AIME 2025 through the technique.

Official sources

  1. Issues
  2. Project website
  3. README
  4. Releases
  5. TencentCloudADP/youtu-agent on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/tencentcloudadp-youtu-agent.svg)](https://hysenlabs.com/projects/tencentcloudadp-youtu-agent)