Model or dataset
microsoft/RD-Agent avatar
microsoft/RD-Agent

microsoft/RD-Agent: An LLM Agent Framework That Runs Its Own Data Science Loops

Research and development (R&D) is crucial for the enhancement of industrial productivity, especially in the AI era, where the core aspects of R&D are mainly focused on data and models. We are committed to automating these high-value generic R&D processes through R&D-Agent, which lets AI drive data-driven AI. đź”—https://aka.ms/RD-Agent-Tech-Report

14,657 stars1,918 forksPythonMIT

At a glance

What is it?
RD-Agent automates the iterative loop of hypothesis, code, experiment and feedback for data and model work. It is a research-grade Python package, not a finished product, and the docs are thinner than the paper list suggests.
Who is it for?
Adopt RD-Agent if you already run reproducible experiments and want an agent to propose and execute the next iteration inside that harness; the quant and Kaggle scenarios are the most developed parts. Do not adopt it if you need a supported product with a stable API, because pyproject.toml still classifies the package as Development Status 3 - Alpha and the README documents no rollback or recovery path for a bad run.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 15 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 17, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The loop RD-Agent is trying to close

Most of the work in a data science or quant project is not the first model. It is the second, fifth and twentieth iteration: form a hypothesis, write the code, run it, read the metric, decide what to change. RD-Agent targets that loop. The README frames it as automating "these high-value generic R&D processes" so that AI drives data-driven AI, and the package name in pyproject.toml is simply rdagent, described as "Research & Development Agent".

The intended user is not someone who wants a chatbot that writes pandas snippets. It is a team that already has an experiment harness and wants an agent to occupy the loop inside it. The repository ships scenarios rather than a single product, and the two the README leans on hardest are MLE-bench style machine learning engineering and R&D-Agent-Quant for trading factors. The package is Python-only, requires 3.10 or newer, and the platform badge in the README says Linux.

How the agent loop is wired: scenarios, LiteLLM and a workspace

Three mechanisms are visible in the repository files. The first is the scenario concept: rdagent/scenarios/ holds separate agent stacks for different problem classes, and the README points at finetune/llm/README.md for the LLM fine-tuning scenario and rdagent/scenarios/rl/autorl_bench/README.md for the RL benchmark. A scenario is what supplies the domain rules; the core package supplies the loop.

The second is the backend. The README states that LiteLLM is now the default backend for integration with multiple LLM providers, and .env.example confirms it with BACKEND=rdagent.oai.backend.LiteLLMAPIBackend commented out as the default value. That matters because it decouples the agent from any single vendor: the same run can point at an OpenAI-compatible endpoint or at a proxy.

The third is the workspace. The pytest configuration in pyproject.toml lists workspace under norecursedirs, and requirements.txt mentions `tensorboard --logdir git_ignore_folder/RD-Agent_workspace` for visualizing SFT training. So a run writes artifacts to a local directory that is deliberately excluded from test collection, which is also where you would look when a run goes wrong.

Installing RD-Agent and running a first scenario

The package is published as rdagent on PyPI, and pyproject.toml defines the console entry point as `rdagent = "rdagent.app.cli:app"`, so the CLI is available after install. The Makefile also has an `install` target for working from a checkout.

Start by installing into a Python 3.10 or 3.11 environment, since those are the only two classifiers listed in pyproject.toml.

bash
pip install rdagent

Next, configure the model backend. The .env.example file is the template: copy it to .env and fill in values. The minimal set is a chat model, an embedding model and an API base.

bash
cp .env.example .env
bash
OPENAI_API_KEY="sk-chat-key"
OPENAI_API_BASE="https://xxx-litellm.com/v1"
CHAT_MODEL='gpt-4o'
EMBEDDING_MODEL="litellm_proxy/BAAI/bge-large-en-v1.5"

Note the `litellm_proxy/` prefix on the embedding model in that example. The comment in .env.example says to "pay attention to the litellm_proxy prefix" when the embedding service is reached through a proxy, and the example uses siliconflow with LITELLM_PROXY_API_KEY and LITELLM_PROXY_API_BASE. Getting that prefix wrong is the kind of thing that fails at the first embedding call rather than at startup.

Finally, invoke the CLI. The README does not print a full command line for a scenario, so check the scenario README under rdagent/scenarios/ for the arguments your chosen scenario expects before running it.

bash
rdagent

There is also a web frontend. The README states it "can be built and served by `rdagent server_ui` for real-time interaction and trace viewing, currently excluding the `data_science` scenario", which is the clearest statement of what the UI does and does not cover.

Where RD-Agent is the wrong tool

The package is classified as Development Status 3 - Alpha in pyproject.toml. That is the project's own label, and it should shape how you deploy it: not as a service other teams depend on, but as something you run and watch.

The README documents no rollback, no resume-from-checkpoint command and no way to prune a bad branch of the search. If an agent run produces a broken factor or a training script that consumes GPU hours, the recovery path is the workspace directory and your own git history, not a project feature. For a loop whose whole purpose is to try many things and discard most of them, that is a real gap.

The dependency list is also heavy for what is often described as an agent framework. requirements.txt pulls in selenium, kaggle, mlflow, azureml-mlflow, streamlit, plotly, flask, prefect, duckduckgo-search and pydantic-ai-slim, among others. If you only want the quant factor loop, you are still installing a crawler stack and a Streamlit app. There is a constraints/ directory and a requirements/ directory in the repository, so pinning is possible, but the default install is broad.

Finally, if your task is a one-off analysis you will finish by hand in an afternoon, an agent that proposes hypotheses and runs experiments is overhead. The value only appears when the loop repeats.

RD-Agent against AIDE on the same benchmark

The README reports MLE-bench results with R&D-Agent at 30.22 percent on the All column using o3 as the reasoning model and GPT-4.1 as the development model, and 22.4 percent with o1-preview, against AIDE with o1-preview at 34.3 percent on the Low == Lite column. Those numbers are the project's own table, and the two entries are not directly comparable because they use different model pairings.

The difference in approach is the more useful comparison. AIDE is a single-agent tree search over solution edits: one agent, one evolving candidate pool. RD-Agent separates the reasoning model from the development model and organises work into scenarios with their own feedback loops, which is why the README can point at separate implementations for quant factors, LLM fine-tuning and RL post-training under one package. If you want one agent iterating on a Kaggle submission, AIDE's shape is simpler. If you want the same machinery reused across several research domains with different evaluation signals, the scenario split is the reason to pick RD-Agent.

Maintenance, licence and the cost of upgrading

The repository is not archived, and the last push was on 2026-09-04. Releases move in steps rather than continuously: v0.6.1 on 2025-06-28, v0.7.0 on 2025-07-08, v0.8.0 on 2025-11-03. The gap between v0.8.0 and the current head means a lot of the README's news items, including the Agent² RL-Bench preprint and the ICML 2026 and ACL 2026 acceptances, sit outside any tagged release. If you pin to v0.8.0 you are not getting what the README describes at the top of the page.

Versioning is handled by setuptools-scm, with a .bumpversion.cfg at the repository root, so versions are derived from tags. Upgrading therefore means reading the CHANGELOG.md rather than trusting a semantic version promise, and the Alpha classifier means breaking changes between minor versions are within the project's stated posture.

The licence is MIT, declared in LICENSE and in the pyproject.toml classifier "License :: OSI Approved :: MIT License". MIT is permissive: you can use it commercially and modify it. The practical implication to check is not the licence text but the model providers you wire in through LiteLLM, since their terms govern the API calls the agent makes. That is a question for your own legal review, not something the repository settles.

Editorial conclusion

Adopt RD-Agent if you already run reproducible experiments and want an agent to propose and execute the next iteration inside that harness; the quant and Kaggle scenarios are the most developed parts. Do not adopt it if you need a supported product with a stable API, because pyproject.toml still classifies the package as Development Status 3 - Alpha and the README documents no rollback or recovery path for a bad run. Before committing, verify two things yourself: that your chosen LLM endpoint works through the LiteLLM backend with the CHAT_MODEL and EMBEDDING_MODEL pair in .env, and that the scenario you care about is not the data_science one, since the README states the new `rdagent server_ui` frontend currently excludes it.

Frequently asked questions

What is RD-Agent from Microsoft?

It is a Python framework that automates research and development loops for data and model work, published as the rdagent package on PyPI under the MIT licence. The README describes it as letting AI drive data-driven AI, with scenarios for machine learning engineering, quant trading and LLM fine-tuning.

How do I use RD-Agent?

Install the rdagent package, copy .env.example to .env and set CHAT_MODEL, EMBEDDING_MODEL and an API base, then invoke the rdagent console script. The README does not print a full scenario command line, so read the scenario README under rdagent/scenarios/ for the arguments it expects.

What is rd agent in a post office?

That is a different subject. RD Agent in a post office refers to a postal or rural delivery agent role, and has nothing to do with the microsoft/RD-Agent software package described here.

What are the four types of agents?

The README does not describe a taxonomy of four agent types. It instead organises RD-Agent by scenario, listing a general data science agent, a Kaggle agent, a quant scenario and an LLM fine-tuning scenario.

What is pdq rd agent?

Nothing in the repository describes a pdq rd agent. The only RD-Agent covered here is the microsoft/RD-Agent package, whose README, pyproject.toml and requirements.txt are the sources used throughout this article.

Official sources

  1. License: MIT
  2. microsoft/RD-Agent on GitHub
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/microsoft-rd-agent.svg)](https://hysenlabs.com/projects/microsoft-rd-agent)