# FoundationAgents/ReCode: recursive code generation as a plan-and-act agent loop

> ReCode replaces separate planning and acting stages with one tree of Python code that expands itself at runtime. It is a research reference implementation for ALFWorld, WebShop and ScienceWorld, not a drop-in agent framework.

**FoundationAgents/ReCode** — Next paradigm for LLM Agent. Unify plan and action through recursive code generation for adaptive, human-like decision-making.

- Repository: https://github.com/FoundationAgents/ReCode
- Stars: 570 · Forks: 68
- Language: Python
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/foundationagents-recode

## The problem ReCode targets: plan and action as two representations

Most LLM agent designs keep planning and acting apart. A planner produces a step, the executor runs it, and the observation goes back to the planner. The granularity is fixed by the prompt: either the model emits many small steps, or it emits one large step and loses the ability to react mid-way. ReCode's README frames this as the problem it solves, describing a design that unifies plan and action into a single representation so the agent can move between strategic thinking and concrete action without switching formats. The intended audience is researchers working on agent decision-making, not application teams looking for an SDK. The repository layout confirms that: agents/recode/ holds the implementation and prompt templates, envs/ holds wrappers for alfworld, webshop and sciworld, and run.py is the CLI that ties them together for experiments.

## How the recursive expansion loop actually works

The mechanism is divide and conquer over code. A partial program is stored as a tree, and each node holds one sub-task plus its execution trace. A node can be a placeholder function, meaning a call the model has not yet resolved. The LLM expands that placeholder into more specific calls or into smaller subroutines, using prompts and few-shots that are specific to the environment. Expansion is not a separate planning pass: the dynamic execution loop runs each node immediately, and the fresh observation decides whether to expand further, retry, or finish. Underneath sits a shared executor state, described as a constrained Python executor that maintains environment variables, validates code blocks, and exposes the toolset available to the agent. That constraint matters, because the model is generating executable Python rather than a JSON action schema, so validation is the boundary between a bad generation and a crashed run. The practical consequence is that depth is chosen at runtime. A task the agent recognizes gets executed directly; a task it does not gets decomposed until the leaves are executable primitives.

## Installing ReCode and running a first dry run

The README requires Python 3.10 or newer and a conda environment. It also warns that dependencies for the three environments have not been confirmed to coexist, and suggests configuring them in separate environments, so treat the environment name as a per-benchmark choice rather than a shared one.

```bash
conda create -n recode-envname python=3.10
conda activate recode-envname
```

For ALFWorld, follow the ALFWorld instructions and point the code at your dataset either through the environment variable or by editing envs/alfworld/base_config.yaml.

```bash
export ALFWORLD_DATA=/path/to/alfworld
```

WebShop is configured from inside its own directory, with a helper script that fetches the goal set and the pre-built search index.

```bash
cd envs/webshop
pip install -e .
conda install -y -c conda-forge openjdk=11
pip install "en_core_web_lg @ https://github.com/explosion/spacy-models/releases/download/en_core_web_lg-3.6.0/en_core_web_lg-3.6.0-py3-none-any.whl"
bash setup.sh
```

ScienceWorld defers to the ScienceWorld repository's own instructions. After the environment is in place, install the top-level requirements. The README flags that this list may be incomplete and asks you to make contact if something is missing, which is an honest signal about the state of packaging.

```bash
pip install -r requirements.txt
```

Before running anything, copy the profile template and put a real credential in it. The profiles file holds named profiles, and run.py --profile selects which one is forwarded to AsyncLLM.

```bash
cp configs/profiles_example.yaml configs/profiles.yaml
```

The example profile in the README uses api_key, base_url, model, temperature and track_costs keys, with a second profile showing max_tokens. Cost tracking reads configs/prices.json, and setting track_costs to false turns it off. If you skip the file entirely, the default profile falls back to the OPENAI_API_KEY environment variable. Then run one instance of the agent in one environment.

```bash
python run.py -a recode -e alfworld -n 1 --split test --profile default
```

Replace alfworld with webshop or sciworld once their assets are available. Logs go to logs/<run_id>/, and the console prints a condensed summary.

## What the evaluation numbers do and do not cover

The README reports an average score of 60.8 across the three environments, which it says surpasses the best baseline by 10.5, a relative 20.9%, against ReAct, CodeAct, AdaPlanner and ADaPT. It also states that under claude-4-sonnet ReCode reaches a score of 100 in ALFWorld. Those are the authors' reported results, and they come with the usual caveat that the configurations behind them are the ones in this repository, so reproducing them means reproducing the environment setup exactly. The training side is more interesting for anyone judging data efficiency: supervised fine-tuning on Qwen2.5-7B-Instruct gives ReCode+SFT 70.4% average across environments, against 67.6% for ReAct+SFT and 55.8% for CodeAct+SFT. Read that as a claim about the trajectory data ReCode generates being more learnable, not as evidence that the recursive loop is cheaper at inference time. Nothing in the README reports token counts, wall-clock time or cost per task, even though the async client can track costs.

## Where ReCode is the wrong tool

The repository is a paper reference implementation. There are no releases, and the install path is a requirements.txt plus three separately configured benchmark environments. If you want an agent library with versioned packages, ReCode is not that. The constrained executor is also a real boundary: the agent's flexibility is bounded by the toolset the executor exposes, so tasks whose useful actions are not expressible as Python calls in that environment will not benefit from recursive expansion. The README's own note that the dependency list may be incomplete, and its request to contact the author about difficulties in using or reproducing the code, point to a project maintained by its authors around a paper rather than a support team. A further constraint is that the README does not document rollback or recovery behaviour when an expansion produces invalid code beyond saying that the executor validates code blocks; if your workload needs a defined failure path, that is unspecified here.

## How ReCode differs from ReAct and CodeAct

ReAct interleaves reasoning text with single actions and keeps the plan in prose, so the granularity of a step is whatever the model writes next. CodeAct moves the action space into code, which makes actions composable, but the plan is still emitted as a flat sequence of code blocks. ReCode's difference is the tree: a plan is itself a function in the same code representation, and it stays a placeholder until execution forces it to be resolved. That is why the README describes granularity as universal rather than fixed. AdaPlanner and ADaPT, the other baselines named, also work on adaptive planning, and the README positions ReCode against them on the same three benchmarks rather than on a separate task suite. The comparison that matters in practice is the one the SFT table makes: at a fixed 7B model size, the representation the agent uses during data collection changes how much the fine-tuned model learns.

## Maintenance, licensing and the cost of upgrading

The repository is MIT licensed, which permits commercial use and modification provided the licence and copyright notice are retained; that is the general shape of MIT, and anything beyond it is a question for your own counsel. The last push to the default branch was on 2026-04-21, and no releases have been published, so there is no version number to pin and no changelog to read before upgrading. In practice, upgrading means pulling main and re-reading the README, because the install steps and the profiles format are the only contract the project states. The dependencies are pinned in requirements.txt (openai==2.6.1, rich==14.2.0, torch==2.9.0), and the three environment wrappers carry their own external setup, so a change in any of those upstream projects can break a working checkout without a ReCode commit. Budget for that: the conda environments are the expensive part to rebuild, and the README's suggestion to keep them separate means three of them.

## Conclusion

Adopt ReCode if you are reproducing or extending the paper's recursive plan-and-action loop and are willing to wire up ALFWorld, WebShop or ScienceWorld assets yourself; the README states the three environments are best configured in separate conda environments because conflicts in one environment have not been confirmed either way. Do not adopt it as a general production agent runtime: there are no releases, no packaging, and the README itself asks you to contact the author if dependencies turn out to be incomplete. Before committing, verify that configs/profiles.yaml resolves a working API credential and that a single-instance dry run on your chosen environment writes a run summary under logs/<run_id>/. If the tree expansion and shared executor state do not match how your task decomposes, stop there.

## FAQ

### What is FoundationAgents/ReCode?

It is the reference implementation for a paper on recursive code generation in LLM agents, where high-level plans are placeholder functions that expand into executable primitives. The repository includes the agent, prompt templates, environment wrappers for ALFWorld, WebShop and ScienceWorld, and the CLI entry point run.py.

### Which Python version does ReCode require?

The README states Python 3.10 or newer, created through conda. It also suggests configuring the three benchmark environments separately because it has not been confirmed whether their dependencies conflict in one environment.

### How do I point ReCode at an LLM API key?

Copy configs/profiles_example.yaml to configs/profiles.yaml and fill in a named profile with api_key, base_url, model and temperature, then select it with the run.py --profile flag. As a fallback, the README says the default profile will use the OPENAI_API_KEY environment variable if the file is omitted.

### Does ReCode track how much a run costs?

Yes, if track_costs is enabled for the profile, in which case it loads pricing metadata from configs/prices.json. Setting track_costs to false disables cost recording for that profile.

### Where do run logs go?

The README states that logs are written to logs/<run_id>/, and that the console prints a condensed summary for quick diagnostics.

## Sources

- [FoundationAgents/ReCode on GitHub](https://github.com/FoundationAgents/ReCode)
- [Issues](https://github.com/FoundationAgents/ReCode/issues)
- [License: MIT](https://github.com/FoundationAgents/ReCode/blob/main/LICENSE)
- [README](https://github.com/FoundationAgents/ReCode/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/foundationagents-recode
