Open-source project
huggingface/Repo2RLEnv avatar
huggingface/Repo2RLEnv

Repo2RLEnv: turning repositories into Harbor RL tasks

Convert any Repo into an RL Environment

680 stars102 forksPythonApache-2.0

At a glance

What is it?
Repo2RLEnv generates coding, terminal and reasoning tasks in the Harbor format from repositories, pull requests and task seeds. It is a Python 3.12+ CLI aimed at people building training or evaluation data, and its own classifiers still mark it pre-alpha.
Who is it for?
Adopt Repo2RLEnv if you already work in the Harbor format and want PR-derived tasks without writing your own extraction pipeline, starting with the stable pr_diff and pr_runtime pipelines on Linux or macOS. Do not adopt it if you need Windows-native Tasksmith or research-recipe runs, if you have no sandbox provider budget, or if you need a task format other than Harbor.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository received new commits within the last day.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The gap Repo2RLEnv fills between a Git history and a runnable task

A repository is a record of work that already happened. An RL environment is a prompt, a starting state, a reference solution and a program that decides whether the attempt was correct. The distance between the two is normally covered by hand: someone picks a pull request, writes an instruction, builds a container, and writes a verifier. Repo2RLEnv automates that conversion. The README describes the output as a Harbor task with an instruction, a starting environment, a reference solution and an executable verifier. That four-part shape is the product. If you only need a prompt and an answer, this tool is heavier than the problem.

The intended audience is narrow but real: people producing coding, terminal and reasoning data for training or evaluation, who already accept the Harbor task format. The repository ships 14 experimental recipes that adapt published methods into code owned by the repository, and the README states that no upstream research package is installed at runtime. That matters for reproducibility. A recipe named after a paper is not the paper's code; it is this project's reimplementation, and the licence files listed in pyproject.toml show that upstream terms travel with several of them.

How a repository becomes a task: pipelines, recipes and the quality loop

The generation path is a funnel. A source (repository, pull request or task seed) enters one of three routes: native pipelines, Tasksmith, or a research recipe. A pipeline names the generation family; a recipe selects its method. The native pipelines are the most legible part of the design. pr_diff asks the model to reproduce a real PR change and scores the result by diff similarity plus an optional LLM judge. pr_runtime turns a PR regression into a test-based task where failing tests must pass and existing tests must stay green. commit_runtime builds the same kind of task from commit history. The README marks those first three stable and the rest experimental.

Tasksmith is a different mechanism. The README states that LangGraph orchestrates Pi or OpenCode to investigate a merged PR, establish a working environment, and design the instruction and a private verifier, with the merged implementation supplying the oracle. It targets testable Python changes and is experimental. Research recipes are the third route, and they are the least predictable: 14 of them, each mapping a published method onto a pipeline family such as repo_mutate, pr_to_env, terminal_synth or env_repair.

What separates this from a prompt generator is the verification stage. After emission, tasks pass static checks, a baseline and oracle run, review and learner rollouts, with bounded repair feeding back into the loop. The README is explicit that exporting a task alone does not establish its quality. That sentence is the most useful thing in the documentation, because it tells you the tool will happily produce tasks that have not been verified.

Installing Repo2RLEnv and generating your first PR-diff tasks

The package requires Python 3.12+ and Git. The README gives a PyPI install and a worked example that generates PR-diff tasks without building a container, which is the cheapest way to see whether the output shape suits you. Authentication for cloning and the GitHub API resolves in a documented order: `gh auth token` first, then a GITHUB_TOKEN environment variable.

bash
pip install repo2rlenv

# GitHub access; alternatively set GITHUB_TOKEN
gh auth login

repo2rlenv generate \
  --repo pallets/click --pipeline pr_diff \
  --pipeline-opt limit=3 --out ./workspace/click-tasks

repo2rlenv validate ./workspace/click-tasks --deep
repo2rlenv pipelines list

The generate command writes tasks to the directory given by --out; --pipeline-opt limit=3 caps the run at three tasks, which is what you want on a first attempt. validate --deep then runs the deeper checks over the emitted directory, and pipelines list shows what is available on your install. If you want test-based repository tasks rather than diff reproduction, switch to pr_runtime or Tasksmith.

LLM stages need a provider key. The .env.example template lists ANTHROPIC_API_KEY, OPENAI_API_KEY and HF_TOKEN, and notes that every variable is optional and that the CLI resolves credentials from the environment without a flag. Extras are installed per route.

bash
pip install 'repo2rlenv[tasksmith,daytona,harbor]'
# Other extras: modal, mutation
repo2rlenv tasksmith install-runtime  # Pi / OpenCode; requires Node.js 22.19+

The template file is copied and sourced before running commands, and the example it gives is `cp .env.example .env` followed by `set -a && source .env && set +a`. The HF_TOKEN variable does double duty: model inference through Hugging Face providers and dataset publishing to the Hub.

Where Repo2RLEnv breaks: Windows, sandboxes and unverified output

The platform boundary is stated plainly in the README. Windows CI covers CLI startup, recipe discovery, native task emission and static validation, but the README instructs users to run Tasksmith, research-recipe generation and the quality controller on Linux, macOS or WSL because full native Windows execution is not yet supported. A Windows user can therefore install the package and emit tasks, then hit a wall the moment they want the quality loop. That is a hard split, not a rough edge.

The runtime requirement is the second constraint. Native runtime pipelines build and cache Docker environments locally. Tasksmith and the owned recipes build and execute remotely through Daytona or Modal, and the README notes that the implemented GPU route uses Modal L4 GPUs. So the more advanced routes carry a third-party dependency and a bill. The Tasksmith walkthrough is described as including an explicit spending limit, which implies runs can cost enough to need one.

Third, the project's own metadata is cautious. pyproject.toml carries the classifier Development Status :: 2 - Pre-Alpha, and the version is 0.9.1. The README's own warning that exporting a task does not establish its quality means the default output is a candidate, not a dataset. Anyone who treats a generated directory as ground truth is misusing the tool. The recipes add a quieter risk: they are reimplementations of published methods with upstream licence files shipped alongside, so behaviour may diverge from the papers they are named after.

Repo2RLEnv compared with SWE-bench-style harnesses and hand-built task sets

The closest thing to an alternative is the pattern this project descends from: a fixed benchmark harness such as the SWE-bench style setup, where tasks are curated once, frozen, and scored by a known test command. The difference is in where the work happens. A frozen harness gives you a stable, comparable set and no generation step; you cannot point it at a new repository on Monday and have tasks on Tuesday. Repo2RLEnv inverts that. It trades comparability for throughput, and it accepts that each generated task needs its own verification before it is worth anything.

The second alternative is building the pipeline yourself, which is what many teams do: clone, diff, prompt, containerize, test. That gives you full control over the verifier and no dependency on someone else's recipe semantics. The cost is that you reimplement the Harbor task shape, the credential resolution and the repair loop. Repo2RLEnv's argument is that those parts are boring and shared. Its counterargument against itself is the pre-alpha classifier: you are adopting a moving target in exchange for not writing the plumbing.

A third comparison is worth naming because the README makes it. The research recipes deliberately do not install upstream packages at runtime, so you are not comparing against running the original research code. You are comparing against a reimplementation with its own licence obligations, which is a different trade than it first appears.

Maintenance, releases and what the Apache-2.0 AND MIT licence actually covers

The repository is not archived, and the last push was on 2026-09-17. Recent releases are close together: v0.9.0, described as Tasksmith and owned generation recipes, and v0.9.1, a Windows CLI fix, both dated 2026-09-15, following v0.8.8 on 2026-07-22. The cadence suggests active work, and the 0.9.x line means interfaces can still move between releases. Pinning a version is the obvious mitigation, and pyproject.toml plus uv.lock in the repository root give you what you need to do it.

The licence situation needs care rather than legal advice. pyproject.toml declares `license = "Apache-2.0 AND MIT"` and lists licence files that include LICENSE, THIRD_PARTY_NOTICES.md, and one UPSTREAM_LICENSE per recipe for cli_gym, endless_terminals, r2e, r2e_gym, scaler, seta_evol, seta_seed2synth, swe_flow, swe_gen, swe_next, swe_smith, terminalworld and tmax, plus an UPSTREAM_NOTICE for scaler. The badge in the README says Apache-2.0 and MIT. The practical reading is that the top-level code is Apache-2.0, and that individual recipes carry upstream terms you should read before redistributing generated tasks. Which obligations attach to generated output is not something the README answers, and that is the question to put to whoever owns licensing on your side.

Editorial conclusion

Adopt Repo2RLEnv if you already work in the Harbor format and want PR-derived tasks without writing your own extraction pipeline, starting with the stable pr_diff and pr_runtime pipelines on Linux or macOS. Do not adopt it if you need Windows-native Tasksmith or research-recipe runs, if you have no sandbox provider budget, or if you need a task format other than Harbor. Before committing, run repo2rlenv validate --deep on a small batch and check the licence files that pyproject.toml declares for the recipes you intend to use.

Frequently asked questions

What Python version does Repo2RLEnv need?

The README states the requirement as Python 3.12+ and Git, and pyproject.toml sets requires-python to >=3.12 with classifiers for 3.12, 3.13 and 3.14.

Can I run Repo2RLEnv on Windows?

Partly. Windows CI covers CLI startup, recipe discovery, native task emission and static validation, but the README says to use Linux, macOS or WSL for Tasksmith, research-recipe generation and the quality controller because full native Windows execution is not yet supported.

Do I need an API key to use Repo2RLEnv?

Only for the stages you actually run. The .env.example template lists ANTHROPIC_API_KEY, OPENAI_API_KEY and HF_TOKEN for LLM stages, and notes that every variable is optional and that the CLI auto-resolves credentials from the environment without a flag.

Does Repo2RLEnv need Docker or another sandbox?

Native runtime pipelines build and cache Docker environments, while the README states that owned recipes and Tasksmith build and execute remotely through Daytona or Modal, with the implemented GPU route using Modal L4 GPUs.

What task format does Repo2RLEnv produce?

Harbor format: a task instruction, a starting environment, a reference solution and an executable verifier, which the README describes as runnable with Harbor agents and publishable to the Hugging Face Hub.

Official sources

  1. huggingface/Repo2RLEnv on GitHub
  2. License: Apache-2.0
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/huggingface-repo2rlenv.svg)](https://hysenlabs.com/projects/huggingface-repo2rlenv)