# Flow-GRPO's fix for over-optimization starts from an importance ratio biased below one

> Flow-GRPO is a research implementation for training flow matching models with online reinforcement learning. It supports SD3.5-M, FLUX, Qwen-Image, Bagel-7B and Wan2.1, ships three checkpoints on Hugging Face, has no GitHub release at all, and still calls its package version 0.0.1.

**yifan123/flow_grpo** — [NeurIPS 2025] An official implementation of Flow-GRPO: Training Flow Matching Models via Online RL

- Repository: https://github.com/yifan123/flow_grpo
- Website: https://arxiv.org/pdf/2505.05470
- Stars: 2,542 · Forks: 170
- Language: Python
- License: MIT
- Published: 2026-09-28 · Updated: 2026-09-28 · Language: en
- Canonical page: https://hysenlabs.com/projects/yifan123-flow-grpo

## The over-optimization story starts with a ratio whose mean is below one

The diagnosis the project publishes is specific. The importance ratio in this policy update is not centred where theory assumes: its mean sits consistently below 1, and the gap widens at low noise steps, with step 8 in SD3.5-M given as the example. The variance also moves from one denoising step to the next. Clipping is supposed to fix overconfident samples by truncating anything outside the region from 1 minus epsilon to 1 plus epsilon, which only works if the distribution has mean 1 and a stable spread. Under the observed bias that mechanism fails on the positive side: gradients for positive samples stop being constrained, the policy pushes past the reward it is chasing, and the symptom is a proxy score that keeps climbing while the gold score falls, with image quality degrading along with it.

That failure has a name in this repository: implicit over-optimization in flow matching. The response is GRPO-Guard, added on 2025-11-04 and documented with its own paper and project page.

## GRPO-Guard corrects the distribution first, then reweights the steps

Two mechanisms are stacked, and the order matters. RatioNorm handles the distributional problem: it corrects the bias in the importance ratios and unifies their statistics across denoising steps, so the mean and spread stop depending on where in the trajectory you are. Gradient Reweight then works on top of that output, reweighting the gradients of different denoising steps according to RatioNorm so that no single noise level dominates the update. The stated effect is narrow and checkable: the proxy score still rises the way it does without the guard, but the gold score no longer collapses.

The training entry point takes the node rank as an argument, zero for the machine that coordinates and one for each of the others:

```bash
# Master node
bash scripts/multi_node/sd3_grpo_guard.sh 0
# Other nodes
bash scripts/multi_node/sd3_grpo_guard.sh 1
```

Note what the README does not attach to this: a link to GRPO-Guard model weights or a base checkpoint to download. It says only that after you download the base model and set up the reward model, you run this script for the SD3.5-M text rendering task.

## The table comparing FlowGRPO with GRPO-Guard argues with itself

The comparison table for the guard is broken in a way that hides its own point. The header names two columns, FlowGRPO and GRPO-Guard. The first data row is empty in both. The only populated row repeats the same sentence in both columns:

| FlowGRPO | GRPO-Guard|
| - | - |
|  |   |
| The clipping mechanism is imbalanced, failing to constrain overconfident positive samples. | The clipping mechanism is imbalanced, failing to constrain overconfident positive samples.|

A table built to contrast the unfixed method with the fixed one therefore states the failure twice and shows nothing about the fix. The same section opens with the sentence To mitigates implicit over-optimization in flow matching, our team propose, which carries a stray third person verb. Neither problem touches the code, but both sit in the section a reader reaches first when deciding whether the guard is worth the detour, and the prose around the figures is where the actual argument lives.

## Training speed comes from dropping CFG and training part of the trajectory

Three adjustments are offered for throughput. The first is to run no CFG during training or testing at all, on the reasoning that the reinforcement learning step ends up distilling CFG into the policy instead of paying for a second forward pass per step. The second is to use the window mechanism from Flow-GRPO-Fast or from MixGRPO, which trains on only some of the denoising steps rather than all of them. The third is Coefficients-Preserving Sampling, with noise_level = 0.8 given as a setting that works without retuning per model or step count, and reported to improve GenEval while producing higher quality samples.

The curves backing this are named down to line numbers: geneval_sd3_fast_nocfg and pickscore_sd3_fast_nocfg in config/grpo.py, at lines 163 and 323, driven by scripts from scripts/multi_node/sd3_fast. Both training and evaluation in that comparison run without CFG, which is the condition the first tip recommends, so the three tips and the evidence are consistent with each other.

## Fast variant injects noise at one random step and keeps the rest deterministic

Flow-GRPO-Fast works by shrinking the stochastic part of a trajectory rather than by changing the objective. For each prompt it first generates a deterministic trajectory with ODE sampling. At a randomly chosen intermediate step it injects noise and switches to SDE sampling to produce the group that the reinforcement learning update needs. Everything after that point returns to ODE sampling. The result confines stochasticity to a single step, or two, out of the full trajectory.

The changelog records how that turned into a maintained variant rather than an experiment. On 2025-07-31 Flow-GRPO-Fast was added. On 2025-08-14 a reward curve was published comparing it against Flow-GRPO, and the claim is that with the PickScore reward the fast variant is comparable to the full one after only 2 steps of training. On 2025-10-14 it was refactored for compatibility with Flow-GRPO, which brought Coefficients-Preserving Sampling and no-CFG training on SD3 at the same time. One sampling knob is documented in the same run of entries: config.sample.same_latent, added 2025-07-28 to control whether identical prompts reuse the same noise, which is the fix the changelog attributes to Issue #7.

## The multi-GPU example launches exactly one process

The one training command printed in full is for Wan2.1, contributed on 2025-08-15:

```bash
accelerate launch --config_file scripts/accelerate_configs/multi_gpu.yaml --num_processes=1 --main_process_port 29503 scripts/train_wan2_1.py --config config/grpo.py:general_ocr_wan2_1
```

The config file it points at is named multi_gpu.yaml, and the launch passes one process. Whatever that configuration does with the GPU count, the example as written trains on a single process, and the README gives no command showing more. There is also a second convention for multi node work in the same document: the GRPO-Guard scripts take a node rank as a positional argument, 0 for the master and 1 for the others, with no host list, no rendezvous address and no indication of how machines are discovered. Two different mechanisms for scaling out, neither fully specified on the page.

## Thirty pinned packages, one commented out, and a version that never moved

The packaging metadata is the least maintained corner of the repository. The distribution is named flow-grpo, the version is 0.0.1, and Python 3.10 is the floor:

```python
setup(
    name="flow-grpo",
    version="0.0.1",
    packages=find_packages(),
    python_requires=">=3.10",
    install_requires=[
        "torch==2.6.0",
        "torchvision==0.21.0",
        "torchaudio",
        "transformers==4.40.0",
        "accelerate==1.4.0",
        "diffusers==0.33.1",
```

Most entries carry exact pins, including torch 2.6.0, diffusers 0.33.1 and transformers 4.40.0, but a handful float: torchaudio, xformers, absl-py, ml_collections, sentencepiece and openai are unpinned, which means an install resolves whatever is current for those. Web serving is not optional either, since fastapi, uvicorn and aiohttp sit in install_requires while the dev extras hold only ipython, black and pytest. flash-attn is present as a comment, so a machine that needs it has to edit the file. And with no GitHub release published at any point, the version string has had nowhere to go: three Hugging Face checkpoints for GenEval, text rendering and PickScore alignment exist, but the package itself is still numbered 0.0.1.

## Editing support arrives runnable, with 800 samples and an open question

FLUX.1-Kontext-dev support was added on 2025-08-04 with an unusual admission attached. The counting task uses the GenEval reward to detect object counts, and CLIP feature similarity to keep the edit consistent with the original image. The entry then states that the implementation offers a runnable pipeline, but that the training set contains only 800 samples, and that making Flow-GRPO genuinely effective for editing tasks still requires further exploration by the community.

That is the honest framing, and it sits inside a project whose base model coverage has grown quickly: FLUX.1-dev on 2025-07-28, Qwen-Image and Qwen-Image-Edit on 2025-08-15, Wan2.1 on 2025-08-15, and Bagel-7B on 2025-11-04. Breadth across models has grown faster than evidence per task. The changelog also repeats a single day twice, giving 2025-11-04 to both the guard and the Bagel-7B support, which is a small hint at how these entries were assembled.

## Conclusion

Flow-GRPO is worth reading if you are training flow matching models with reinforcement learning, mostly because it explains its own failure modes rather than hiding them: it names the importance ratio bias, admits the editing task has 800 training samples, and ships GRPO-Guard to counter over-optimization. What to verify before spending GPU hours is the ground you cannot see from the README. The reproduction weights for GRPO-Guard are not linked from this page, the multi-GPU example launches a single process, the multi-node scripts take a node rank as a bare argument with no host list, and the dependency set is pinned hard enough that one comment line, flash-attn, has to be switched on by hand. Check whether your base model, reward model and step budget match the configurations named in the training speed section before assuming the numbers transfer.

## FAQ

### flow grpo vs dance grpo

Nothing in this repository mentions DANCE-GRPO. What the changelog records is coverage for SD3.5-M, FLUX.1-dev, FLUX.1-Kontext-dev, Qwen-Image, Qwen-Image-Edit, Bagel-7B and Wan2.1, plus the Flow-GRPO-Fast variant, Coefficients-Preserving Sampling, CLIPScore and PickScore as reward models, and GRPO-Guard against over-optimization.

### What does GRPO stand for in AI?

The README uses GRPO, GRPO-Guard, GRPO-Fast and Flow-GRPO throughout and never expands the acronym, so the repository is not the place to look it up. It does name the concrete mechanisms around it: RatioNorm and Gradient Reweight inside GRPO-Guard, and Coefficients-Preserving Sampling with a noise_level of 0.8 as a typical setting.

### Which base models does Flow-GRPO support?

The changelog records FLUX.1-dev, FLUX.1-Kontext-dev, Qwen-Image, Qwen-Image-Edit, Bagel-7B and Wan2.1, with SD3.5-M used for the text rendering task, the GenEval task and the PickScore task. Three fine-tuned checkpoints are published on Hugging Face, one per task.

### How does Flow-GRPO-Fast differ from Flow-GRPO?

Flow-GRPO-Fast trains on only one or two denoising steps per trajectory. Each prompt gets a deterministic ODE trajectory, noise is injected at one randomly chosen intermediate step, SDE sampling produces the group there, and the rest continues with ODE sampling. The changelog reports it comparable to Flow-GRPO on PickScore after 2 steps of training.

### What is GRPO-Guard and when was it added to Flow-GRPO?

GRPO-Guard is the project's answer to implicit over-optimization in flow matching, added on 2025-11-04 with its own paper and project page. It adds RatioNorm to correct the biased importance ratio distribution across denoising steps, and Gradient Reweight to balance step contributions on top of it.

### Is there a packaged release of Flow-GRPO?

No GitHub release has ever been published for the repository, and setup.py still declares version 0.0.1 with Python 3.10 as the minimum. The only downloadable artifacts linked are the three Hugging Face checkpoints for GenEval, text rendering and PickScore alignment, and the training scripts are expected to be run from a clone.

## Sources

- [Issues](https://github.com/yifan123/flow_grpo/issues)
- [License: MIT](https://github.com/yifan123/flow_grpo/blob/main/LICENSE)
- [Project website](https://arxiv.org/pdf/2505.05470)
- [README](https://github.com/yifan123/flow_grpo/blob/main/README.md)
- [yifan123/flow_grpo on GitHub](https://github.com/yifan123/flow_grpo)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/yifan123-flow-grpo
