# FireRed-Image-Edit: an open image editing model with a LoRA training path

> FireRed-Image-Edit is a Python image editing foundation model from FireRedTeam, released under Apache-2.0 with weights on Hugging Face and ModelScope. Its selling points are identity consistency, multi-element fusion and a released LoRA training pipeline, but the README documents the training route far better than the inference route.

**FireRedTeam/FireRed-Image-Edit** — FireRed-Image-Edit is a powerful image editing foundation model achieving open-source state-of-the-art performance with precise instruction following, high-fidelity generation, superior identity consistency, and seamless multi-element fusion.

- Repository: https://github.com/FireRedTeam/FireRed-Image-Edit
- Stars: 1,375 · Forks: 81
- Language: Python
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/fireredteam-firered-image-edit

## What FireRed-Image-Edit is for, and who should care

Instruction-driven image editing means you hand the model a picture plus a sentence and get an edited picture back. FireRed-Image-Edit is a foundation model for that job. The README describes FireRed-Image-Edit-1.0 as a general-purpose editing model and 1.1 as an update built on top of 1.0 that targets portrait consistency, multi-element fusion, stylized text reference and portrait makeup effects. The repository is Python and the licence is Apache-2.0.

The audience is narrower than the feature list suggests. If you want to type a prompt into a web page, the README points at the Hugging Face Space and the ModelScope studio, and you never touch this code. The repository matters to people who need to run the model locally, wire it into a ComfyUI graph, or train a LoRA on their own style. The topics list (aigc, diffusion-models, image2image, pytorch) and the presence of train/, agent/ and rededit_bench/ directories confirm that this is aimed at people who intend to do engineering work, not just consume a demo.

## How the editing pipeline is put together

The README makes one architectural claim that is worth taking seriously: editing capability is injected through a Pretrain to SFT to RL pipeline and is described as backbone-agnostic, meaning it can be transferred to any text-to-image foundation model. That is a claim about training methodology, and the repository layout backs the existence of the training side with a train/ directory. What the README does not give is a layer-by-layer diagram of the inference model, so anyone integrating it will be reading inference.py rather than a spec.

Two mechanisms are named concretely. First, multi-element fusion is described as combining 10 or more elements with agent-powered automatic cropping and stitching, which explains the agent/ directory. Second, the training methodology pre-extracts features offline so that VLM inference is decoupled from the training loop. That second point matters more than it reads: if the vision-language model is not in the loop during training, your GPU is not spending time on caption generation, and the README frames this as removing generation overhead to speed convergence.

The optimization stack is the other half of the design. The README lists distillation, quantization, static compilation and db_cache, and says the combination reaches roughly 4.5 seconds per sample at 30GB VRAM. Those numbers come from the project's own announcement, not from an independent run, so treat them as the vendor's figure. The requirements.txt hints at which pieces are real: cache_dit and optimum-quanto are listed, which corresponds to caching and quantization respectively.

## Installing FireRed-Image-Edit and running a first edit

The README does not contain a pip install section. The repository ships a requirements.txt at the top level, and the listed dependencies are torch, Pillow, diffusers, accelerate, google-generativeai, tqdm, cache_dit, optimum-quanto and openai. The last entry carries an inline comment saying it is optional and required only for MiniMax or OpenAI-compatible recaption providers. Installing from that file is the only dependency step the repository documents.

```bash
pip install -r requirements.txt
```

Weights are not in the repository. The README links to Hugging Face and ModelScope for both FireRed-Image-Edit-1.0 and FireRed-Image-Edit-1.1, plus a separate FireRed-Image-Edit-LoRA-Zoo repository. You have to fetch the checkpoint from one of those hosts before inference will work, and the README does not spell out the download command, so follow the model card on whichever host you choose.

The one inference command the README does give is the optimized path, announced on 2026.03.01. The news entry says it covers a distilled LoRA, quantization, db_cache and static compilation, and that it needs only 30GB VRAM at roughly 4.5s per sample.

```bash
python inference.py --optimized True
```

That single flag is the whole documented interface. The README does not list other arguments, does not show where input images are read from or where output is written, and does not give a sample prompt. Before running it, open inference.py and read the argument parser, because the flag name is the only thing the README commits to. The examples/ directory contains reference outputs (edit_ours.png, edit_qwen2511.png, edit_seedream4.png, edit_nano_pro.png and others), which is the closest thing to an expected-result baseline in the repository.

## Where the documentation leaves you on your own

The README is a feature announcement more than an operating manual. Several practical questions have no answer in it. There is no stated VRAM requirement for the non-optimized path, only the 30GB figure attached to the optimized one, so a smaller card is an unknown until you test it. There is no list of supported input formats or resolutions. There is no documented rollback procedure or checkpoint versioning scheme, which matters if you plan to pin a model version in production and later move to 1.1.

The 1.0 to 1.1 relationship is also a gap. The README says 1.1 optimizes portrait consistency, multi-element fusion, stylized text reference and portrait makeup on top of 1.0, but it does not say whether 1.1 fully supersedes 1.0 or whether some tasks still favour the older checkpoint. Both weight repositories remain linked, which suggests both are intended to be usable, but the README never states a selection rule.

The agent workflow is the most opaque part. Automatic cropping and stitching for multi-element compositions is described as removing the need for long prompt engineering, but the README does not document what the agent does, which model drives it, or what happens when cropping picks the wrong region. If your work depends on precise control over which elements are fused, that is a mechanism you will have to inspect in agent/ rather than configure from documentation.

## ComfyUI, GGUF and the deployment options people ask about

The README claims native ComfyUI node support and GGUF lightweight format compatibility, and the project's own search traffic is dominated by exactly those two topics alongside LoRA. What the README does not provide is a node installation procedure, a workflow JSON, or a link to a ComfyUI custom node repository. The Hugging Face organization does host a FireRed-Image-Edit-1.0-ComfyUI repository containing a Lightning 8-step LoRA file, referenced in the 2026.03.01 news entry, so the ComfyUI path exists as published artifacts even though the README stops short of walking through it.

The same pattern applies to LoRA training. The README states that full training code is released for custom style creation and that a LoRA Zoo exists on Hugging Face. The train/ directory is in the repository and the README notes support for HSDP/FSDP and disaggregated setups, though that sentence is cut off in the README as shown. ModelScope added LoRA training support for FireRed-image-edit on 2026.03.25, which gives a second route that does not require your own training loop. For anyone whose goal is a custom style rather than a general editor, that hosted training option is probably the faster path than standing up the local train/ code.

The honest summary is that the ecosystem pieces are published as artifacts (weights, LoRA files, a benchmark dataset) while the connective documentation lives outside the README. Budget time for reading repository code, not just the front page.

## FireRed-Image-Edit versus Qwen-Image-Edit and other editors

The examples/ directory contains a file named edit_qwen2511.png alongside edit_ours.png, which tells you the project positions itself against Qwen-Image-Edit class models. The difference the README draws is not raw editing quality but the surrounding engineering: a released LoRA training ecosystem, an acceleration suite the project claims brings end-to-end generation to about 4.5 seconds at 30GB VRAM, and an agent layer for multi-image composition. A general-purpose editor release typically ships weights and an inference script; FireRed-Image-Edit additionally ships train/, agent/ and rededit_bench/.

That extra surface is the real differentiator and also the real cost. If you only want to edit a photo with a sentence, a model with a simpler, better-documented inference path will get you there with less reading. The FireRed-Image-Edit argument is strongest when you intend to fine-tune, when you need the identity-consistency behaviour for repeated characters, or when you want the multi-element fusion to be handled by the agent rather than by a very long prompt. The project also published REDEdit-Bench on 2026.03.09, described as covering more diverse scenarios and instructions that align better with human language, which gives you a benchmark to run rather than a claim to trust. Running that benchmark on your own hardware is the most defensible way to compare it against whatever editor you use now.

## Licence, maintenance and the cost of upgrading

The repository is Apache-2.0, the same licence the README badge points to. That is a permissive licence and it is a meaningful difference from editing models released under custom or non-commercial terms, but it covers the code and the repository's own licence file. The model weights live on Hugging Face and ModelScope as separate artifacts, and the README does not restate licence terms for those checkpoints, so check the model card for each weight repository you download. This is a factual observation about where the licence text sits, not legal advice.

The last push to the repository was on 2026-04-03. That is roughly five and a half months before today, so this is a project with recent but not current activity, and it should not be described as under active development on the strength of that date alone. The news section in the README ends with entries from March 2026, the most recent being the ModelScope LoRA training support on 2026.03.25.

Upgrade cost is the practical worry. There are no retrieved releases, so versioning happens through the two checkpoint names (1.0 and 1.1) rather than through tagged releases you can diff. The README does not describe a migration path between them or state whether 1.1 changes inference arguments. If you pin a checkpoint today, plan to re-read inference.py and the model card before moving, because the README will not tell you what changed at the code level.

## Conclusion

Adopt FireRed-Image-Edit if you already run a GPU box with roughly 30GB of VRAM, you want an Apache-2.0 editing model you can fine-tune with the released train/ code, and you are comfortable reading inference.py and the docs/ directory because the README stops short of a full walkthrough. Do not adopt it if you need a hosted API with a published rate limit, a documented VRAM floor for the non-optimized path, or a rollback and versioning story, because none of those appear in the README. Before you commit, check the two weight repositories (FireRed-Image-Edit-1.0 and FireRed-Image-Edit-1.1) and confirm which one your pipeline should load, then run the optimized inference script on a single example image and compare the output against examples/edit_ours.png.

## FAQ

### What is FireRed-Image-Edit?

It is an image editing foundation model from FireRedTeam, written in Python and released under Apache-2.0. The README describes version 1.0 as a general-purpose editing model and 1.1 as an update that improves portrait consistency, multi-element fusion, stylized text reference and portrait makeup effects.

### Is FireRed-Image-Edit good?

The README claims leading open-source results for instruction following, image quality and visual coherence, and the project publishes REDEdit-Bench as an evaluation set. Those are the project's own claims; the repository does not include independent benchmark results, so the defensible way to judge it is to run REDEdit-Bench or your own images.

### Is FireRed-Image-Edit censored?

The README does not mention content filtering, safety classifiers or any moderation layer, so there is nothing in the available material that answers this. Note that the README does list google-generativeai and openai as dependencies, with the openai entry marked optional and tied to recaption providers.

### Which ComfyUI tool is best for image editing?

The README claims native ComfyUI node support and GGUF lightweight format compatibility for FireRed-Image-Edit, and a FireRed-Image-Edit-1.0-ComfyUI repository holds a Lightning 8-step LoRA file. It does not compare ComfyUI nodes against each other, so it cannot answer which tool is best.

### What are the best open-source image edit models?

The README describes FireRed-Image-Edit as delivering leading open-source results with accurate instruction following and high image quality, and the examples/ directory includes an edit_qwen2511.png reference alongside edit_ours.png. The repository does not rank other open-source editors, so it gives no basis for a general comparison.

## Sources

- [FireRedTeam/FireRed-Image-Edit on GitHub](https://github.com/FireRedTeam/FireRed-Image-Edit)
- [Issues](https://github.com/FireRedTeam/FireRed-Image-Edit/issues)
- [License: Apache-2.0](https://github.com/FireRedTeam/FireRed-Image-Edit/blob/main/LICENSE)
- [README](https://github.com/FireRedTeam/FireRed-Image-Edit/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/fireredteam-firered-image-edit
