# Step1X-Edit: a reasoning image editing model where the thinking happens before the edit

> StepFun's open weights editor adds an explicit reasoning pass, ships the benchmark scores that justify it, and pins its own transformers version twice in one requirements file. Both facts matter.

**stepfun-ai/Step1X-Edit** — A SOTA open-source image editing model, which aims to provide comparable performance against the closed-source models like GPT-4o and Gemini 2 Flash.

- Repository: https://github.com/stepfun-ai/Step1X-Edit
- Website: https://step1x-edit.github.io
- Stars: 2,267 · Forks: 106
- Language: Python
- License: Apache-2.0
- Published: 2026-10-07 · Updated: 2026-10-07 · Language: en
- Canonical page: https://hysenlabs.com/projects/stepfun-ai-step1x-edit

## What Step1X-Edit is and how it differs from a plain editor

Step1X-Edit is an open source image editing model from StepFun, released under Apache-2.0. The repository description states the ambition plainly: a state of the art open source image editing model aiming for performance comparable to closed models such as GPT-4o and Gemini 2 Flash. The repository topics are image-editing, reasoning and visual-reasoning, which is the right summary of the design.

The distinguishing idea is that editing is treated as a reasoning problem rather than a transformation problem. Most editors map an instruction plus an image straight to an output image. Step1X-Edit inserts an explicit reasoning step, and the README describes the preview checkpoint released on 2025-09-08 as a native reasoning edit model that combines instruction reasoning with reflective correction. The vocabulary in the benchmark tables reflects that: a checkpoint appears in variants labelled base, thinking, and thinking plus reflection, so you can see the model's internal deliberation treated as a switchable component.

Two papers back the work. The project page is at step1x-edit.github.io, the first report is arXiv 2504.17761, and a second report, 2511.22625, covers the ReasonEdit line. Weights are on Hugging Face at stepfun-ai/Step1X-Edit, with separate repositories for the v1p2 and v1p2-preview checkpoints. The benchmark dataset GEdit-Bench is published too, at the stepfun-ai org. Publishing the evaluation set alongside the model is a rarer courtesy than it should be, and it means the numbers below can be re-derived rather than taken on faith.

## The benchmark tables the project publishes about itself

The README carries evaluation tables for two suites, GEdit-Bench and KRIS-Bench, and it is worth reading them closely because they are the project's own argument. On KRIS-Bench, whose overall column aggregates factual, conceptual and procedural knowledge, Flux-Kontext-dev scores 49.54 and Qwen-Image-Edit-2509 scores 56.15. Step1X-Edit v1.1 scores 51.59, which places it above Flux on this suite and below Qwen.

Then the reasoning variants appear. The v1p2 base checkpoint scores 56.33, thinking alone reaches 58.64, and thinking plus reflection reaches 60.93. That progression is the clearest evidence in the repository for the central claim, since it isolates the contribution of the reasoning pass while holding the model fixed. Note also that v1p2 base at 56.33 edges past Qwen-Image-Edit-2509 at 56.15 before any reasoning is involved, so part of the gain comes from the checkpoint and part from the deliberation.

On GEdit-Bench, scored with both GPT-4.1 and Qwen2.5-VL-72B as judges, v1.1 reaches 7.66 for semantic consistency and the v1p2 thinking-plus-reflection variant reaches 8.18. One methodological detail deserves credit: the maintainers published intermediate evaluation results as a separate dataset to support reproducibility, and they name the judge models rather than leaving the evaluation opaque.

The February 2026 entry adds a third-party result worth separating from the first-party ones. RegionE, an external project, delivers a 2.5x inference speedup with no accuracy degradation, which the README attributes to a five-line code change. The 2025-06-17 entry notes native TeaCache and parallel inference support, and `xfuser==0.4.3.post3` appears in the requirements specifically for the parallel path.

## Checkpoint timeline and the naming that will confuse you

Four checkpoints exist and the labels overlap in ways worth mapping before you download anything. v1.0 is the original release. v1.1 followed on 2025-07-09 and is the release that added text-to-image generation alongside editing, so if you want a single model for both jobs, this is the one. v1p2-preview arrived on 2025-09-08 with the reasoning capability. v1p2, released 2025-11-26, is referred to in the paper as ReasonEdit-S and is described as a native reasoning edit model with better performance on both KRIS-Bench and GEdit-Bench.

So the preview and the final v1p2 are separate weights with different scores, and the difference between them is not trivial: the preview's overall KRIS-Bench score is 52.51 while the final v1p2 thinking-plus-reflection variant reaches 60.93. If you are reading a tutorial that references v1p2-preview, that advice is two generations behind.

There is also a diffusers port for v1p1 at Step1X-Edit-v1p1-diffusers, which matters if you would rather not use the project's own inference code. And the most recent entry, dated 2026-04-29, announces Step Image Edit 2 as a lightweight hosted model on the StepFun Open Platform that completes generation and editing within two seconds. That is a different thing from this repository: it is an API product, not weights you run yourself. The repository is not archived, has 34 open issues, 2,264 stars and 104 forks, and was last pushed on 2026-04-29.

## Dependencies, and the version conflict inside requirements.txt

There is no installation section in this README, so the dependency story has to be read out of `requirements.txt`, which is organised into labelled groups. The inference group is the core:

```
torch>=2.3.1
torchvision>=0.18.1
liger_kernel==0.5.4
transformers==4.51.3
gradio==5.29.0
```

Finetuning adds accelerate, diffusers, bitsandbytes, lion-pytorch, schedulefree, pytorch-optimizer and several prodigy optimisers, plus tensorboard and easygui for a training UI. The parallel group is a single line, `xfuser==0.4.3.post3`. The final group is labelled for diffusers and pins `transformers==4.55.0` alongside `peft==0.17.0`.

That is the thing to notice: `requirements.txt` pins transformers to two different exact versions in one file. Installing it top to bottom will not give you a consistent environment. The sensible reading is that the first pin is the original inference stack and the last group was appended later for the diffusers port without revisiting the earlier line. Pick the path you want, install that group's pins, and treat the other as documentation of the other stack rather than as a requirement.

The tree confirms what the entry points are. `inference.py` and `sampling.py` handle generation, `gradio_app.py` provides the interface, `finetuning.py` handles LoRA work, and `modules/` and `library/` hold the model code. The `scripts/` directory holds auxiliary tooling and `GEdit-Bench/` is the evaluation suite checked into the repository. There is an `examples/` folder with sample images and three prompt files, `prompt_cn.json`, `prompt_en.json` and `prompt_fix_hand.json`, which is a useful sign that Chinese-language prompts are a first-class input rather than an afterthought.

## Finetuning, ComfyUI, and the paths around the core repository

The 2025-05-22 entry is the one that changes who can actually use this. Step1X-Edit gained LoRA finetuning on a single 24GB GPU, and a hand-fixing LoRA was released alongside it. Before that, adapting the model was a distributed-training proposition. The finetuning dependency set backs the claim up: bitsandbytes for quantised training, liger_kernel for fused kernels, schedulefree and the prodigy optimisers, and peft for the adapter layer.

The ecosystem integrations are documented as news entries rather than as a supported feature list, and it is worth knowing which are first-party. The 2025-04-30 entry credits two community contributors for ComfyUI plugins, quank123wip and raykindle, with links to both a hyphenated and an underscored repository name. Node-based workflows are therefore available for people who want them, but they are maintained outside this repository and will drift independently of it.

Replicate hosts a hosted version too, so there is a no-GPU route for evaluation before you commit hardware. Between that, the diffusers port, and the Gradio app, the project has covered the three common entry points: a hosted demo for looking, a diffusion-native port for integrating into an existing pipeline, and community ComfyUI nodes for visual assembly.

What the repository does not do is ship an install guide. The README runs through the news timeline and then stops, with the setup instructions apparently living on the project page instead. For a model this size that is a real documentation gap, though the labelled requirement groups and the three entry-point scripts make the shape of the code discoverable without it.

## Conclusion

Step1X-Edit is one of the few open image editing releases where the maintainers publish the numbers that argue against their own earlier checkpoints, which makes it easier to evaluate than most. The v1.1 to v1p2 progression shows the reasoning pass doing real work: on KRIS-Bench, overall accuracy moves from 51.59 to 60.93, and the table separates a plain checkpoint from a thinking checkpoint from thinking plus reflection so you can see which part buys what. Two practical notes before you commit. The requirements file pins transformers to 4.51.3 for inference and 4.55.0 for the diffusers path in the same file, so you pick one path and install deliberately rather than following the file top to bottom. And LoRA finetuning is advertised as fitting a single 24GB GPU, which is a real constraint to design around rather than a limit to discover later. Start from inference.py with the pinned stack, treat the reasoning variants as a separate experiment, and read the GEdit-Bench tables directly rather than trusting any summary of them.

## FAQ

### What is Step1X-Edit and who released it?

Step1X-Edit is an open source image editing model released by StepFun under Apache-2.0, with weights published on Hugging Face at the stepfun-ai organisation. The project describes it as aiming for performance comparable to closed models such as GPT-4o and Gemini 2 Flash, and the repository topics are image-editing, reasoning and visual-reasoning.

### Which Step1X-Edit checkpoint should I download?

v1.1 is the version that added text-to-image generation alongside editing. v1p2-preview from 2025-09-08 introduced reasoning edit capability, and the final v1p2 from 2025-11-26, called ReasonEdit-S in the paper, scores better on both KRIS-Bench and GEdit-Bench than the preview. A diffusers port exists for v1p1 if you prefer not to use the project's own inference code.

### Can Step1X-Edit be finetuned on a single GPU?

Yes. A 2025-05-22 release added LoRA finetuning on a single 24GB GPU, and a hand-fixing LoRA was published at the same time. The finetuning dependency group in requirements.txt includes bitsandbytes, liger_kernel, schedulefree, the prodigy optimisers and peft for the adapter layer.

### How much GPU memory does Step1X-Edit need to run inference?

The requirements file pins torch at 2.3.1 or newer with liger_kernel 0.5.4 and gradio 5.29.0, and parallel inference is supported through xfuser 0.4.3.post3. The README does not state a minimum VRAM figure for inference, but it does state that LoRA finetuning fits a single 24GB GPU, and a third-party project called RegionE reports a 2.5x inference speedup with no accuracy loss.

### Is there a free online version of Step1X-Edit I can try before running it locally?

There is a Replicate-hosted version linked from the README, and the project itself offers Step Image Edit 2 as a hosted API on the StepFun Open Platform that completes generation and editing within two seconds. That second option is a different, lightweight API model rather than the weights in this repository.

## Sources

- [Issues](https://github.com/stepfun-ai/Step1X-Edit/issues)
- [License: Apache-2.0](https://github.com/stepfun-ai/Step1X-Edit/blob/main/LICENSE)
- [Project website](https://step1x-edit.github.io)
- [README](https://github.com/stepfun-ai/Step1X-Edit/blob/main/README.md)
- [stepfun-ai/Step1X-Edit on GitHub](https://github.com/stepfun-ai/Step1X-Edit)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/stepfun-ai-step1x-edit
