Open-source project
VectorSpaceLab/OmniGen2 avatar
VectorSpaceLab/OmniGen2

OmniGen2: Image Generation and Editing, with ComfyUI Support

OmniGen2: Exploration to Advanced Multimodal Generation. https://arxiv.org/abs/2506.18871

4,114 stars36 forksJupyter NotebookApache-2.0

At a glance

What is it?
OmniGen2 is an open-source multimodal model from VectorSpaceLab that covers text-to-image generation, instruction-guided image editing, in-context generation, and visual understanding from one set of model weights. Its dual-pathway architecture separates text and image decoding with unshared parameters, a design that differs from its predecessor OmniGen v1.
Who is it for?
OmniGen2 suits researchers and practitioners who want a single open-source model that covers text-to-image generation, image editing, and visual understanding without juggling separate tools for each task. Its ComfyUI integration and optional CPU offload lower the hardware barrier.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Activity is slowing. The repository last received commits 6 months ago.
What is it written in?
Mainly Jupyter Notebook, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

Four Capabilities, One Dual-Pathway Architecture

OmniGen2 covers four distinct tasks. The first is visual understanding, inherited from the Qwen-VL-2.5 foundation model that backs the text pathway. The second is text-to-image generation, producing images from natural language prompts. The third is instruction-guided image editing, where the model modifies an existing image according to a text instruction. The fourth is in-context generation, which takes diverse inputs such as humans, reference objects, and scenes and combines them into a new coherent image.

The architecture behind all four tasks uses two separate decoding pathways with unshared parameters: one for text and one for images. OmniGen2 also uses a decoupled image tokenizer, which the VectorSpaceLab team identifies as a key departure from how OmniGen v1 was structured. This separation means the model does not have to reconcile two modalities through a single shared set of weights, which the technical report (arxiv 2506.18871) attributes to the approach's performance on instruction-guided editing benchmarks.

The project is also the basis for EditScore, a family of open-source reward models ranging from 7B to 72B parameters designed for evaluating instruction-guided image editing quality. The EditScore release includes OmniGen2-EditScore7B, which enables online reinforcement learning for image editing by providing a scoring signal. LoRA weights for OmniGen2-EditScore7B are available on Hugging Face and ModelScope.

Installing OmniGen2 and Running the Bundled Example Scripts

The repository requires Python 3.11, PyTorch 2.6.0, and CUDA 12.4. The README documents an optional dependency on flash-attn for performance; as of June 2025, OmniGen2 runs without it, though the authors recommend installing it when available.

Clone the repository and set up a clean environment:

bash
git clone [email protected]:VectorSpaceLab/OmniGen2.git
cd OmniGen2
conda create -n omnigen2 python=3.11
conda activate omnigen2

Install PyTorch first with the correct CUDA wheel, then install the remaining dependencies from requirements.txt:

bash
pip install torch==2.6.0 torchvision --extra-index-url https://download.pytorch.org/whl/cu124
pip install -r requirements.txt

The requirements.txt pins key packages including torch==2.6.0, torchvision==0.21.0, transformers==4.51.3, diffusers, accelerate, and timm, among others. Once the environment is ready, the repository ships four shell scripts that exercise each capability:

bash
bash example_understanding.sh
bash example_t2i.sh
bash example_edit.sh
bash example_in_context_generation.sh

Each script runs the corresponding inference path using example images stored in the example_images directory. A Jupyter notebook (example.ipynb) covers the same tasks interactively. The model weights are hosted on Hugging Face under the OmniGen2/OmniGen2 repository; the inference scripts download them automatically on first run.

For machines with limited VRAM, the README documents CPU offload support, added in June 2025. This extends the range of hardware that can run the model at the cost of inference speed.

Faster Inference with TeaCache and TaylorSeer

Two inference acceleration methods are available as opt-in additions. TeaCache, integrated in July 2025 following a community pull request, reduces redundant computation by caching intermediate attention states across steps. TaylorSeer, also added in July 2025, approximates attention outputs using Taylor expansion to skip full computation on less critical steps.

Both methods are covered in the Usage Tips section of the README, which refers users to the respective upstream repositories for the underlying technique. Neither method is enabled by default; they require explicit configuration before running inference. The README notes speed improvements from both but does not give specific benchmark numbers, so expected throughput gains on any particular GPU should be measured independently.

For ComfyUI users, OmniGen2 received official support in the ComfyUI examples gallery in July 2025. This means the model can be loaded and run through the node-based ComfyUI workflow interface without writing Python code, which makes it accessible to users who work primarily in that environment.

Fine-Tuning with the Training Code and OmniContext Benchmark

Training code for OmniGen2 was released in June 2025 and is documented in docs/FINETUNE.md inside the repository. This allows practitioners to fine-tune the model on custom datasets rather than relying solely on the pre-trained weights.

Two training datasets are publicly available on Hugging Face. X2I2 is a training dataset released in July 2025. OmniContext is a benchmark for in-context generation, released in June 2025, with evaluation code in the omnicontext/ directory of the repository. The OmniContext benchmark measures how well a model can perform in-context generation tasks, providing a systematic way to compare approaches on that capability.

The training data construction pipeline was listed as a pending TODO item in the repository's public roadmap at the time of the last push, so assembling a custom dataset in the exact format used for pre-training requires working from the released datasets as a reference rather than following a documented pipeline.

Diffusers integration was also listed as a pending TODO, meaning users who expect to load OmniGen2 through the standard Hugging Face diffusers API cannot do so without additional bridging work.

Limitations: VRAM Demands and Incomplete Roadmap Items

OmniGen2 is not a lightweight model. The README includes a dedicated Resources Requirement section acknowledging that the model demands significant GPU memory. CPU offload is available for devices with limited VRAM, but this trades speed for accessibility. The README does not give a precise minimum VRAM number, so testing on the target hardware before committing to a deployment is necessary.

The in-context generation capability depends on providing coherent reference inputs. The Limitations and Suggestions section of the README notes cases where results degrade, though the specific failure modes are described in that section rather than reproduced here. Instruction-following for complex multi-step edits can also produce unintended changes.

Flash-attn, while optional, affects performance. The version pinned in requirements.txt (2.7.4.post1) is specified for compatibility with CUDA 12.4. Users on CUDA 12.6 or later may be able to use a newer version, per the README note, but this is not guaranteed.

The diffusers integration and the training data construction pipeline remain incomplete per the repository's own TODO list. If either is required for a use case, it represents work the adopter must complete independently.

OmniGen2 vs. Flux Kontext: Two Different Approaches

Flux Kontext is an image generation model from Black Forest Labs that focuses on precise, context-aware image editing through a flow-matching diffusion approach. It operates as a diffusion model with a single task focus on generation and editing rather than as a multimodal system with text understanding built in.

OmniGen2 takes a different position: it is a unified model that inherits a vision-language understanding capability from Qwen-VL-2.5 and extends it with image generation. This means OmniGen2 can answer questions about an image (visual understanding) within the same system that generates or edits images. Flux Kontext does not include a text understanding or question-answering pathway.

The practical difference is scope. If a project requires only high-quality text-to-image generation or targeted editing with a well-tested diffusion pipeline, Flux Kontext or similar dedicated diffusion tools may suit the workflow better. OmniGen2 is the more relevant choice when a single model needs to handle both understanding and generation, or when the instruction-guided editing benchmark results documented in the technical report are specifically relevant to the task.

Maintenance Status and Apache-2.0 Licence

The last push to the VectorSpaceLab/OmniGen2 repository was on 2026-03-20. The repository is not archived. No GitHub releases have been published; model weights are distributed through Hugging Face rather than release tags.

The Apache-2.0 licence permits commercial use, modification, and distribution, provided that the licence notice and any NOTICE file are included with redistributions. The licence does not impose copyleft restrictions, so derived works do not need to be released under the same terms. Users incorporating OmniGen2 into products should verify that the Qwen-VL-2.5 foundation model's licence is also compatible with their intended use, as that model's weights have their own terms separate from the OmniGen2 repository licence.

The EditScore models (OmniGen2-EditScore7B) are distributed through Hugging Face separately and may carry their own licence terms.

Editorial conclusion

OmniGen2 suits researchers and practitioners who want a single open-source model that covers text-to-image generation, image editing, and visual understanding without juggling separate tools for each task. Its ComfyUI integration and optional CPU offload lower the hardware barrier. Teams focused purely on photorealistic generation with a simpler pipeline may find a diffusion-only tool more straightforward. Before adopting, verify that your GPU meets the VRAM requirements stated in the README and that the Apache-2.0 licence fits your deployment terms. The last code push was on 2026-03-20, so any downstream fixes or feature additions will need to come from your own fork until the project resumes activity.

Frequently asked questions

How does OmniGen2 compare to Flux for image generation?

OmniGen2 is a unified multimodal model that handles both visual understanding and image generation from one set of weights, while Flux is a dedicated diffusion model focused on image generation and editing. OmniGen2 claims state-of-the-art performance on instruction-guided image editing benchmarks among open-source models, according to the technical report.

How does OmniGen2 compare to Qwen for image editing?

OmniGen2 is built on a Qwen-VL-2.5 foundation and extends it with image generation capabilities, including instruction-guided image editing. Qwen-VL-2.5 is the vision-language model that provides OmniGen2's visual understanding pathway, so OmniGen2 adds generation on top of what Qwen-VL-2.5 provides.

How does OmniGen2 compare to HiDream?

The OmniGen2 repository and technical report do not include a direct comparison with HiDream. OmniGen2's technical report (arxiv 2506.18871) documents its benchmark results on text-to-image and image editing tasks, but the README does not describe HiDream's architecture or approach.

Official sources

  1. Issues
  2. License: Apache-2.0
  3. README
  4. VectorSpaceLab/OmniGen2 on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/vectorspacelab-omnigen2.svg)](https://hysenlabs.com/projects/vectorspacelab-omnigen2)