# JoyAI-Image: one model family for image understanding, generation and editing

> JD Open Source packages an 8B multimodal language model and a 16B multimodal diffusion transformer into a single editing and understanding line, released as four Hugging Face checkpoints and a small Python inference repo.

**jd-opensource/JoyAI-Image** — JoyAI-Image is the unified multimodal foundation model for image understanding, text-to-image generation, and instruction-guided image editing.

- Repository: https://github.com/jd-opensource/JoyAI-Image
- Stars: 2,157 · Forks: 165
- Language: Python
- License: Apache-2.0
- Published: 2026-10-06 · Updated: 2026-10-06 · Language: en
- Canonical page: https://hysenlabs.com/projects/jd-opensource-joyai-image

## An 8B language model bolted to a 16B diffusion transformer

The README describes JoyAI-Image as a unified multimodal foundation model covering image understanding, text-to-image generation and instruction-guided image editing. The architecture is concrete: an 8B Multimodal Large Language Model paired with a 16B Multimodal Diffusion Transformer. That pairing is the whole pitch, because it explains why the project can read a scene and then repaint it without a second model being bolted on afterwards.

The stated principle is a closed loop between understanding, generation and editing. The argument is that better spatial understanding improves grounded generation and controllable editing through scene parsing, relational grounding and instruction decomposition, and that generative transforms such as viewpoint changes feed evidence back into spatial reasoning. Whether that loop delivers on a real benchmark is a question for the technical report, which the README links on arXiv as 2605.04128 and dates to 2026-05-07.

What is less clear from the repository is how much of this is one set of weights and how much is several checkpoints behind a shared interface. The README says one model family through a shared MLLM-MMDiT interface, and the model zoo table that would answer it breaks off partway through the row for JoyAI-Image-Edit. Two entries are legible: JoyAI-Image-Und, described as a text and image understanding backbone for spatial reasoning and editing-aware perception, and JoyAI-Image-Edit itself.

## Ten entries at the top of the repository tree

The repository is small, and the tree explains why. Alongside `README.md`, `LICENSE` and `.gitignore` there is `pyproject.toml`, `requirements.txt`, two inference scripts named `inference.py` and `inference_und.py`, an `src/` package directory, an `assets/` directory and a `test_images/` directory.

Those names line up with the two halves of the model. `inference_und.py` is almost certainly the understanding entry point and `inference.py` the generation or editing one, and `test_images/` holds sample inputs for the latter. There is no `train/` directory, no training script and no evaluation harness at the top level, which is the clearest signal about what this repository is for: consuming released checkpoints, not producing new ones.

The package name in `pyproject.toml` is `joyai-image-release`, and the description field says it plainly, `Package for JoyAI-Image image inference.` Version 0.1.0, `requires-python` set to `>=3.10`, Apache-2.0 license text. This is a release wrapper around model weights hosted elsewhere, not the model codebase.

## Two dependency lists that pin different things

The repository ships two dependency specifications and they do not agree, which is the first thing worth reading carefully. `pyproject.toml` is the stricter one and pins the two libraries where a version mismatch actually bites:

```toml
dependencies = [
    "accelerate",
    "diffusers==0.36.0",
    "einops",
    "kernels",
    "loguru",
    "packaging",
    "pillow",
    "safetensors",
    "sentencepiece",
    "torch==2.8.0",
    "torchvision",
    "transformers>=4.57.0,<4.58.0",
]
```

`diffusers` is pinned to an exact patch release and `torch` to an exact minor release, while `transformers` is constrained to a single minor version band. That combination is a signal: the code depends on behaviour that Diffusers changed between releases, and the transformers pin exists so the text and vision stack matches what the checkpoint was converted against. `requirements.txt` is the looser sibling, asking for `diffusers>=0.34.0` and unversioned `torch`.

The bigger difference is that `requirements.txt` adds a web service layer that `pyproject.toml` does not: `fastapi`, `uvicorn` and `pydantic>=2` alongside `requests`. So the plain requirements file expects you to expose the model over HTTP, while the packaged project exposes only inference and carries a single optional extra, `rewrite = ["openai"]`. If you are serving the model behind an API, install the requirements file; if you are importing it, take the pins from `pyproject.toml`.

## Installing the inference code and reaching for the weights

Because the repository itself contains no weights, an install is two steps: get the code and its dependencies, then fetch a checkpoint from a model hub. The dependency half is unremarkable:

```bash
git clone https://github.com/jd-opensource/JoyAI-Image
cd JoyAI-Image
pip install -r requirements.txt
```

The weight half is the part with choices. The README points at four Hugging Face repositories under the `jdopensource` account: `JoyAI-Image-Edit-Diffusers` and `JoyAI-Image-Edit-Plus-Diffusers` for use through the Diffusers library, and `JoyAI-Image-Edit-ComfyUI` and `JoyAI-Image-Edit-Plus-ComfyUI` for ComfyUI. The same four are mirrored on ModelScope, which matters for anyone behind a network where Hugging Face is slow or blocked.

There are two separate product lines hiding in those names. JoyAI-Image-Edit is the original instruction-guided editing model, weights released on 2026-04-02. JoyAI-Image-Edit-Plus, announced on 2026-06-23, is the newer one. They are separate downloads at separate sizes, and the README does not give a parameter count for either, so pick based on which checkpoint the Diffusers integration you install expects rather than on the model zoo labels.

## What changed between April and August 2026

The README carries a dated news list rather than a release history, and GitHub shows no published releases for this repository. That distinction matters: there is no versioned changelog to diff, so the news list is the maintenance record. The last push was on 2026-08-05, which lines up with the top entry announcing JoyAI-Video-Edit as a separate repository for real-time open-ended video editing.

Reading the list backwards gives the shape of the project. On 2026-04-02 the editing weights appeared. On 2026-04-06 the demos went live on Hugging Face Spaces. On 2026-04-10 the SpatialEdit training dataset and the SpatialEdit-Bench benchmark were published as separate datasets, and on 2026-04-11 JoyAI-Image-Edit gained Diffusers support through a dedicated repository. On 2026-04-15 the OpenSpatial data engine and the OpenSpatial-3M dataset were released, with the engine's code in a separate repository under a different owner.

After that came distribution work rather than new modelling. On 2026-05-08 the README notes that Diffusers merged the project's pull request 13444, which moved JoyAI-Image-Edit into the library itself. On 2026-06-23 came the Edit-Plus release, and on 2026-07-17 native ComfyUI support, including a workflow JSON file to load without extra dependencies. For a reader deciding whether to adopt this, that sequence is more informative than the feature bullets: the effort has moved from research releases toward making the model usable inside existing image tooling.

## Where the README stops and the model hubs take over

The README in this repository is an index, not a manual. It names the model, describes the architecture in a paragraph, lists four checkpoints and two Space demos, and then points outward. Installation, usage snippets, hardware requirements, prompt format and the ComfyUI workflow details all live in the individual Hugging Face model repositories, the ComfyUI workflow file, the arXiv report and the OpenSpatial data engine repository.

That split has a practical consequence for evaluation. Nothing in this repository will tell you what a single edit costs in VRAM, whether Edit-Plus is a drop-in replacement for Edit, or how the spatial benchmark scores compare to other editing models. The technical report is the place for the claims about layout fidelity and long-text typography; the benchmark dataset is the place to reproduce them.

What the repository does settle is the integration surface, and that is genuinely useful. Diffusers support is merged upstream rather than requiring a fork, ComfyUI support is native with a loadable workflow, the licenses are Apache-2.0 in the code and published weights, and the surrounding data, benchmarks and data generation engine are all released rather than kept internal. A project of this size that keeps its training data engine and its benchmark sets public is easier to evaluate than the many that publish weights alone.

## Conclusion

JoyAI-Image is a reasonable choice if your work is instruction-guided image editing with a spatial angle, and you are already running Diffusers on hardware that can hold a 16B diffusion transformer. It is a poor choice as a text-to-image generator you plan to fine-tune, because the repository has no training code, and a poor choice if you need the 8B understanding backbone to run without the diffusion side. Read `pyproject.toml` first: the pins there tell you more about what will actually work than the feature list does, then check whether the Diffusers checkpoint you want carries a diffusers `model_index.json` before you budget the download.

## FAQ

### How does JoyAI-Image-Edit Plus compare when used in ComfyUI?

The README documents both models as native ComfyUI checkpoints rather than describing differences between them. `JoyAI-Image-Edit-ComfyUI` and `JoyAI-Image-Edit-Plus-ComfyUI` are separate repositories, and the Edit-Plus one is the later release, published 2026-06-23 against the original Edit weights of 2026-04-02. Each ships a workflow JSON, and the README says the workflow loads in current ComfyUI with no additional dependencies. Any quality comparison has to come from running both, since the README does not make one.

### How do I install the JoyAI-Image inference code?

Clone the repository and install `requirements.txt`, which asks for Python 3.10 or newer and includes `fastapi` and `uvicorn` if you intend to serve the model over HTTP. Two entry points are provided: `inference.py` for the generation and editing side and `inference_und.py` for the understanding backbone. The model weights themselves are not in the repository, they download from Hugging Face or ModelScope.

### What checkpoints does JoyAI-Image publish?

Four Hugging Face repositories under `jdopensource`: `JoyAI-Image-Edit-Diffusers`, `JoyAI-Image-Edit-Plus-Diffusers`, `JoyAI-Image-Edit-ComfyUI` and `JoyAI-Image-Edit-Plus-ComfyUI`, all mirrored on ModelScope. There is also a JoyAI-Image-Und understanding backbone hosted inside the Edit repository tree. The model zoo table in the README lists the same families by task, covering multimodal understanding and image editing.

## Sources

- [Issues](https://github.com/jd-opensource/JoyAI-Image/issues)
- [jd-opensource/JoyAI-Image on GitHub](https://github.com/jd-opensource/JoyAI-Image)
- [License: Apache-2.0](https://github.com/jd-opensource/JoyAI-Image/blob/main/LICENSE)
- [README](https://github.com/jd-opensource/JoyAI-Image/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/jd-opensource-joyai-image
