Open-source project
TencentARC/ColorFlow avatar
TencentARC/ColorFlow

ColorFlow: Diffusion-Based Sequential Image Colorization with Retrieval-Augmented Identity Preservation

The official implementation of paper "ColorFlow: Retrieval-Augmented Image Sequence Colorization". ColorFlow:基于检索增强的图像序列上色

463 stars44 forksPythonNOASSERTION

At a glance

What is it?
ColorFlow is an official Python implementation from TencentARC of a three-stage diffusion framework for colorizing black-and-white cartoon and comic image sequences, using reference image retrieval to preserve character and object color identity without requiring per-character finetuning. The training code is not yet released.
Who is it for?
ColorFlow is aimed at researchers and studios working on automatic colorization of sequential artwork, particularly cartoons and manga where character identity must remain consistent across dozens or hundreds of frames. It is not suited for teams that need to train their own models: the training code is not yet released, and the repository notes it as a pending TODO.
Can I use it commercially?
Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
Is it still maintained?
Yes. The repository last received commits 47 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The Identity Consistency Problem in Sequential Colorization

Colorizing a single black-and-white image with a generative model is well-studied, but colorizing an entire manga chapter or cartoon episode consistently is much harder. The same character must appear in the same colors across every panel. Existing methods either require per-character finetuning, which is expensive at scale, or use explicit identity embedding extraction, which relies on predefined character templates and breaks when a character changes pose or is partially obscured.

ColorFlow addresses this by treating it as a retrieval and in-context learning problem. Given a reference image showing a colored version of a character, the model retrieves relevant color information and applies it to a new panel without needing a separate finetuning step for each character. The README describes this as a Retrieval Augmented Colorization pipeline. The target audience is researchers in visual colorization and studios that colorize manga or animation frames at scale.

Three-Stage Framework: RAP, ICP, and GSRP

The README describes the framework as having three primary components. The Retrieval-Augmented Pipeline, or RAP, locates relevant color references for the instances present in the input image. The In-context Colorization Pipeline, or ICP, uses a dual-branch diffusion model where one branch extracts color identity from the retrieved references and the other branch performs the actual colorization. The two branches share information through the self-attention mechanism in the diffusion model, which the README describes as enabling strong in-context learning and color identity matching.

The Guided Super-Resolution Pipeline, or GSRP, is the third stage. Its role is to produce high-quality output while maintaining the color assignments established in the previous stage. Together, these three components form a pipeline that the paper claims sets a new standard on the ColorFlow-Bench benchmark.

The dual-branch design is the key architectural choice. Unlike methods that require an explicit identity embedding from a fixed template, ColorFlow uses the retrieved reference images directly as conditioning context. This means the model can generalize to characters it has not seen during training, as long as a suitable colored reference image is provided at inference time.

Installing ColorFlow and Running the Gradio Interface

The README documents a local setup using Anaconda or Miniconda. Clone the repository and create a dedicated environment:

bash
git clone https://github.com/TencentARC/ColorFlow
cd ColorFlow
bash
conda create -n colorflow python=3.8.5
conda activate colorflow
pip install -r requirements.txt

The requirements.txt file pins specific versions: torch==2.0.0, torchvision==0.15.1, gradio==4.44.1, transformers==4.46.3, and a custom local install of the diffusers package from the repository's own diffusers/ subdirectory. After installation, launch the Gradio web interface:

bash
python app.py

The interface is then accessible at http://localhost:7860. For remote servers, the README notes that replacing localhost with the server IP or domain works, and that the port can be changed by editing the server_port parameter in the demo.launch() call inside app.py. A Hugging Face Space at huggingface.co/spaces/TencentARC/ColorFlow provides a hosted demo without local installation.

ColorFlow-Bench and Evaluation Results

The repository introduces a benchmark called ColorFlow-Bench for reference-based sequential colorization evaluation. The README states that ColorFlow outperforms existing models across multiple metrics on this benchmark. Because the benchmark was introduced alongside the model in the same paper, independent validation from third parties is limited at this stage.

The paper is arXiv:2412.11815, authored by Junhao Zhuang, Xuan Ju, Zhaoyang Zhang, Yong Liu, Shiyi Zhang, Chun Yuan, and Ying Shan. The model weights for the inference pipeline and a Sketch_Shading model variant are available on the Hugging Face model hub at huggingface.co/TencentARC/ColorFlow. The Sketch_Shading model and its associated code were added in an update dated December 23, 2024, several days after the initial release on December 17, 2024.

Limitations: Training Code, Python Version, and License

The repository's TODO list shows the training code as not yet released. This is a significant constraint for researchers who want to replicate the paper's results, fine-tune the model on their own dataset, or adapt the architecture. Only the inference pipeline is available. If training reproducibility matters, this repository does not yet support it.

The environment requires Python 3.8.5. This is an older version. Many dependencies in requirements.txt are pinned to specific versions, which means compatibility with newer Python or PyTorch versions is not guaranteed. Setting up an isolated conda environment as documented is not optional; attempting to install into an existing modern Python environment will likely encounter dependency conflicts.

The repository license is NOASSERTION, meaning it does not carry a standard SPDX identifier. The README directs users to review the LICENSE file for details. Teams that need a clear open source grant before using or integrating the model weights should read that file before proceeding.

Comparison with ScreenStyle

The README's acknowledgments section names ScreenStyle as one of the projects that inspired ColorFlow. ScreenStyle, from the paper by Msxie92, addresses manga colorization using a style-transfer approach with screen tone extraction. The key difference is in how identity is handled: ScreenStyle focuses on transferring a reference style to the target image, while ColorFlow uses retrieval to match specific character instances and preserve their color identity across a sequence.

ScreenStyle does not have a three-stage super-resolution pipeline and does not use the retrieval-augmented approach. For a single image or a style transfer task, ScreenStyle may be sufficient. ColorFlow targets the sequential consistency problem, where the same character must be colored consistently across many frames. If the use case is a single image or an artistic style transfer rather than a frame-by-frame sequence, ScreenStyle or simpler diffusion-based colorization tools may be more proportionate choices.

Editorial conclusion

ColorFlow is aimed at researchers and studios working on automatic colorization of sequential artwork, particularly cartoons and manga where character identity must remain consistent across dozens or hundreds of frames. It is not suited for teams that need to train their own models: the training code is not yet released, and the repository notes it as a pending TODO. Before deploying it, verify that your environment meets the Python 3.8.5 and CUDA requirements, and review the LICENSE file directly rather than assuming a standard open source license, since the repository identifies it as NOASSERTION.

Frequently asked questions

Does ColorFlow require per-character finetuning for each new character?

No. ColorFlow's Retrieval Augmented Colorization pipeline works by retrieving a relevant colored reference image at inference time, without a separate finetuning step for each character. This is described in the README as the key difference from existing per-ID finetuning methods.

Is the training code for ColorFlow available?

Not yet. The repository's TODO list marks the release of training code as a pending item. Only the inference code and model weights are currently released, along with a Sketch_Shading model variant added in December 2024.

Where can I access ColorFlow model weights?

The model weights are hosted on Hugging Face at huggingface.co/TencentARC/ColorFlow. A hosted demo is available at huggingface.co/spaces/TencentARC/ColorFlow without requiring local installation.

Official sources

  1. Issues
  2. Project website
  3. README
  4. TencentARC/ColorFlow on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/tencentarc-colorflow.svg)](https://hysenlabs.com/projects/tencentarc-colorflow)