Open-source project
xinntao/Real-ESRGAN avatar
xinntao/Real-ESRGAN

Real-ESRGAN restores from a degradation pipeline it invented, and that is the trade

Real-ESRGAN aims at developing Practical Algorithms for General Image/Video Restoration.

36,960 stars4,485 forksPythonBSD-3-Clause

At a glance

What is it?
Real-ESRGAN extends ESRGAN into a restoration application trained on pure synthetic data, shipping four families of weights, one documented noise control, and a Python package that sits well past its last release tag. Here is what each part gives you and where it stops working.
Who is it for?
Adopt Real-ESRGAN when the source is photographs, anime frames or video you are willing to have reinterpreted, and reach for the ncnn-vulkan archives when you want a first look without touching Python. Leave it alone if you need bit-accurate recovery, if your workload is medical or forensic, or if you need a release cadence newer than v0.3.0 from 2022-09-20.
Can I use it commercially?
Yes. BSD-3-Clause is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Probably not. The repository last received commits 26 months ago, on August 6, 2024.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

Blind super-resolution built on a degradation pipeline nobody measured

Real-ESRGAN is an extension of ESRGAN, and the extension is the whole point. The paper behind it is titled Training Real-World Blind Super-Resolution with Pure Synthetic Data, and the project says the model is trained with pure synthetic data rather than on matched pairs of real photographs. That one decision fixes what the tool can promise. A supervised super-resolution model wants a low and high image of the same scene. Real capture pairs are expensive and hard to match, so this project simulates the degradation instead and lets the network learn to invert it.

The cost of that shortcut lands on you at inference time. Because the degraded input never came from the sensor that produced a clean version, the model is estimating which blur, noise and compression chain produced your file, then undoing its own guess. When the guess is close, a 4x result reads as recovered detail. When it is not, the network writes texture that has no source in the original: plausible grain across skin, invented leaf patterns on a scanned print, an edge sharpened past anything the lens caught. On a photograph you intend to keep, that difference is the whole argument. User complaints about exactly this land in docs/feedback.md, which is where the project records them.

realesr-general-x4v3 has a -dn denoising switch, the x4plus and anime weights do not

Model choice does more work here than flag choice. The update list introduces realesr-general-x4v3 as a small model for general scenes and gives it one documented control: -dn, short for denoising strength, to balance noise and avoid over-smooth results. That option belongs to this model in the entry that introduces it, and no equivalent switch is named for RealESRGAN_x2plus, for the x4plus family, or for the anime weights, which are AnimeVideo-v3 and RealESRGAN_x4plus_anime_6B. Reach for an anime or x4plus file on noisy phone footage because it looks sharper on paper and you have no documented way to trade sharpness against noise on that model, so over-smoothed faces and flattened foliage are failures you have to go looking for yourself.

For anime frames the comparison target is named outright. RealESRGAN_x4plus_anime_6B is described as optimized for anime images with much smaller model size, and the comparison is run against waifu2x in docs/anime_model.md, pointing at its ncnn-vulkan build. The difference in approach is packaging. waifu2x sits in its own project, while this one ships a PyTorch package, a Cog entry point and a separate Real-ESRGAN-ncnn-vulkan repository, with video model comparisons in docs/anime_comparisons.md. Which one reconstructs line work better is settled inside those documents, not here.

--outscale is a LANCZOS4 resize after the network has already run

A flag called --outscale does not scale the way the name suggests. Arbitrary scale is supported with --outscale, and the detail that matters is in the same sentence: it further resizes outputs with LANCZOS4. The network runs at its trained factor first, then a conventional Lanczos resize trims the result to the size you asked for. Point an x4 model at a 2x request and the machine renders four times the linear dimension before throwing three quarters of it away by interpolation. You pay full 4x compute and buy resampling where you wanted inference, which is the wrong trade for a queue of frames and the wrong claim to put in a methods section.

Input handling is wider than the scaling story. The inference code is listed as supporting tile options, images with an alpha channel, gray images and 16-bit images. Accepting a format is a statement about the loader, not about fidelity, and nothing in the repository lists 16-bit training data or evaluation numbers for it. A 16-bit archival scan is therefore taken without any stated account of how the model handles noise at that bit depth, so try one crop before pointing a whole collection at it.

setup.py writes realesrgan/version.py, and it stamps unknown without a .git directory

The repository root explains a mechanics question. inference_realesrgan.py and inference_realesrgan_video.py are the two entry points, one for stills and one for video, and cog_predict.py with cog.yaml sit beside them for container use. The Python package lives in realesrgan/, with options/ for configuration, experiments/ for training settings, inputs/ and weights/ as working directories, docs/ holding model_zoo.md, Training.md, FAQ.md, anime_video_model.md and feedback.md, and tests/ and scripts/ next to them.

Version identity here is generated rather than hand written. setup.py reads the VERSION file and writes realesrgan/version.py containing __version__, __gitsha__ and version_info, stamping the current commit into the build. That commit comes from a shell call inside setup.py:

python
out = _minimal_ext_cmd(['git', 'rev-parse', 'HEAD'])

get_hash() returns the string 'unknown' when there is no .git directory, which is the situation for a source distribution or a wheel built outside a checkout. Practical consequence: the wheel you install from PyPI reports a version with no commit attached, and the only honest answer to which build is running sits in the generated realesrgan/version.py inside the installed package. .github/workflows/publish-pip.yml is what pushes that build outward.

requirements.txt sets floors, no ceilings, and installs a face stack nobody asked for

requirements.txt lists basicsr>=1.4.2, facexlib>=0.2.5, gfpgan>=1.3.5, numpy, opencv-python, Pillow, torch>=1.7, torchvision and tqdm. Read that as a description of what the code imports, not as a lockfile. Nothing caps basicsr, so a fresh install resolves whatever the newest compatible release of the restoration toolbox happens to be on the day you build, and nothing in the repository would catch a breaking change in the library the code was written against 1.4.2.

Two floors are face specific and unconditional. gfpgan and facexlib enter the environment whether or not you turn on face enhancement, and GFPGAN appears in the recommended projects as a practical algorithm for real-world face restoration, a different task from the general one this repository performs. A video-only deployment inherits a face stack it never calls.

BasicSR holds the mirror position. It is a hard dependency here and also a recommended project, described as an open source image and video restoration toolbox, so slimming an image by uninstalling it is not on the table. On hardware, the dependency list names PyTorch >= 1.7 and Python >= 3.7 and recommends Anaconda or Miniconda, but no CUDA build and no supported GPU appear anywhere in it. Driver and CUDA compatibility is therefore the first thing to verify on a new machine, before a long render queue.

The portable ncnn-vulkan zips are 20220424 builds living outside the Python package

A second distribution channel skips pip entirely. The README links portable executable archives for Windows, Linux and macOS, described as executable files for Intel, AMD and Nvidia GPU, and each file name carries its own build date: realesrgan-ncnn-vulkan-20220424-windows.zip, realesrgan-ncnn-vulkan-20220424-ubuntu.zip and realesrgan-ncnn-vulkan-20220424-macos.zip, all served from the v0.2.5.0 release page. The ncnn implementation is a separate repository, Real-ESRGAN-ncnn-vulkan, so a reader on Windows can try the model with no Python at all, and a reader on Linux can hold that build against the package.

The dates are the catch. Those archives are an April 2022 build, while the default branch was last pushed on 2024-08-06 and the update list advertises realesr-general-x4v3 and an updated AnimeVideo-v3. Nothing in the README states which weight files each archive bundles, so the catalog in docs/model_zoo.md describes the Python path and leaves the portable one unstated. If you compare results between a zip and a fresh clone, the build date is the first variable to hold constant.

Hosted demos exist for the same reason: a Replicate demo, two Colab notebooks, one for Real-ESRGAN and one for anime videos, and a Gradio demo on Huggingface Spaces. A commented out line in the README points at the Tencent ARC demo page and notes it supports only RealESRGAN_x4plus_anime_6B, a narrower lineup than anywhere else. For inspecting a face result before spending GPU hours, the recommended projects also name HandyView, a PyQt5 based image viewer built for viewing and comparison.

Last push 2024-08-06, last tag v0.3.0 from 2022-09-20, two years of untagged commits

The maintenance picture is uneven in a way worth stating plainly. The repository is not archived, and the last push was on 2024-08-06. The most recent tag is v0.3.0 from 2022-09-20, ahead of v0.2.5.0 from 2022-04-24 and v0.2.4.0 from 2022-02-15. Master therefore carries roughly two years of commits that no release note describes, and the two distribution paths sit at opposite ends of that gap: a pip install hands you the v0.3.0 era, a clone of master hands you the 2024 state.

Training is documented more thoroughly than inference. The training codes have been released with a guide in docs/Training.md, and finetuning is supported on your own data or on paired data, described as finetuning ESRGAN. That is the route to take when your source is one camera or one compression scheme rather than a general mixture, and it asks for matched degradation data of your own before it gives anything back.

Licensing is the plainest line in the repository. LICENSE sits at the top level under BSD-3-Clause, no CLA file appears in the top level entries, and the weights are fetched from GitHub releases rather than vendored, so a commercial deployment should read both the LICENSE and the terms attached to those model downloads before shipping.

Editorial conclusion

Adopt Real-ESRGAN when the source is photographs, anime frames or video you are willing to have reinterpreted, and reach for the ncnn-vulkan archives when you want a first look without touching Python. Leave it alone if you need bit-accurate recovery, if your workload is medical or forensic, or if you need a release cadence newer than v0.3.0 from 2022-09-20. First thing to check in any environment: open realesrgan/version.py in the installed package and confirm which commit it stamped, because a PyPI wheel and a clone of master are not the same code.

Frequently asked questions

How do I install Real-ESRGAN?

The installation section starts by cloning https://github.com/xinntao/Real-ESRGAN.git. Dependencies are Python >= 3.7 and PyTorch >= 1.7, with Anaconda or Miniconda recommended, and the version floors live in requirements.txt. A PyPI package is also linked from the README.

How do I use Real-ESRGAN for video?

Video runs through inference_realesrgan_video.py at the repository root, and AnimeVideo-v3 is the model named for it. The README points to docs/anime_video_model.md for those models and to docs/anime_comparisons.md for the comparisons.

How do I use Real-ESRGAN in Python?

The package is realesrgan/, published to PyPI under the name realesrgan. setup.py reads the VERSION file and generates realesrgan/version.py, and cog_predict.py with cog.yaml sit at the root for container deployment.

Is Real-ESRGAN open source?

The repository ships LICENSE at the top level under BSD-3-Clause, and no CLA file is listed among the top level entries. The model weights are downloaded from GitHub releases rather than stored in the repository.

What does Real-ESRGAN do?

It extends ESRGAN into a restoration application trained with pure synthetic data, aimed at practical image and video restoration. The paper behind it is Training Real-World Blind Super-Resolution with Pure Synthetic Data.

Official sources

  1. Official README
  2. Project repository
  3. Release notes
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/xinntao-real-esrgan.svg)](https://hysenlabs.com/projects/xinntao-real-esrgan)