Open-source project
XingangPan/DragGAN avatar
XingangPan/DragGAN

DragGAN: point-based image editing with StyleGAN2 weights

Official Code for DragGAN (SIGGRAPH 2023)

35,746 stars3,395 forksPythonNOASSERTION

At a glance

What is it?
DragGAN is the SIGGRAPH 2023 reference implementation of interactive point-based manipulation on the generative image manifold. It is a research codebase with a GUI, a Gradio front end and a 25GB Docker image, not a consumer app.
Who is it for?
Adopt DragGAN if you have an NVIDIA GPU, a Conda environment and a reason to edit StyleGAN latents directly; the README's own path is conda env create -f environment.yml, then python scripts/download_model.py, then sh scripts/gui.sh. Do not adopt it if you need to edit photographs straight from a camera roll: the README states that real images require GAN inversion with an external tool such as PTI first, and that step is not shipped here.
Can I use it commercially?
Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
Is it still maintained?
Probably not. The repository last received commits 28 months ago, on May 18, 2024.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What DragGAN actually solves, and for whom

Most image editors ask you to move pixels. DragGAN asks you to move a point on a generated image and then solves for a latent code that puts content there. The README describes the project as interactive point-based manipulation on the generative image manifold, and the workflow follows that description: you place a handle point on a feature, a target point where it should go, and the optimiser updates the latent vector so the handle travels toward the target while the rest of the image stays plausible. The intended user is a researcher or graphics engineer who already works with StyleGAN2 checkpoints and wants a visual front end for latent optimisation. It is not aimed at someone who wants to retouch a JPEG. The README is explicit that editing a real image requires GAN inversion through an external tool such as PTI, after which you load the resulting latent code and model weights into the GUI. That single sentence sets the boundary of the audience: if you cannot produce a latent code, the tool has nothing to edit.

The mechanism: latent optimisation behind a GUI

The repository layout shows the split clearly. The DragGAN algorithm lives in the top-level visualiser scripts, visualizer_drag.py and visualizer_drag_gradio.py, with supporting code in gui_utils/, gradio_utils/ and viz/. Everything else is inherited: dnnlib/, torch_utils/, training/, legacy.py and the stylegan_human/ directory come from the StyleGAN3 codebase, which the README names as the base, with additional code borrowed from StyleGAN-Human. That inheritance matters when you read the licence section, because it means two licence regimes coexist in one tree. The generator itself is a StyleGAN2 network loaded from checkpoints under ./checkpoints, and the drag operation is an optimisation loop over the latent vector rather than a diffusion or inpainting pass. There is no separate model for editing. The same generator that produced the image is the one being steered, which is why the results stay on the manifold and why the method is limited to whatever that particular generator can already represent. A checkpoint trained on faces will not grow a second head convincingly, and no amount of dragging changes that.

Installing DragGAN and running the GUI for the first time

The README gives two installation routes. With an NVIDIA CUDA card it points to the requirements of NVlabs/stylegan3 and provides an environment file, followed by a separate requirements install. Run these two commands from the repository root and you should end up in a Conda environment named stylegan3 with the CUDA build of PyTorch and the GUI packages present.

bash
conda env create -f environment.yml
conda activate stylegan3
pip install -r requirements.txt

If you are on an Apple Silicon Mac or have no GPU, the README filters the NVIDIA entries out of the environment file before creating it, and sets an environment variable so that unsupported MPS operations fall back to CPU.

bash
cat environment.yml | \
  grep -v -E 'nvidia|cuda' > environment-no-nvidia.yml && \
    conda env create -f environment-no-nvidia.yml
conda activate stylegan3

# On MacOS
export PYTORCH_ENABLE_MPS_FALLBACK=1

Before the GUI can do anything you need weights. The repository ships a download script, and the README notes that StyleGAN-Human and the Landscapes HQ dataset weights are fetched separately from Google Drive links and placed under ./checkpoints.

bash
python scripts/download_model.py

Then start the interface. The shell script is the documented entry point on Linux and macOS, and the batch file is the Windows equivalent. Both launch the same GUI, which the README describes as supporting editing of GAN-generated images.

bash
sh scripts/gui.sh

There is also a Gradio front end that the README calls universal for both Windows and Linux, started directly with python visualizer_drag_gradio.py. If you would rather not build the environment at all, the README offers a Docker route based on the NGC PyTorch image. The build produces a container tagged draggan:latest, and the run command publishes port 7860 and mounts the working directory. The README warns that the image takes about 25GB of disk space, and that adding --gpus all is what lets the container use an NVIDIA card.

bash
docker build . -t draggan:latest
docker run -p 7860:7860 -v "$PWD":/workspace/src -it draggan:latest bash
cd src && python visualizer_drag_gradio.py --listen

Where DragGAN stops being the right tool

The hardest limit is the one stated in the README rather than discovered in the code: real photographs need GAN inversion before DragGAN can touch them, and the inversion tool is not part of this repository. You are expected to run something like PTI yourself, then import the latent code and weights. That is a research pipeline, not a feature. The second limit is hardware. The CPU and Apple Silicon path exists, but the README frames it as a fallback for machines without CUDA, and the Docker image is built on an NVIDIA base image with an explicit note about adding --gpus all for acceleration. Anyone expecting interactive latency on a laptop without a discrete GPU should read those lines before installing. The third limit is scope: the GUI edits GAN-generated images, and the quality of any edit is bounded by the checkpoint you loaded. The README suggests other pretrained StyleGAN models are fair game, which sounds open-ended but in practice means you own the job of finding a generator whose latent space covers the edit you want. There is no rollback mechanism documented in the README, so an edit that goes wrong is recovered by reloading a latent, not by an undo stack.

How DragGAN differs from prompt-driven editors

The obvious alternative for most people is a text-conditioned diffusion editor, where you describe the change in words and the model regenerates the region. The difference is control versus specification. DragGAN takes an explicit spatial handle and target, so the user decides exactly where content moves, and the optimiser is constrained to the manifold of one generator. A prompt-driven tool takes a description and lets the model decide what the pixels should become, which handles photographic input natively and needs no inversion step. Neither dominates. If your source is a real photo and you cannot run inversion, the diffusion route is the only one that starts. If you need a precise, repeatable displacement on a latent you already control, dragging points is a more direct interface than writing a prompt and hoping the model lands on the geometry you had in mind. Within the same family, the README points to StyleGAN-Human and LHQ checkpoints as alternative weight sets, which is a reminder that the editing method is fixed while the domain of the generator is not.

Licence, watermarking and what adoption costs

The licence situation is the part most likely to end a commercial evaluation. The README states that the DragGAN algorithm code is under CC-BY-NC, a non-commercial licence, while code used or modified from StyleGAN3 falls under the Nvidia Source Code License. The repository's LICENSE.txt is the file to read, and the metadata does not resolve to a standard SPDX identifier. The README also states that any form of use and derivative of this code must preserve the watermarking functionality showing AI Generated. That is a condition on outputs, not just on the source, and it is unusual enough that it should be checked against your intended use before you build anything on top. This is a description of what the README says, not legal advice; if the terms affect a product, that is a question for a lawyer. On maintenance, the metadata records the last push date as unknown and lists no releases, so there is no basis here for claiming the project is actively developed. Plan on the code as it stands, and expect to pin your own dependency versions, because requirements.txt allows floating minimums such as torch>=2.0.0 and gradio>=3.35.2 that will resolve differently over time.

Editorial conclusion

Adopt DragGAN if you have an NVIDIA GPU, a Conda environment and a reason to edit StyleGAN latents directly; the README's own path is conda env create -f environment.yml, then python scripts/download_model.py, then sh scripts/gui.sh. Do not adopt it if you need to edit photographs straight from a camera roll: the README states that real images require GAN inversion with an external tool such as PTI first, and that step is not shipped here. Also skip it if CC-BY-NC and the Nvidia Source Code License terms do not fit your use, or if a 25GB Docker image and the watermarking requirement are unacceptable. Verify first that your CUDA setup matches what NVlabs/stylegan3 requires, and check that the pretrained weights download completes before you plan any work around the GUI.

Frequently asked questions

Is DragGAN AI free to use?

The code is available at no cost, but the README states that the DragGAN algorithm code is licensed under CC-BY-NC, which is a non-commercial licence, and that parts derived from StyleGAN3 fall under the Nvidia Source Code License. The README also requires any use or derivative to preserve watermarking that shows AI Generated.

What is a DragGAN alternative?

The README does not list alternatives. It does point to StyleGAN-Human and the Landscapes HQ dataset as other pretrained weights you can load under ./checkpoints, and to PTI as the tool for GAN inversion when you want to edit a real image. The closest change in approach is a text-conditioned diffusion editor, which takes a written description instead of handle and target points.

How do I install DragGAN?

The README gives conda env create -f environment.yml followed by conda activate stylegan3 and pip install -r requirements.txt for CUDA machines, and a filtered environment file for Macs or CPU-only machines. Weights come from python scripts/download_model.py, and the GUI starts with sh scripts/gui.sh.

Can DragGAN edit a real photograph?

The README states that to edit a real image you must first perform GAN inversion using a tool such as PTI, then load the new latent code and model weights into the GUI. That inversion step is not included in this repository.

Does DragGAN run without an NVIDIA GPU?

The README provides a route that strips nvidia and cuda entries from environment.yml for Apple Silicon Macs or plain CPU use, and sets PYTORCH_ENABLE_MPS_FALLBACK=1 on MacOS. The Docker image is built on an NVIDIA base image, and the README notes that --gpus all is what enables GPU acceleration inside the container.

Official sources

  1. Official documentation
  2. Official README
  3. Project repository
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/xingangpan-draggan.svg)](https://hysenlabs.com/projects/xingangpan-draggan)