ControlNet: a locked copy, a trainable copy, and the 1.1 merge that never landed here
Let us control diffusion models!
At a glance
- What is it?
- ControlNet is the official implementation of adding conditional control to text-to-image diffusion models, built by duplicating network blocks into a locked copy and a trainable copy connected by zero convolutions. Nine Gradio apps show it working, version 1.1 lives in a separate repository, and the last push to this one was 2024-02-25.
- Who is it for?
- ControlNet fits a team that already runs Stable Diffusion 1.5 and wants to condition generation on a specific structure: an edge map, a line drawing, a pose, a segmentation, or a scribble. It does not fit a project looking for a packaged library, because what ships here is nine Gradio scripts plus training tutorials, and the Gradio interface is described in the README as difficult to customize.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Probably not. The repository last received commits 31 months ago, on February 25, 2024.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.
Editorial analysis
Every block is copied twice, and only one copy trains
The whole idea fits in three sentences. ControlNet copies the weights of neural network blocks into a locked copy and a trainable copy. The trainable one learns your condition. The locked one preserves your model.
That division is what makes the method usable on an existing checkpoint. Because the original weights sit in a copy that is not trained, training on a small dataset of image pairs will not destroy a production-ready diffusion model. The README is blunt about the consequence: no layer is trained from scratch, you are still fine-tuning, and your original model is safe. From that follows the practical claim, which is that training works on small-scale or even personal devices, and that the result stays friendly to merging, replacement, offsetting of models, weights, blocks and layers.
For a team, that is the difference between a research artefact and something you can point at a customer checkpoint. You are not forking a diffusion model. You are attaching a second set of blocks to one you already have.
Zero convolution starts at zero, and the obvious objection lives in docs/faq.md
The connector between the two halves is the zero convolution: a 1x1 convolution whose weight and bias are both initialised as zeros. Before any training happens, every zero convolution outputs zeros, so ControlNet adds no distortion to the model it is attached to. Step zero of your run produces exactly the behaviour of the base model.
Then there is the objection that the project answers in its own FAQ. If the weight of a convolution layer is zero, the gradient is also zero, so the network should not learn anything. Why does the zero convolution work? The answer given is two words: this is not true. The explanation itself is not in the README. It is in docs/faq.md.
So if you are evaluating the method, that file is the one to read, and knowing it exists saves you from concluding the design is broken. It is also the pattern to expect elsewhere in this repository: the README states the mechanism and links the reasoning, and the reasoning lives under docs/.
Fourteen copies of the structure, hung off an encoder that keeps no gradients
Repeating the locked and trainable structure fourteen times is what produces a controllable Stable Diffusion. The ControlNet side reuses the Stable Diffusion encoder as a deep backbone for learning diverse controls, and the README points at two pieces of external work as evidence that the encoder is a good one.
The memory argument is the part that matters when you are choosing hardware. The way the layers are connected is computationally efficient, and the original SD encoder does not need to store gradients, specifically the locked original SD encoder blocks 1234 and the middle block. So although many layers are added, the required GPU memory is not much larger than the original model.
Read that as a budget, not a guarantee. It says the overhead is not proportional to the depth you added, which is what makes 8GB cards plausible, and it is why the low VRAM mode exists as a separate document rather than as a default everyone uses.
Version 1.1 is in another repository and the merge is described as pending
The first line of the README is about a version that is not in this repository. ControlNet 1.1 is released as ControlNet-v1-1-nightly, and the new models are described as waiting to be merged into this repository after the authors make sure everything is good. The news entry for the release is dated 2023/0/14 in the file itself, and the page above it is explicitly labelled as ControlNet 1.0.
So there are two lines to reason about, and the repository has no GitHub releases at all to tell you which one a given artifact came from. The last push to this repository was 2024-02-25. If you are choosing a base to build on, that gap is the first question: are you on 1.0 as published here, or on the 1.1 models hosted in the nightly repository, and which of the Gradio scripts or training tutorials match the weights you downloaded. The README does not reconcile the two, so treat the choice as explicit rather than as a detail someone already made for you.
Nine Gradio scripts, and a UI the README calls hard to customize
The pretrained models ship with nine Gradio apps, one per control type, each launched as a plain Python script. Canny edges run with:
python gradio_canny2image.pyStraight lines use M-LSD detection with gradio_hough2image.py, soft HED boundaries use gradio_hed2image.py and are described as preserving detail well enough for recolouring and stylising, human pose uses gradio_pose2image.py, semantic segmentation uses gradio_seg2image.py with the ADE20K protocol, and there is a depth and a normal map script alongside them. The Canny and M-LSD apps expose their detector thresholds in the interface, and test images sit in test_imgs.
The caveat is stated repeatedly and it is about the tool, not the model. The interface is based on Gradio, and Gradio is described as somewhat difficult to customize. For the scribble app you are told to draw outside the interface, in your own drawing software such as MS Paint, and import the result. For pose, Openpose detects the skeleton for you. An interactive scribble canvas is provided in a separate script, and a paragraph about its limitations is struck through and marked as fixed. Plan for a round trip through an external image editor rather than a drawing surface in the browser.
Setup is a conda environment and two weight directories, both of which must be filled by hand
Installation is two commands in a fresh environment:
conda env create -f environment.yaml
conda activate controlAfter that the work is manual. All models and detectors are downloaded from the project's Hugging Face page, and they have to land in two specific places: base models in ControlNet/models, and detectors in ControlNet/annotator/ckpts. The instruction is to download all necessary pretrained weights and detector models from that page, naming the HED edge detection model, the Midas depth estimation model and Openpose among them.
That is the step where a first run usually goes wrong, because nothing checks it for you. A missing base model or a missing detector produces a script that starts, serves a page, and then fails on the request. Downloading the full set up front, rather than the one model the script you happen to be running uses, is the difference between a working environment and a debugging session. The training side has the same shape, with tutorial_train.py, tutorial_dataset.py, and SD 2.1 variants of both, and the helper scripts tool_add_control.py, tool_add_control_sd21.py and tool_transfer_control.py for adding or moving control into an existing model.
Low VRAM mode, non-prompt mode, and a 45% speedup left as a question
Three smaller capabilities are worth knowing before you plan a deployment. Low VRAM mode was added in February 2023 and lives in docs/low_vram.md, with the instruction to use it if you are on an 8GB GPU or want a larger batch size. Non-prompt mode, also called guess mode, was released later that month as an implementation for generating without a prompt.
The third is a question rather than a feature. In March 2023 the project published a discussion titled Precomputed ControlNet: Speed up ControlNet by 45%, but is it necessary? The title carries its own verdict. Reading that as a caution rather than a recommendation: the number is in a discussion post, not in the README's own feature list, and the project framed the whole point as whether the speedup is worth it at all.
None of this changes the basic shape of the work. What you get here is the official implementation, a set of demonstration scripts, a training tutorial, and a body of design notes. Whether that is enough depends on whether you intended to build a product or to study a method.
Editorial conclusion
ControlNet fits a team that already runs Stable Diffusion 1.5 and wants to condition generation on a specific structure: an edge map, a line drawing, a pose, a segmentation, or a scribble. It does not fit a project looking for a packaged library, because what ships here is nine Gradio scripts plus training tutorials, and the Gradio interface is described in the README as difficult to customize. Before you build on it, decide which line you are using, since 1.1 is in the separate ControlNet-v1-1-nightly repository and the README still says those models have yet to be merged here; place the base model in ControlNet/models and the detectors in ControlNet/annotator/ckpts; turn on low VRAM mode from docs/low_vram.md if you are on an 8GB card; and read docs/faq.md, because the obvious objection to zero convolutions is answered in a file rather than in the README.
Frequently asked questions
How do I use ControlNet?
Create the conda environment with `conda env create -f environment.yaml` and `conda activate control`, then download the base model into ControlNet/models and the detectors into ControlNet/annotator/ckpts from the project's Hugging Face page. Nine Gradio apps are provided, each launched as its own script, such as `python gradio_canny2image.py` for Canny edges and `python gradio_pose2image.py` for human pose.
How do I install ControlNet?
There is no package to install. The setup is a conda environment created from environment.yaml, followed by manually placing the Stable Diffusion models in ControlNet/models and the detector models, including HED, Midas and Openpose, in ControlNet/annotator/ckpts.
How do I use ControlNet with human pose?
Run `python gradio_pose2image.py`, which uses the Openpose detector on an image you supply rather than letting you manipulate a skeleton directly in the interface. The README notes that this model deserves a better interface and that Gradio is hard to customize, so the pose is detected for you.
How do I use ControlNet with Stable Diffusion?
The structure is repeated fourteen times on top of the Stable Diffusion encoder, which is reused as the backbone for learning controls while the original encoder blocks keep their locked copies and store no gradients. The result adds many layers without needing much more GPU memory than the original model.