Model or dataset
CompVis/stable-diffusion avatar
CompVis/stable-diffusion

Stable Diffusion v1 is four checkpoints and a 2024 commit date, and the README tells you not to ship the weights

A latent text-to-image diffusion model

73,490 stars10,563 forksJupyter NotebookNOASSERTION

At a glance

What is it?
CompVis's Stable Diffusion repository is the reference implementation for the original latent text-to-image model: an 860M UNet with a frozen CLIP ViT-L/14 text encoder, four published checkpoints, and a script that watermarks everything it renders. The code has not moved since 2024-06-18 and the project describes its own weights as research artifacts.
Who is it for?
Use this repository when you want the original v1 model with its documented checkpoint lineage and a sampling script you can read, not when you want the current text-to-image capability, since the last push was 2024-06-18 and the newest checkpoint here is sd-v1-4.ckpt.
Can I use it commercially?
Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
Is it still maintained?
Probably not. The repository last received commits 27 months ago, on June 18, 2024.
What is it written in?
Mainly Jupyter Notebook, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The last push was 2024-06-18 and the newest checkpoint here is sd-v1-4.ckpt

Set the maintenance expectation first, because it determines what this repository is for. The default branch is main and the last push to it was on 2024-06-18. The repository has no GitHub releases, so there is no tag to pin and no release note to read. What you have instead is a fixed set of four published checkpoints, described in the README as sd-v1-1.ckpt, sd-v1-2.ckpt, sd-v1-3.ckpt and sd-v1-4.ckpt.

That set has a lineage worth knowing, because the names alone do not tell you which training data produced the images you get. sd-v1-1 was trained for 237k steps at 256x256 on laion2B-en and then continued for 194k steps at 512x512 on laion-high-resolution, a set of 170M examples drawn from LAION-5B at resolution 1024x1024 or above. sd-v1-2 resumed from sd-v1-1 and ran 515k further steps on a filtered subset. sd-v1-3 and sd-v1-4 both resumed from sd-v1-2, at 195k and 225k steps respectively, both adding 10 percent dropping of the text conditioning to improve classifier-free guidance sampling.

So a reader who installs this today gets a model from 2024 and a weight lineage frozen at v1.4. Anything newer is not in this repository, and the README's own pointer for where movement happens is the separate diffusers integration, which it says it expects to see more active community development.

setup.py names the package latent-diffusion at version 0.0.1

The editable install does not install anything called Stable Diffusion. The packaging file is short and unambiguous:

python
    name='latent-diffusion',
    version='0.0.1',
    install_requires=[
        'torch',

Three consequences follow. The importable and installable name is latent-diffusion, inherited from the earlier latent-diffusion project this work builds on, so a pip freeze in your environment will not show a line called stable-diffusion and a script looking for one will not find it. The version is 0.0.1 and has no relationship to the checkpoint numbering, so there is no way to ask pip which of the four weights a given environment was set up for. And the declared dependencies are only torch, numpy and tqdm.

That third point is the one that bites. The README's own update instructions install transformers, diffusers and invisible-watermark as a separate pip line. Running pip install -e . on its own satisfies the packaging metadata and leaves the three packages the sampling path actually imports missing, with no error to tell you which ones.

The install is a conda file plus a hard pin on transformers 4.19.2

There are two documented paths, and they are not the same thing. The first creates a fresh environment from the repository's own file and names it:

bash
conda env create -f environment.yaml
conda activate ldm

The second is explicitly for updating an environment that already exists, and it is the one that carries the pin:

bash
conda install pytorch torchvision -c pytorch
pip install transformers==4.19.2 diffusers invisible-watermark
pip install -e .

Read the difference carefully before you run either. The first path is a fresh environment named ldm. The second is for a latent diffusion environment you already have, and the exact pin on transformers is 4.19.2, with no upper or lower range around it. If you follow the update path inside an environment that also serves something else, that single equality constraint is the most likely line in the whole procedure to fail or to force a downgrade of an unrelated package.

The reason for the split is that environment.yaml and the pip line are not redundant. The file gives you the working set; the pip line adds diffusers for the alternate integration and invisible-watermark for the watermarking the sampling script applies.

The weights are research artifacts under a licence that permits commercial use

The licence is the CreativeML OpenRAIL M license in the LICENSE file at the repository root, adapted from the work BigScience and the RAIL Initiative carry together. It is permissive in form and carries use-based restrictions, and commercial use is permitted under its terms.

The same paragraph then says something that matters more than the permission. The project does not recommend using the provided weights for services or products without additional safety mechanisms and considerations, because the weights have known limitations and biases, and because research on safe deployment of general text-to-image models is ongoing. The final sentence is the one to quote back to a stakeholder: the weights are research artifacts and should be treated as such.

That is a genuine tension inside one document, and it is not resolved anywhere else in the repository. The bias discussion is delegated to Stable_Diffusion_v1_Model_Card.md, which the README also points to for training details and intended use, and the README separately notes that v1 is a general text-to-image model and therefore mirrors biases and misconceptions present in its training data. For an engineer, the practical reading is that the licence clears you legally and the project is telling you it has not cleared you technically. Any decision to put these weights behind a product needs work the repository asks you to supply.

The reference script watermarks its own output and checks for explicit results

Two behaviours are built into the sampling path rather than offered as options. The reference sampling script incorporates a Safety Checker Module whose stated purpose is to reduce the probability of explicit outputs, and it applies invisible watermarking to the outputs so that viewers can identify the images as machine-generated.

Reduce is the operative word in both cases. The checker lowers a probability rather than refusing an input, and the watermark marks a picture rather than stopping one. No flag that turns either off appears anywhere in the quickstart, so anyone who needs unmarked output is editing the reference script, and anyone who needs a hard refusal is adding logic the project does not provide.

For an organisation that has to answer questions about image provenance, the watermark is the more consequential of the two, and it is worth being clear that it is not a separate tool you opt into. It ships in scripts/txt2img.py as documented, which means output from the reference path is marked whether or not that is what you wanted. The dependency that makes it work, invisible-watermark, is installed by the pip line in the update path, so it is present in the environment the README builds.

Ten gigabytes of VRAM, 512x512, fifty PLMS steps, scale 7.5

These are the only performance figures committed to anywhere in the repository, and together they describe the entire stated envelope. The model is an 860M UNet with a 123M text encoder, described as relatively lightweight, running on a GPU with at least 10GB VRAM. The default sampling settings are a guidance scale of 7.5, Katherine Crowson's implementation of the PLMS sampler, and 50 steps, rendering 512x512 images, the size the model was trained on.

The architecture detail that matters for reproducing the model is the downsampling-factor 8 autoencoder. Stable Diffusion v1 is defined as a specific configuration using that autoencoder with the 860M UNet and the CLIP ViT-L/14 text encoder, pretrained at 256x256 and finetuned at 512x512, with the text encoder frozen. The model is conditioned on the non-pooled text embeddings of that encoder, which is the difference from an implementation that pools them.

What the README does not give is any timing, any throughput figure, any figure for a batch size other than one, and any statement about what happens below 10GB of VRAM. If you are sizing hardware, this page is the whole of the project's own guidance, and there is no lower bound to fall back on.

Sampling means linking a checkpoint by hand, then calling one script

The first-run sequence is short and entirely manual. After obtaining the stable-diffusion-v1-*-original weights, the checkpoint is symlinked into a fixed path that the scripts expect:

bash
mkdir -p models/ldm/stable-diffusion-v1/
ln -s <path/to/model.ckpt> models/ldm/stable-diffusion-v1/model.ckpt

Generation is then a single call, and the prompt in the example is the one the README uses:

bash
python scripts/txt2img.py --prompt "a photograph of an astronaut riding a horse" --plms

Two things about that arrangement are easy to miss. The path is hard-coded to a single model directory, so keeping several checkpoints around means editing the link between runs, and no flag for choosing a checkpoint is offered. And the --plms flag selects the sampler rather than being implied, even though PLMS is what the defaults describe, which tells you the script exposes the sampler as a choice.

The README also names a second way in, a diffusers integration, and says it expects to see more active community development there. That is a useful signal about where the moving parts are, but it also splits the documentation: the reference script is the path this repository explains, and the path it expects to be superseded is the one it explains less.

Editorial conclusion

Use this repository when you want the original v1 model with its documented checkpoint lineage and a sampling script you can read, not when you want the current text-to-image capability, since the last push was 2024-06-18 and the newest checkpoint here is sd-v1-4.ckpt. Read the CreativeML OpenRAIL M licence and Stable_Diffusion_v1_Model_Card.md before any product decision, because the terms permit commercial use while the project explicitly declines to recommend the weights for services or products without added safety work. Verify first that you can satisfy the 10GB VRAM floor and the transformers 4.19.2 pin, since those are the two constraints committed to anywhere in the repository and environment.yaml does not relax either.

Frequently asked questions

What are stable diffusions?

Stable Diffusion v1 is a latent text-to-image diffusion model built from a downsampling-factor 8 autoencoder with an 860M UNet and a frozen CLIP ViT-L/14 text encoder. It was pretrained on 256x256 images and then finetuned on 512x512, on a subset of the LAION-5B database.

How safe is Stable Diffusion?

The reference sampling script includes a Safety Checker Module to reduce the probability of explicit outputs, and it watermarks outputs so viewers can identify them as machine-generated. The project also states that the weights mirror biases and misconceptions present in their training data, and that they are research artifacts which should be treated as such.

how to install stable diffusion

The README's path is a conda environment created from environment.yaml and named ldm, followed by an editable install of the package. It also gives a separate update route for an existing latent diffusion environment, which pins transformers to 4.19.2 and adds diffusers and invisible-watermark.

how to use stable diffusion locally

You symlink the checkpoint into models/ldm/stable-diffusion-v1/ and then run scripts/txt2img.py with a prompt. The model needs a GPU with at least 10GB VRAM, and the documented defaults are a guidance scale of 7.5, 50 PLMS sampling steps and 512x512 output.

Can I use Stable Diffusion for free?

The weights are published through the CompVis organization on Hugging Face under the CreativeML OpenRAIL M license, which permits commercial use while carrying use-based restrictions. The project states that the weights are research artifacts and does not recommend using them for services or products without additional safety mechanisms and considerations.

Official sources

  1. Official documentation
  2. Official README
  3. Project repository
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/compvis-stable-diffusion.svg)](https://hysenlabs.com/projects/compvis-stable-diffusion)