Model or dataset
rupeshs/fastsdcpu avatar
rupeshs/fastsdcpu

FastSD CPU: Stable Diffusion without a GPU, and what it costs

Fast stable diffusion on CPU and AI PC

2,167 stars217 forksPythonMIT

At a glance

What is it?
Latent consistency distillation and OpenVINO make sub-second 512x512 generation on an Intel Core i7 plausible. The project has also quietly accumulated a 1-bit quantization path, an MCP server, a GIMP plugin and a dependency stack that reveals where it really runs.
Who is it for?
FastSD CPU is one of the few diffusion projects where the hardware assumption is inverted on purpose rather than as a fallback. Everything in its architecture follows from that: one to four step inference instead of fifty, TAESD instead of the full VAE decoder, OpenVINO instead of PyTorch eager execution, and GGUF quantization down to one bit.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 75 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 7, 2026, and from our analysis. They are not legal advice.

Editorial analysis

Two compression techniques, and why one step is possible

The project describes itself as a faster version of Stable Diffusion on CPU, and names its two foundations: Latent Consistency Models and Adversarial Diffusion Distillation. Both do the same job from different directions, which is why they can be combined.

A standard diffusion sampler runs the same network repeatedly, stepping from noise toward an image over dozens of iterations. Latent consistency training produces a model that predicts the final latent in a single forward pass rather than a trajectory leading to it, so inference collapses to a handful of steps. Adversarial distillation attacks the same problem from a game-theoretic angle, training the model against a discriminator so it lands on its output rather than toward it. The project's own writing on fast stable diffusion on CPU describes the pairing as the reason a sub-second 512x512 generation is reachable on a processor.

The practical consequence is the speed figure the README leads with: 0.82 seconds for a single 512x512 image on a Core i7-12700, using OpenVINO with the SDXS-512-0.9 model. That number is real and specific, which is better than a vague claim about being fast, but it carries a caveat that the project itself documents further down. The default model was changed to SDXS-512-0.9 and then reverted to SDTurbo, so the configuration that produced the headline timing is not the one you get on a fresh install.

There is a longer arc in the feature list about the same trade. TAESD, a tiny autoencoder, replaces the SD decoder and is credited with a 1.4x speed boost at moderate quality. Tiny AutoEncoder 1.3 support came later, under the name Mocha Croissant. Combined with OpenVINO at two steps and a tiny decoder, the README claims 5.7x over the baseline path. Every one of these multipliers is quality given up, and the README is honest that the quality is moderate rather than claiming otherwise.

Memory is the gate, and the table says so

Most projects bury memory requirements in an issue thread. This one puts a table at the top of the README, and it is the first thing worth reading.

The reference configuration for the first two rows is SD Turbo at one step and 512x512, and Dreamshaper v8 at three steps for LCM-LoRA. Against that: LCM needs 2 GB of system RAM, LCM-LoRA needs 4 GB, OpenVINO with Flux2 needs 8 GB, and OpenVINO generally needs 11 GB. Enabling the tiny decoder saves roughly 2 GB, so OpenVINO mode drops to about 9 GB.

One more line in that section matters more than the numbers: guidance scale above 1 increases both RAM usage and inference time. That is a genuine tension in the feature list, where OpenVINO models gained negative prompt support with the instruction to set guidance above 1.0. So enabling the feature costs you the two things you are using the project for. Whether that trade is worth it depends entirely on whether your prompt depends on negative guidance.

The table also explains the platform list. Windows, Linux and Mac are the mainstream targets, and Android with Termux plus PRoot and Raspberry Pi 4 are listed as supported, which follows from a 2 GB LCM ceiling being reachable on a Pi. Requirements are Python 3.10 or higher and uv as the package manager.

A dependency detail worth flagging, since it constrains the Python version more tightly than the stated floor suggests. The pinned requirements include numpy 1.26.4 alongside transformers 5.0.0, diffusers 0.37.0 and openvino 2026.2.0. numpy 1.26 predates the newest CPython releases and ships no wheels for them, so on an interpreter where that pin has to build from source, installation is where you will find out. The bundled Dockerfile pins the problem in place by using python3.11-trixie-slim, which is a good reason to prefer the container over a local environment if you are on a very new Python.

Three interfaces with different jobs

The interfaces are split by capability, not by preference. The Qt desktop GUI covers basic text to image and is described as faster. The Gradio-based WebUI carries the advanced features: LoRA, ControlNet, image to image. The CLI is the third option.

That division shows up in the launch scripts at the repository root, which are split into Windows batch files and shell scripts for the main app, the web UI, the web server, the realtime mode and the MCP server. The realtime mode generates images while you type and is labelled experimental in the table of contents.

Feature coverage by interface is uneven, which is the thing to check before picking one. ControlNet v1.1 support exists in LCM-LoRA mode with a set of annotators listed: Canny, Depth, LineArt, MLSD, NormalBAE, Pose, SoftEdge and Shuffle. Multiple LoRA support and basic LoRA support in the CLI and WebUI are both listed. Image to image, including Turbo model support in both PyTorch and OpenVINO, is described as a WebUI feature. SDXL-Lightning has both a PyTorch and an int8 OpenVINO path. HyperSD and SDXL and SD 1.5 single-step support are all present.

Upscaling is a separate cluster worth naming: a 2x EDSR in ONNX, tiled SD upscale marked experimental, Aura SR at 4x, Aura SR v2, and GigaGAN. The CLIP skip and token merging options exist for prompt adherence at some cost in speed.

Maximum inference steps was raised to 25, which is a telling number in a project built around one to four steps. It exists so you can dial quality upward when you have time, and the guidance scale warning applies to the same knob.

One default deserves a plain mention rather than a euphemism: the safety checker is disabled by default, though a setting to control it was added. Whatever you think of that default, it is a deliberate choice you are accepting on install rather than discovering later.

GGUF quantization down to one bit, and the models it applies to

The 2026-07-05 release added bonsai image GGUF support and Flux2 Klein GGUF support, with quantization levels from 4-bit down to 1-bit. The release notes link a bonsai image model described as 1-bit.

This is the part of the project where the usual relationship between size and quality stops being smooth. Four-bit quantization on a diffusion model is aggressive but defensible. One-bit quantization on an image model is not a rounding of the same idea; it means the weights have been reduced to something close to a single bit per weight, and the README provides no quality measurement for it. No accuracy figure, no comparison images, no note about which layers were excluded. The honest way to treat the 1-bit claim is as an experiment worth running rather than a setting worth enabling by default.

Earlier in the same line of development, GGUF support for Flux arrived on 2024-10-02, and a 1-bit bonsai model plus Flux2 Klein on 2026-07-05. There is also a separate tiny autoencoder for FLUX.1, TAEF1, published on Hugging Face as an OpenVINO export, which sits in the same family of ideas as TAESD: replace the expensive decoder with a cheap one.

Model coverage has widened substantially since the 2024 baseline. SDXL Turbo and SD1B 1B LCM models are there, LCM-LoRA works with fine-tuned SD 1.5 and SDXL models and can be configured from a text configuration file, and experimental support exists for single-file Safetensors models from Civitai by adding a local path to `configs/stable-diffusion-models.txt`. Safety in loading is not discussed anywhere in the README, so a local model file from an unfamiliar source deserves the same caution you would apply to any executable.

The June 2026 release also brought image editing with prompt presets, and those four presets are more revealing than they look: restore old photo, restore and colorize old photo, colorize photo, enhance photo. That is a photo restoration product description arriving inside a text-to-image tool, at two to four steps.

Deployment surfaces: Docker, MCP, GIMP, Home Assistant of AI PCs

The project has grown well past a local script, and the tree shows it: a Dockerfile, a docker-compose.yml, a `.dockerignore`, a THIRD-PARTY-LICENSES file, and a docs directory.

Docker support arrived on 2026-05-01 alongside a Hugging Face minimal demo. The compose file exposes a single service on port 7860 with the Hugging Face caches pointed at a named volume, which is the right shape for this workload since model weights are the bulk of the disk:

yaml
    ports:
      - "7860:7860"

    environment:
      - HF_HOME=/data/huggingface
      - TRANSFORMERS_CACHE=/data/huggingface/transformers
      - HF_DATASETS_CACHE=/data/huggingface/datasets
      - PYTHONUNBUFFERED=1

    volumes:
      - hf_cache:/data/huggingface

And the image builds on uv with CPU-only torch from the PyTorch CPU index, then layers requirements on top:

dockerfile
FROM ghcr.io/astral-sh/uv:python3.11-trixie-slim

RUN apt-get update && apt-get install -y \
    git curl build-essential libgl1 libglib2.0-0 \
    && rm -rf /var/lib/apt/lists/*

The MCP server support, added 2025-04-20 along with faster uv-based installation, Claude desktop and Open WebUI support, means the model can be driven as a tool rather than only through a UI. Both a `start-mcpserver.sh` and a `.bat` variant exist at the root, so the Windows path is first class rather than an afterthought.

The integration that surprised me most is the GIMP route: on 2025-12-22 the FastSD engine was integrated into Intel's OpenVINO AI Plugins for GIMP. That is a third-party endorsement of the engine specifically, not a feature this project added, and it is a different kind of adoption than an MCP server because it puts generation inside an existing image editor workflow.

Hardware support widened on the same axis. Intel AI PC GPU and NPU support arrived 2024-09-03, and Intel Core Ultra Series 2 Lunar Lake NPU support on 2024-11-03. The NPU path is described as power-efficient inference and covers text to image, image to image and image variations, with a device check added later. Integrated GPU support uses OpenVINO with `DEVICE=GPU`.

The ComfyUI support is listed in the table of contents as well, which matters for anyone who already has a node graph they do not want to rebuild.

Editorial conclusion

FastSD CPU is one of the few diffusion projects where the hardware assumption is inverted on purpose rather than as a fallback. Everything in its architecture follows from that: one to four step inference instead of fifty, TAESD instead of the full VAE decoder, OpenVINO instead of PyTorch eager execution, and GGUF quantization down to one bit. Those choices trade image fidelity for latency and for a memory ceiling low enough to run on a Raspberry Pi 4, which is a real trade rather than a marketing one. Four things are worth settling before you build on it. The headline 0.82 second figure describes SDXS-512-0.9, which was made the default and then reverted, so the number and the out-of-box behaviour no longer match. The safety checker is disabled by default, which is a deliberate choice you should confirm you want. The Python floor is documented as 3.10 or higher while the dependency set includes a numpy pin that has no wheels for the newest interpreters. And the memory table is the honest ceiling: 11 GB for OpenVINO mode without a tiny decoder, 9 GB with one. Read that table before the interface, because it decides whether a given machine is in scope at all.

Frequently asked questions

What is FastSD CPU?

FastSD CPU is a Python project that runs Stable Diffusion, SDXL and Flux models on a processor rather than a discrete GPU. It applies Latent Consistency Models and Adversarial Diffusion Distillation so that generation takes one to four steps instead of dozens, and adds OpenVINO for execution. The README reports 0.82 seconds for a 512x512 image on a Core i7-12700 using OpenVINO with SDXS-512-0.9. It is MIT licensed and ships a Qt desktop GUI, a Gradio web UI and a CLI.

Do AI models run on CPU or GPU?

Both, and the choice is usually about memory and speed rather than possibility. FastSD CPU is built for the CPU case: it reduces diffusion to one to four inference steps, swaps in a tiny autoencoder decoder, and uses OpenVINO so the same work runs on an integrated GPU or NPU. The README reports 11 GB of system RAM for OpenVINO mode, dropping to about 9 GB with the tiny decoder enabled, and 2 GB for plain LCM mode.

What is CPU in AI?

In this context the CPU is the general purpose processor of the machine, as distinct from a GPU or NPU that is built for parallel arithmetic. FastSD CPU takes the position that a general processor is enough for image generation if the model is distilled down to very few steps and executed through a runtime that targets the hardware, which is why the project reports sub-second 512x512 results on a Core i7 and lists Raspberry Pi 4 and Android with Termux as supported platforms.

Official sources

  1. Issues
  2. License: MIT
  3. README
  4. Releases
  5. rupeshs/fastsdcpu on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/rupeshs-fastsdcpu.svg)](https://hysenlabs.com/projects/rupeshs-fastsdcpu)