Open-source project
apple-aiml-research/ml-stable-diffusion avatar
apple-aiml-research/ml-stable-diffusion

apple/ml-stable-diffusion: Core ML Stable Diffusion for Apple Silicon

Stable Diffusion with Core ML on Apple Silicon

17,966 stars1,081 forksPythonMIT

At a glance

What is it?
Apple's repository converts PyTorch Stable Diffusion checkpoints into Core ML packages and ships a Swift pipeline for iOS and macOS apps. It is a conversion and deployment toolchain, not a desktop image generator, and the README is explicit about the hardware and OS floors.
Who is it for?
Adopt it if you are shipping image generation inside an iOS, iPadOS or macOS app and you need the model to run on the Neural Engine rather than in a server round trip. Skip it if you want a prompt box on a Mac today: this is a conversion and deployment toolchain, and the README points at the Hugging Face Diffusers app for that use.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 19 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What apple/ml-stable-diffusion actually ships

The repository contains two distinct pieces. The first is `python_coreml_stable_diffusion`, a Python package that converts PyTorch models into Core ML format and then generates images through Hugging Face diffusers in Python. The second is `StableDiffusion`, a Swift package meant to be added to an Xcode project as a dependency. The Swift package does not carry model weights; it consumes the Core ML files that the Python side produced. That split matters when you plan a build pipeline, because the conversion step and the shipping step have different requirements and can run on different machines. The README describes the project as Stable Diffusion with Core ML on Apple Silicon, and the setup.py description reads "Run Stable Diffusion on Apple Silicon with Core ML (Python and Swift)". The intended audience is developers embedding generation in an app, not end users looking for a chat-style interface.

The conversion path from PyTorch checkpoint to .mlpackage

The data flow is one-directional. You start from a PyTorch checkpoint, run the conversion in `python_coreml_stable_diffusion`, and get Core ML model files. The Swift package then loads those files at runtime. Two flags shape the output. `--compute-unit` selects where inference runs and, per the README's benchmark notes, does not modify the Core ML model, so it can be changed at runtime. `--attention-implementation` does modify the model, so it is baked in at conversion time. That asymmetry is the single most important thing to understand before you convert, because changing attention later means re-running conversion. The README also notes that the performance optimizations in the repository are generally applicable to Transformers rather than tuned specifically for Stable Diffusion, and that better performance may be observed with custom kernel tuning. In other words, the shipped numbers are not a ceiling. The text encoder is converted with a static shape covering all 77 tokens regardless of prompt length, so short prompts cost the same as long ones.

Installing python_coreml_stable_diffusion and converting a first model

The README points at the System Requirements section before installation. For model conversion it lists macOS 13.1, Python 3.8 and coremltools 7.0; the package metadata in setup.py and requirements.txt pins `coremltools>=8.0`, `diffusers[torch]==0.30.2`, `transformers==4.44.2` and `numpy<1.24`. Install from the repository root:

bash
pip install -r requirements.txt
pip install -e .

The editable install exposes the conversion entry point. The README's conversion workflow is driven by a Python module, and the compute unit and attention implementation are passed as flags:

bash
python -m python_coreml_stable_diffusion.torch2coreml \
  --compute-unit CPU_AND_NE \
  --attention-implementation SPLIT_EINSUM_V2

What you should see is a set of Core ML model files on disk, ready to be handed to the Swift package. Verify the output before wiring it into an app, because a conversion that silently picks the wrong attention implementation will only show up as a latency surprise later. The README does not document a rollback path for a conversion, so keep the source checkpoint.

Hardware and OS floors you cannot design around

The System Requirements table is unusually specific and worth reading as a constraint list rather than a formality. Model conversion needs macOS 13.1, Python 3.8 and coremltools 7.0. Building the project needs macOS 13.1, Xcode 14.3 and Swift 5.8. Target device runtime needs macOS 13.1 or iPadOS and iOS 16.2. The memory improvements tied to 6-bit compression raise those floors to macOS 14.0 and iPadOS or iOS 17.0. Minimum hardware is an M1 Mac, an M1 iPad or an A14 iPhone. If your deployment target sits below iOS 17.0, you are choosing between the compression improvements and your install base. The README does not offer a middle path here, and that is a real product decision rather than a footnote.

What the benchmark tables say, and what they do not

The performance tables are the most quoted part of the README and the easiest to misread. For `stabilityai/stable-diffusion-2-1-base` at 512x512, Apple and Hugging Face report median latency across five back-to-back end-to-end executions, measured in August 2023 on public beta versions of iOS 17.0, iPadOS 17.0 and macOS 14.0 Seed 8, using 20 inference steps, 77 text tokens and classifier-free guidance with a unet batch size of 2. Weights are compressed to 6 bits; activations are float16 on both GPU and Neural Engine. The Swift code is described as not fully optimized, adding up to roughly 10 percent overhead unrelated to Core ML execution. An asterisk marks runs where `reduceMemory` was enabled, which loads and unloads models just-in-time and added up to 2 seconds. The README states plainly that these numbers do not represent peak hardware capability. For the larger 768x768 SDXL model the picture changes sharply: the fastest listed device, iPhone 15 Pro Max, is reported at 31 seconds end to end, and several older devices sit above a minute. That is the honest boundary of this approach on phones.

Where this is the wrong tool

If you want to type a prompt on a Mac and get an image, this repository is the wrong entry point. It gives you a conversion package and a Swift library, not an application, and the README points elsewhere for a ready-made experience. The same applies if your target hardware is older than the stated floors: an A13 iPhone or an Intel Mac is outside the supported matrix, and nothing in the repository suggests a fallback. There is also a conversion cost that does not appear in the latency tables. Every checkpoint you support has to be converted, and because `--attention-implementation` changes the model files, supporting two attention implementations doubles the artifacts you ship. Teams that expect to swap models at runtime will find that the model, not the app, is the unit of release. Finally, the README does not document rollback or versioning for converted artifacts, so provenance is on you.

How it differs from a diffusers-on-MPS setup

The obvious alternative for Mac users is running diffusers directly against PyTorch with MPS acceleration, which is what a typical local Python setup does. The difference in approach is where the work happens. A diffusers-on-MPS setup keeps the PyTorch graph and executes it through Metal Performance Shaders on the GPU. This repository instead compiles the graph into Core ML and lets the runtime schedule work across CPU, GPU and Neural Engine. The practical consequences follow from that. Core ML artifacts are static-shaped, which is why the text encoder always processes 77 tokens. Core ML artifacts are also what an iOS app can load without shipping a Python runtime, which is the entire reason the Swift package exists. If your output is a Python script on a laptop, diffusers-on-MPS is simpler. If your output is an app on the App Store, the Core ML route is the one that fits the platform.

Licence, maintenance and upgrade cost

The repository is MIT licensed, which permits commercial use and modification provided the copyright notice and permission notice are retained. That is a permissive licence, but it covers the code in this repository, not the model weights you convert. Checkpoint licences come from their own publishers and are a separate question; nothing here changes them. Maintenance signals are mixed. The last push was on 2026-09-11, which is recent, but the latest release is 1.1.1 from 2024-05-04, and the package metadata still classifies the project as "Development Status :: 4 - Beta". The dependency pins are tight and deliberate: `diffusers[torch]==0.30.2`, `transformers==4.44.2` and `numpy<1.24`. Those exact pins are what the conversion was validated against, and loosening them is your risk to take. Upgrading coremltools, diffusers or transformers is not a routine bump here, because a version change can alter the converted graph.

Editorial conclusion

Adopt it if you are shipping image generation inside an iOS, iPadOS or macOS app and you need the model to run on the Neural Engine rather than in a server round trip. Skip it if you want a prompt box on a Mac today: this is a conversion and deployment toolchain, and the README points at the Hugging Face Diffusers app for that use. Before committing, verify that your target device generation and OS version clear the floors in the System Requirements section, and confirm that the checkpoint you want has a conversion recipe in the repository.

Frequently asked questions

Can I run Stable Diffusion on a Mac with apple/ml-stable-diffusion?

Yes, for the Python side. The README lists macOS 13.1 as the target device runtime and an M1 Mac as the minimum hardware generation, and the Python package performs image generation through Hugging Face diffusers. The repository also ships a Swift package for embedding generation in iOS, iPadOS and macOS apps.

Is Stable Diffusion still a thing in this repository?

The project is still present and was pushed to on 2026-09-11, but the latest release is 1.1.1 from 2024-05-04 and setup.py classifies it as Development Status 4 - Beta. The README covers both Stable Diffusion 2.1 base and SDXL base 1.0 for iOS.

Does apple/ml-stable-diffusion cost anything to use?

The repository is MIT licensed, so the code can be used and modified commercially as long as the copyright and permission notices are kept. The model weights you convert carry their own licences from their publishers, which this licence does not cover.

Which iPhone and iOS version does apple/ml-stable-diffusion need?

The README lists an A14 iPhone as the minimum hardware generation and iOS 16.2 as the target device runtime. Using the memory improvements tied to 6-bit compression raises the requirement to iOS 17.0.

Can I change the attention implementation after converting a model?

No. The README states that `--attention-implementation` modifies the Core ML model, while `--compute-unit` does not and can be applied at runtime. Changing attention means converting again.

Official sources

  1. apple-aiml-research/ml-stable-diffusion on GitHub
  2. Issues
  3. License: MIT
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/apple-aiml-research-ml-stable-diffusion.svg)](https://hysenlabs.com/projects/apple-aiml-research-ml-stable-diffusion)