# BentoDiffusion: Deploy Stable Diffusion and FLUX Models as Production APIs with BentoML

> BentoDiffusion is a collection of BentoML example projects that shows how to serve FLUX.1, Stable Diffusion 3.5, SDXL, and ControlNet as HTTP endpoints rather than interactive UIs. It is aimed at engineers who need to call image generation from an application backend rather than through a browser interface.

**bentoml/BentoDiffusion** — BentoDiffusion: A collection of diffusion models served with BentoML

- Repository: https://github.com/bentoml/BentoDiffusion
- Website: https://bentoml.com
- Stars: 389 · Forks: 29
- Language: Python
- License: Apache-2.0
- Published: 2026-09-17 · Updated: 2026-09-17 · Language: en
- Canonical page: https://hysenlabs.com/projects/bentoml-bentodiffusion

## What BentoDiffusion Is and Who It Serves

BentoDiffusion is not a standalone diffusion application. It is a set of example projects, each in its own subdirectory, showing how to wrap a model from the Stable Diffusion family inside a BentoML Service and expose it as an HTTP endpoint. The README describes the Stable Diffusion family as models specialised in generating and manipulating images or video clips from text prompts.

The audience is engineers integrating text-to-image generation into backend services. Because each project exposes a REST endpoint with an OpenAPI spec at `/swagger`, the generated image is returned programmatically, making it straightforward to call from application code without building a separate UI layer. The Swagger interface at `localhost:3000` is also available for manual testing during local development. The repository uses SDXL Turbo as its walkthrough example but covers seven distinct models, from the fast single-step SDXL Lightning to the higher-quality Stable Diffusion 3.5 Large.

## Repository Layout and the Seven Model Variants

The repository groups examples by model in top-level subdirectories. The README lists the seven available variants:

- `flux-timestep-distilled/` for FLUX.1
- `sd3-medium/` for Stable Diffusion 3 Medium
- `sd3.5-large-turbo/` for Stable Diffusion 3.5 Large Turbo
- `sd3.5-large/` for Stable Diffusion 3.5 Large
- `sdxl-lightning/` for Stable Diffusion XL Lightning
- `sdxl-turbo/` for Stable Diffusion XL Turbo
- `controlnet/` for ControlNet

Each subdirectory contains its own `requirements.txt` and a `service.py` defining the BentoML Service. The walkthrough in the README uses `sdxl-turbo` as the example throughout. Detailed explanations of the service code for each variant live in the BentoML documentation rather than inline in the repository.

## Running SDXL Turbo Locally: Install and Serve

The README recommends Python 3.11 and an Nvidia GPU with at least 12 GB VRAM for local testing. To set up the SDXL Turbo example:

```bash
git clone https://github.com/bentoml/BentoDiffusion.git
cd BentoDiffusion/sdxl-turbo
pip install -r requirements.txt
```

A `service.py` file in the subdirectory defines the BentoML Service. Start it with:

```bash
bentoml serve
```

The server listens on `http://localhost:3000`. The README shows the startup log output and notes that the Swagger UI is available at that address for interactive testing. The pipeline loading progress appears in the terminal before the server accepts requests.

## Calling the Endpoint from Code or curl

Once the server is running, the `txt2img` endpoint accepts a POST request with a JSON body. The README provides a curl example:

```bash
curl -X 'POST' \
  'http://localhost:3000/txt2img' \
  -H 'accept: image/*' \
  -H 'Content-Type: application/json' \
  -d '{"prompt": "A cinematic shot of a baby racoon wearing an intricate italian priest robe.", "num_inference_steps": 1, "guidance_scale": 0}'
```

The BentoML Python SDK provides a typed client for the same call:

```python
import bentoml

with bentoml.SyncHTTPClient("http://localhost:3000") as client:
        result = client.txt2img(
            prompt="A cinematic shot of a baby racoon wearing an intricate italian priest robe.",
            num_inference_steps=1,
            guidance_scale=0.0
        )
```

The client receives the generated image as a return value. The SDXL Turbo example uses `num_inference_steps=1` and `guidance_scale=0` because the model is a distilled variant designed for single-step inference.

## Deploying to BentoCloud

The README covers a two-step path for deploying to BentoCloud, the managed hosting service. First, log in from the CLI:

```bash
bentoml cloud login
```

Then deploy from the project directory:

```bash
bentoml deploy
```

BentoCloud handles infrastructure, scaling, and the exposed URL. The README notes that a BentoCloud account is required, with a sign-up link at bentoml.com. For custom deployments on private infrastructure, the README refers to BentoML's OCI image packaging documentation, which builds a container image from the Service definition.

## Limitations: What BentoDiffusion Is Not

BentoDiffusion is an example collection, not a production-ready library. Each subdirectory pins its own dependencies in a `requirements.txt`, but there are no GitHub releases and no versioned packages on PyPI. Upgrading to a newer model variant means switching subdirectory and reinstalling dependencies manually.

The 12 GB VRAM requirement stated in the README applies to local testing of SDXL Turbo. Larger models in the collection, such as Stable Diffusion 3.5 Large, will require more VRAM. The repository does not document minimum hardware requirements for each variant beyond that single note, so engineers should consult the BentoML documentation for each model's specifications before provisioning hardware.

The examples are each a starting point rather than a complete production deployment. Error handling, authentication, rate limiting, and storage of generated images are outside what the example services demonstrate. The `service.py` files show the BentoML Service pattern, inputs, and outputs, but not the surrounding infrastructure a real deployment needs. Engineers building production pipelines will need to add those layers on top of the template.

ComfyUI is a widely-used alternative for running diffusion models locally. Where BentoDiffusion produces an HTTP API endpoint intended for backend integration, ComfyUI offers a browser-based node workflow editor for interactive image generation, making it better suited to visual exploration than to programmatic pipelines. The two tools address different use cases: BentoDiffusion targets services that need to call generation from application code, ComfyUI targets users who iterate on prompts and workflow configurations manually.

## Maintenance and License

The repository's last push was on 2026-07-14. The code is licensed under the Apache-2.0 License, which permits commercial use, modification, and distribution.

The repository has no GitHub releases. BentoML's broader example catalog, linked in the README at docs.bentoml.com, provides context for where BentoDiffusion fits among the other available project types. Because this repository acts as examples rather than a library, updates tend to track the BentoML SDK's breaking changes and new model releases rather than following a regular release schedule.

For teams that need a container-based deployment, the README refers to BentoML's OCI image packaging guide, which covers building a portable image from a `service.py` definition. This means the examples can be containerised for deployment to Kubernetes or any OCI-compatible runtime, not only to BentoCloud. The actual containerisation steps live in BentoML's own documentation rather than in this repository.

## Conclusion

BentoDiffusion is the right starting point for engineers who already use BentoML and want to expose FLUX.1, Stable Diffusion 3 or 3.5, SDXL, or ControlNet as HTTP services callable from application code. It is not a standalone product: each subdirectory is an example, not a maintained library, and there are no GitHub releases. Teams that want a browser-based generation tool, or whose hardware does not meet the 12 GB VRAM requirement, will find ComfyUI a better fit. The repository's last push was on 2026-07-14. Check the full BentoML example list at docs.bentoml.com before choosing a subdirectory, as the README links seven model variants and their configuration details live in the BentoML documentation.

## FAQ

### Which diffusion models does BentoDiffusion support?

The repository includes examples for FLUX.1 (timestep-distilled), Stable Diffusion 3 Medium, Stable Diffusion 3.5 Large, Stable Diffusion 3.5 Large Turbo, SDXL Lightning, SDXL Turbo, and ControlNet. Each model lives in its own subdirectory with a separate requirements file and service definition.

### Is BentoML free to use?

The BentoML framework is open source. BentoCloud, the managed deployment service mentioned in the README, requires an account; the README links to a sign-up page at bentoml.com but does not describe pricing. Self-hosting the same code avoids BentoCloud entirely.

### Does BentoDiffusion work without a GPU?

The README states that an Nvidia GPU with at least 12 GB VRAM is recommended for local testing, implying CPU-only setups are technically possible but not the intended configuration. No CPU performance data is provided in the repository.

## Sources

- [bentoml/BentoDiffusion on GitHub](https://github.com/bentoml/BentoDiffusion)
- [Issues](https://github.com/bentoml/BentoDiffusion/issues)
- [License: Apache-2.0](https://github.com/bentoml/BentoDiffusion/blob/main/LICENSE)
- [Project website](https://bentoml.com)
- [README](https://github.com/bentoml/BentoDiffusion/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/bentoml-bentodiffusion
