CLI tool
bghira/SimpleTuner avatar
bghira/SimpleTuner

SimpleTuner: a diffusion fine-tuning kit for image, video and audio models

A general fine-tuning kit geared toward image/video/audio diffusion models.

2,930 stars296 forksPythonAGPL-3.0

At a glance

What is it?
SimpleTuner is a Python training kit for diffusion models, with a web UI, Docker deployment and multi-GPU support. It suits teams that already know what they want to train and can read configuration files.
Who is it for?
Adopt SimpleTuner if you train diffusion models regularly, want one pipeline for image, video and audio, and are comfortable editing configuration rather than clicking through a wizard. Skip it if you need a one-click consumer trainer, if your workflow is built entirely around kohya or OneTrainer config files, or if AGPL-3.0 obligations conflict with how you ship your product.
Can I use it commercially?
Yes, with strict conditions. AGPL-3.0 is a network copyleft licence: if people use a modified version over a network, for example as a hosted service, you must offer them its source code under the same licence.
Is it still maintained?
Yes. The repository last received commits 2 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What SimpleTuner trains, and who it is aimed at

SimpleTuner is a fine-tuning kit for diffusion models. The README describes it as "geared towards simplicity, with a focus on making the code easily understood", and calls the codebase "a shared academic exercise". That framing matters: it is a training harness, not a product with a support contract.

The scope is wider than the name suggests. The feature list covers image, video and audio generative models through what the README calls a "unified pipeline", and the model table spans families including ACE-Step at 3.5B parameters, Auraflow at 6B, and Boogu-Image, alongside the Stable Diffusion lineage that the package description still references. Each row carries a license and a commercial-use column, which is unusual and useful: Anima is listed under CircleStone Labs Non-Commercial License v1.2 with commercial use marked No, while ACE-Step and Auraflow are Apache-2.0 with commercial use marked Yes. You can pick a base model partly on licensing grounds before you download a single weight.

The intended user is someone running training jobs on their own hardware, or on rented GPUs, who wants defaults that mostly work. The README states the aim is "good default settings for most use cases, so less tinkering is required". That is a claim about defaults, not about zero configuration. Every run still needs a dataset, captions and a config.

There is also a multi-user layer, which is the part most people will not expect. The README lists worker orchestration, SSO via LDAP/Active Directory or OIDC, role-based access control with four default roles, organizations and teams, quotas, a job queue with five priority levels, approval workflows, email notifications and audit logging. The README calls this a "complete multi-user training platform" and adds "free and open source, forever". For a lab or a small studio sharing a GPU box, that is a real difference from a single-user script.

How a SimpleTuner run is structured

The architecture visible from the repository is a Python package with a CLI entry point and a separate web application. The pyproject.toml declares four console scripts: simpletuner, simpletuner-train, simpletuner-configure and simpletuner-inference, all mapped into the simpletuner package or the st_cli module. So the same install gives you a training entry point, a configuration entry point and an inference entry point.

Training state is cached. The README lists "Advanced caching" as a core feature, with image, video, audio and caption embeddings written to disk. On a second run over the same dataset, the expensive encoding step is largely already done. This is the single biggest reason to keep your dataset directory stable between runs: move or rewrite it and you pay the encoding cost again.

Datasets with mixed dimensions are handled by aspect bucketing, which the README lists as a core feature supporting "varied image/video sizes and aspect ratios". Memory pressure is handled by two separate mechanisms: DeepSpeed, reached through Hugging Face Accelerate, for optimizer state offload, and FSDP2 for DTensor-based sharding and context parallelism. The README points to a dedicated DEEPSPEED.md and a FSDP2.md in the documentation directory rather than folding those into the main guide, which tells you they are not one-line switches.

The repository layout confirms the split between library and interface: simpletuner/ holds the package, st_cli.py is the CLI module, documentation/ holds the guides referenced throughout the README, and package.json exists purely for JavaScript tests of the web UI components, with Jest as its only dependency. The web UI is a real component with its own test suite, not a thin wrapper. The README also documents a custom tracker hook: drop an accelerate.GeneralTracker into simpletuner/custom-trackers and pass --report_to=custom-tracker with --custom_tracker=<name>.

On the data side, the README states plainly that no data leaves your machine except through opt-in flags: report_to, push_to_hub, or webhooks that you configure manually. That is a design statement worth taking at face value when you are training on client material.

Installing SimpleTuner and running a first job

The package requires Python 3.12 or newer, below 3.14, per the requires-python field in pyproject.toml. The setup.py detects the platform by probing for nvidia-smi, then rocm-smi, then ROCm environment variables, and falls back to CPU, so the install pulls different PyTorch builds depending on what it finds. Install from the repository rather than assuming a bare pip name, since the project's own installation instructions live in the documentation directory.

bash
git clone https://github.com/bghira/SimpleTuner.git
cd SimpleTuner
python -m venv .venv
source .venv/bin/activate
pip install -e .

After that, the four console scripts are on your path. The README points new users at the web UI tutorial or the command-line tutorial, and notes that a manually configured quick start without either interface is described in documentation/QUICKSTART.md.

bash
simpletuner-configure
simpletuner-train

The configure command is where a run is defined. The README does not reproduce the full option list in the main document; the tutorial and quick start files carry that.

If you would rather not manage Python dependencies at all, the repository ships a Dockerfile and a docker-compose.yml. The compose file builds from the local Dockerfile with PYTHON_VERSION "3.12", reserves all NVIDIA GPUs, mounts config, datasets and output into the container, mounts the host Hugging Face cache so models are not re-downloaded, and publishes two ports: 8001 for the web UI and 2222 for SSH. It sets shm_size to 16gb, which matters for dataloader workers.

bash
export HF_TOKEN=your_token_here
docker compose up -d

The compose file reads HF_TOKEN, WANDB_API_KEY, GIT_USER, GIT_PAT, GIT_REPOSITORY, TRAINING_NAME and USE_SSL from the environment. Once the container is up, the web UI is reachable on port 8001. The README's own tutorial path starts with the web UI, so that is the intended first contact for most people.

Where SimpleTuner gets in your way

The memory claims need reading carefully. The README says "Most models trainable on 24G GPU, many on 16G with optimizations". That is a statement about the model set as a whole, not a promise for every architecture in the table. A 6B model and a small SD-family model do not fit in the same envelope. The README does not publish a per-model memory figure in the section reproduced here, so the honest position is that you find out by trying, or by reading the quick start's compatibility section.

The multi-user platform is described in the README as living under documentation/experimental/server/. The directory name is the disclosure. If you build an internal training service on top of the worker orchestration and approval workflows, you are building on an API the project itself files under experimental, and releases are frequent enough that paths and behaviour can shift between minor versions. The release history shows v4.9.0, v4.9.1 and v4.9.2 inside a nine-day window, with v4.9.0 described as bringing torch compile speedup, MegaCache import/export, NextLat, Explorative Modeling and block swap optimisations. That is a fast-moving surface.

Python version support is narrow by design: 3.12 and 3.13 only. If your environment is pinned to 3.11 or earlier, SimpleTuner is not installable without changing that first. The package classifiers also mark the project as Development Status 4 - Beta, which is the maintainer's own label.

Finally, the README does not document rollback. There is no described procedure for reverting a training run, no checkpoint retention policy, and no statement about what happens to cached embeddings when you change the model or the resolution. If you need reproducible runs across months, you will be managing that discipline yourself.

SimpleTuner against kohya and OneTrainer

The two comparisons people actually search for are kohya and OneTrainer, and the difference is mostly about configuration surface and deployment shape.

kohya's scripts are the long-standing reference implementation for Stable Diffusion LoRA training, and their configuration style is a flat set of command-line flags and TOML files aimed at a single training run on a single machine. SimpleTuner keeps the config-file approach but wraps it in a package with a web UI, a job queue and a worker model. If you already have kohya configs you like, moving to SimpleTuner means re-expressing them in SimpleTuner's configuration system, not importing them.

OneTrainer is a desktop application with a graphical configuration workflow. SimpleTuner's web UI is a browser interface served on port 8001 from a process you run, usually in Docker or a virtualenv. On a headless rented GPU, the browser interface is reachable; a desktop application is not, unless you tunnel it. That is the practical split: OneTrainer optimises for a local workstation, SimpleTuner optimises for a machine you reach over the network.

The other axis is model coverage. SimpleTuner's table spans image, video and audio families, and the README frames the pipeline as multi-modal. If you train one SD-family LoRA occasionally, that breadth is overhead. If you train video or audio diffusion models, the alternatives are thinner and the aspect bucketing plus caching machinery starts paying for itself.

One feature worth calling out because it has no obvious equivalent in the alternatives: concept sliders, described in the README as "slider-friendly targeting for LoRA/LyCORIS/full" with positive, negative and neutral sampling and per-prompt strength. There is a dedicated SLIDER_LORA.md. If slider-style control is what you are producing, that is a reason to look here specifically.

Licence, maintenance and the cost of upgrading

SimpleTuner is AGPL-3.0-or-later. The pyproject.toml carries the SPDX identifier and the classifier list matches. AGPL is the network-copyleft variant: if you modify SimpleTuner and let users interact with it over a network, the licence's source-availability obligations are generally understood to reach that modified version. Whether that matters depends entirely on whether you are training models for your own use, which is the common case, or shipping SimpleTuner itself as part of a service. That is a question for your own counsel, not for this article.

Separately, the models you train carry their own licences, and the README's model table is explicit about this. Anima is listed as non-commercial for the model, with outputs allowed. Downloading a base model and training a LoRA on it does not change the base model's terms.

The last push to the repository was on 2026-09-07, and the most recent release, v4.9.2, is described as minor bugfixes and quality-of-life changes. The pace of releases is the main upgrade cost: three releases in the nine days before that, one of which bundled several new training techniques. Upgrading is not free, because cached embeddings and configuration keys are the kind of thing that changes across a minor version. Pin a version for a training campaign, finish it, then move.

The documentation is spread across the repository rather than centralised: DEEPSPEED.md, FSDP2.md, DISTRIBUTED.md, QUICKSTART.md, TUTORIAL.md, SLIDER_LORA.md, CAPTIONFLOW.md, plus the experimental server guides. The README is explicit that you should read it fully before starting the tutorials, and having looked at the structure, that advice is sound: the README is where the model table, the licence columns and the hardware envelope live, and none of those are repeated in the tutorial files.

Editorial conclusion

Adopt SimpleTuner if you train diffusion models regularly, want one pipeline for image, video and audio, and are comfortable editing configuration rather than clicking through a wizard. Skip it if you need a one-click consumer trainer, if your workflow is built entirely around kohya or OneTrainer config files, or if AGPL-3.0 obligations conflict with how you ship your product. Before committing, verify three things: that your target model family appears in the compatibility table, that your GPU fits the documented memory envelope for that model, and that the web UI on port 8001 comes up under your chosen install path.

Frequently asked questions

How does SimpleTuner compare with kohya?

kohya's scripts are a single-machine, flag-and-TOML style trainer, while SimpleTuner packages the same config-file approach with a web UI, a job queue and a worker model for distributed GPUs. Configs are not interchangeable between them, so moving means re-expressing your settings in SimpleTuner's configuration system.

How does SimpleTuner compare with OneTrainer?

OneTrainer is a desktop application with a graphical configuration workflow, while SimpleTuner's web UI is served on port 8001 from a process you run, typically in Docker or a virtualenv. On a headless rented GPU the browser interface is reachable without tunnelling a desktop app.

Which Python version does SimpleTuner need?

The pyproject.toml sets requires-python to 3.12 or newer, below 3.14, and the classifiers list 3.12 and 3.13. An environment pinned to 3.11 or earlier cannot install the package without being changed first.

Does SimpleTuner send my training data anywhere?

The README states that no data is sent to third parties except through the opt-in report_to flag, push_to_hub, or webhooks that must be configured manually. Nothing leaves the machine by default.

Which diffusion model families does SimpleTuner support?

The README's model table lists families including ACE-Step at 3.5B parameters, Anima, Auraflow at 6B, and Boogu-Image, each with a licence and a commercial-use column. The quick start guide carries the detailed per-model training feature compatibility.

Official sources

  1. bghira/SimpleTuner on GitHub
  2. Issues
  3. License: AGPL-3.0
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/bghira-simpletuner.svg)](https://hysenlabs.com/projects/bghira-simpletuner)