ostris AI Toolkit: LoRA Training for FLUX, Video, and Audio Diffusion Models
The ultimate training toolkit for finetuning diffusion models
At a glance
- What is it?
- ostris AI Toolkit is a Python fine-tuning suite for diffusion models that targets artists and researchers who want LoRA adapters for FLUX, SDXL, Wan video, and audio models on consumer-grade GPUs. It ships a GUI and CLI, a Docker image, and platform-specific launch scripts.
- Who is it for?
- Artists and ML researchers who want a single tool for LoRA training across FLUX, SDXL, Wan video, and audio models will find ostris AI Toolkit the most comprehensive open-source option for consumer GPU hardware. Those using FLUX.1-dev should confirm whether its gated Hugging Face license covers their use case before starting any training run.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 3 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.
Editorial analysis
Who ostris AI Toolkit Is For and What It Solves
Fine-tuning large diffusion models has historically required assembling separate tools for each architecture. A researcher training SDXL LoRAs would reach for Kohya_ss; someone working with FLUX.1 needed different scripts; video model training had almost no consumer-accessible tooling at all. ostris AI Toolkit addresses that fragmentation by providing a single codebase that handles image, edit, video, audio, and even LLM fine-tuning under one interface.
The primary audience is artists who want to train custom styles or characters, researchers exploring diffusion model adaptation, and engineers who need reproducible fine-tuning pipelines. The README describes the goal as supporting "all the latest models on consumer grade hardware," which makes it distinct from cloud-only or multi-node training frameworks. The MIT license means there are no commercial use restrictions from the toolkit itself, though the licenses of the base models vary.
The Model Roster: Image, Edit, Video, and Audio
The repository explicitly lists supported models by category. The image category includes FLUX.1, FLUX.2, and the FLUX.2-klein variants at 4B and 9B parameters, along with SDXL, SD 1.5, HiDream I1 and O1, Qwen-Image and its 2.1 variant, OmniGen2, Lumina2, Krea 2, Anima, Mage-Flow, and several others. The edit and instruction category covers FLUX.1-Kontext-dev and multiple Qwen-Image-Edit variants; for Qwen-Image-2.1 specifically, the same model handles both generation and editing: it switches to edit mode when the dataset includes control images.
The video category covers Wan 2.1 (1.3B text-to-video, 14B text-to-video, and 14B image-to-video in 480P and 720P), Wan 2.2 in text-to-video, image-to-video, and TI2V-5B variants, plus three Lightricks LTX releases (LTX-2, LTX-2.3, LTX-2.5) and MiniMax-H3 including a reference-image-to-video mode.
The audio category includes Ace Step 1.5 and Ace Step 1.5 XL. YuE2-3B is listed with a specific caveat: the official audio-to-token encoder from the model's authors is unreleased. Training with YuE2 currently depends on a community tokenizer. This is a concrete dependency risk that the README documents but does not resolve.
The LLM category contains Qwen2.5-Omni-7B. Zeta Chroma is listed separately as experimental.
Installing ostris AI Toolkit: Manager, Scripts, and Docker
The README describes three routes. The AI Toolkit Manager, built into the repository, is the recommended starting point. It detects the local GPU, sets up the correct PyTorch build, creates an isolated Python environment, downloads local copies of Node.js and FFmpeg inside the ai-toolkit folder, and starts the web UI at http://localhost:8675. On each launch it checks for updates and applies them; if local files have been modified, the update is skipped with a warning to avoid overwriting changes. The README explicitly marks the manager as experimental.
Platform-specific scripts at the repository root (run_linux.sh, run_mac.zsh, run_windows.bat) provide manual launch paths for users who prefer not to use the manager.
The Docker path uses an official image. The docker-compose.yml in the repository maps port 8675, mounts the Hugging Face hub cache, the SQLite database, the datasets folder, the output folder, and the config folder as volumes, and requires NVIDIA GPU passthrough:
version: "3.8"
services:
ai-toolkit:
image: ostris/aitoolkit:latest
restart: unless-stopped
ports:
- "8675:8675"
volumes:
- ~/.cache/huggingface/hub:/root/.cache/huggingface/hub
- ./aitk_db.db:/app/ai-toolkit/aitk_db.db
- ./datasets:/app/ai-toolkit/datasets
- ./output:/app/ai-toolkit/output
- ./config:/app/ai-toolkit/config
environment:
- AI_TOOLKIT_AUTH=${AI_TOOLKIT_AUTH:-password}
- NODE_ENV=production
- TZ=UTC
deploy:
resources:
reservations:
devices:
- driver: nvidia
count: all
capabilities: [gpu]The AI_TOOLKIT_AUTH environment variable sets the web UI password and defaults to the string "password" if unset, which is suitable for local-only use but should be changed before any network-exposed deployment. Training outputs land in the output/ folder; configuration YAML files go in config/.
The Training Workflow and config/ Structure
AI Toolkit uses YAML configuration files to describe each training job. The config/ folder in the repository holds examples. A training run reads the config, loads the base model from Hugging Face (caching it in the standard Hugging Face hub directory), and writes LoRA weights and checkpoints to output/. The GUI at http://localhost:8675 provides a visual editor for these configs; the CLI path uses run.py directly. The extensions/ and extensions_built_in/ directories at the root indicate the tool has a plugin architecture, though the README portion available does not describe the extension API.
The repository also includes a notebooks/ directory, suggesting Jupyter-based workflows are possible, and a run_modal.py script for running training on Modal cloud infrastructure. The spark_requirements.txt file implies optional distributed or Spark-based workflows, though the README does not document the Spark integration path.
Where AI Toolkit Falls Short
The AI Toolkit Manager is marked experimental in the README. That caveat carries real weight: automated environment setup on NVIDIA, AMD (ROCm), and Apple Silicon can fail in ways that are harder to debug than a manual install.
YuE2's dependency on an unofficial community tokenizer is a fragility. The README acknowledges that the model author's official tokenizer is unreleased, so audio training with YuE2 depends on a third-party artifact that could change or disappear. Anyone building a production pipeline around YuE2 should track that dependency explicitly.
The docker-compose.yml hard-codes NVIDIA GPU passthrough. Users on AMD hardware cannot use Docker as provided and must build a custom image or use the manual install path. The related searches include "AI Toolkit-amd" and "ai-toolkit rocm," which suggests this limitation is a common friction point.
The README does not document specific VRAM requirements per model, and the manual installation steps are described in a portion the README marks as continuing beyond its summary section.
ostris AI Toolkit vs Kohya_ss
Kohya_ss is the most widely used alternative for SDXL and SD 1.5 LoRA training. It is also a Python tool with a web UI, similarly targeting consumer hardware. The difference is scope: Kohya_ss focuses on stable diffusion architectures and has deep support for SD 1.5 and SDXL training scenarios including DreamBooth, fine-tune, and LoRA. ostris AI Toolkit explicitly targets a wider architecture list including FLUX, video models, and audio models, which Kohya_ss does not support in the same integrated way.
The trade-off is that Kohya_ss has a longer track record for SD 1.5 and SDXL specifically, with a larger body of community guides and known-good configurations. If the target is exclusively SD 1.5 or SDXL LoRA training, Kohya_ss is the more tested option. For FLUX, Wan video, or Ace Step audio fine-tuning, ostris AI Toolkit is the more direct path.
License, Maintenance, and Upgrade Cost
The repository is MIT licensed, which allows commercial use, modification, and redistribution without restriction from the toolkit itself. Base models carry their own licenses: FLUX.1-dev, for example, is distributed under a non-commercial gated license on Hugging Face, which means MIT on the toolkit does not extend to the model weights.
The last push to the repository was on 2026-09-27, one day before this analysis. The supported model list continues to expand. The manager's automatic update mechanism means most users get changes on each launch without manual git operations. However, when local config or code files have been modified, the manager skips updates with a warning, which means users who customize the config/ or toolkit/ directories need to merge upstream changes manually.
Editorial conclusion
Artists and ML researchers who want a single tool for LoRA training across FLUX, SDXL, Wan video, and audio models will find ostris AI Toolkit the most comprehensive open-source option for consumer GPU hardware. Those using FLUX.1-dev should confirm whether its gated Hugging Face license covers their use case before starting any training run. Anyone who needs a stable, production-hardened path should note that the AI Toolkit Manager remains explicitly experimental.
Frequently asked questions
What is ostris AI Toolkit?
ostris AI Toolkit is an open-source Python training suite for fine-tuning diffusion models including FLUX, SDXL, Wan video, and audio models. It provides a web GUI at port 8675, a CLI, and a Docker image, and is designed to run on consumer-grade GPU hardware.
How do you install ostris AI Toolkit?
The README recommends using the AI Toolkit Manager, which is bundled in the repository. It detects the local GPU, sets up the Python environment, and starts the web UI at http://localhost:8675. A Docker image (ostris/aitoolkit:latest) and platform-specific scripts (run_linux.sh, run_mac.zsh, run_windows.bat) are also available.
Is AI Toolkit free?
Yes. The README states it is "Free and open source" and the repository is MIT licensed. The base models it trains, such as FLUX.1-dev, carry their own separate licenses and some have non-commercial restrictions.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/ostris-ai-toolkit)