FLUX.2: running the open-weight image models locally
Official inference repo for FLUX.2 models
At a glance
- What is it?
- Black Forest Labs publishes minimal Python inference code for FLUX.2, covering the 4B klein model that fits in about 8GB of VRAM and the 32B dev model that does not. The license split matters more than the architecture.
- Who is it for?
- Adopt FLUX.2 if you need text-to-image and multi-reference editing in one model and can work inside its license terms: the klein 4B and 4B Base weights are Apache 2.0, while every 9B variant and FLUX.2 [dev] fall under the FLUX Non-Commercial License, which rules out most commercial deployment. Skip it if you have no NVIDIA GPU, since the README states the code was tested on GB200 with CUDA 12.9, or if you want a hosted endpoint rather than local weights.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Activity is slowing. The repository last received commits 6 months ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 4, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What FLUX.2 is for, and who should pull the repo
FLUX.2 is the official inference repository for Black Forest Labs' open-weight image models. The README describes it as "minimal inference code to run image generation & editing" with those models. That word minimal is doing real work: this is not a training framework, not a serving stack, and not a node graph. It is a Python package that loads weights and produces images.
The models cover text-to-image, single-reference editing, and multi-reference editing. According to the model table, every listed checkpoint supports all three, which is unusual. Most open image models handle generation well and treat editing as a separate fine-tune. Here the same weights do both, so a pipeline that generates a product shot and then edits it can stay on one checkpoint.
The intended audience is fairly narrow. You need a GPU, you need to be comfortable running Python from a virtual environment, and you need to read a license before you ship anything. If you want a hosted API, the README points at docs.bfl.ai instead. If you want a drag-and-drop interface, this repository does not provide one; ComfyUI integration is not covered by the README.
The klein and dev model split, and why the license decides
The repository ships two model families. FLUX.2 [klein] is the fast one, described in the README as generating and editing images in under a second on modern hardware. FLUX.2 [dev] is a 32B parameter flow matching transformer for maximum quality when latency does not matter.
Within klein there is a second split that matters more than size: distilled versus Base. The distilled models are 4-step and the README recommends them for production apps and real-time generation. The Base models are 50-step and recommended for fine-tuning, LoRA training, and output diversity. A 50-step model is not a slower version of a 4-step model; it is a different starting point for training. If you plan to train a LoRA, starting from a distilled checkpoint is the wrong move.
The license table is the part to read twice. FLUX.2 [klein] 4B and FLUX.2 [klein] 4B Base are Apache 2.0. Every 9B variant, including the KV and Base versions, uses the FLUX Non-Commercial License, and so does FLUX.2 [dev]. The repository's own LICENSE.md is Apache-2.0, but the weights carry their own terms in model_licenses/. A permissive repository license does not relicense the checkpoints.
Installing FLUX.2 and running a first generation
The README states the inference code was tested on GB200 using CUDA 12.9 and Python 3.12. The pyproject.toml pins requires-python to >=3.10,<3.13, so 3.12 is the safe choice. Create the environment and install the package in editable mode with the CUDA 12.9 wheel index:
python3.12 -m venv .venv
source .venv/bin/activate
pip install -e . --extra-index-url https://download.pytorch.org/whl/cu129 --no-cache-dirThe install pulls pinned dependencies rather than ranges: torch 2.8.0, torchvision 0.23.0, transformers 4.56.1, safetensors 0.4.5, einops 0.8.1, accelerate 1.12.0, fire 0.7.1, and openai 2.8.1. The openai package is there for the OpenRouter prompt upsampling path, not because the models are API-backed. Because the pins are exact, expect a large download and expect resolution to fail loudly if your platform has no matching torch wheel.
The README's installation section ends at the venv step and its CLI section is truncated in the README text available here, so the exact invocation is not something this article can state. What the repository does provide is scripts/ and src/, so the entry point lives there. Read the CLI section of README.md in the checkout before assuming a flag name; guessing at arguments for an image model wastes a multi-gigabyte model load when it fails.
For FLUX.2 [dev] specifically, the README warns that the script needs an H100-equivalent GPU. It also notes a partnership with Hugging Face on quantized versions, with instructions for running on an RTX 4090 using a remote text encoder, and points to docs/flux2_dev_hf.md for other quantization sizes.
KV caching is what makes klein 9B the editing pick
The klein family has an odd property that the README states plainly: for image editing, 9B KV is faster than 4B for multi-reference work, via KV caching. That is counterintuitive. A larger model beating a smaller one on latency usually means the comparison is unfair, but here the mechanism explains it. Multi-reference editing feeds several images into the same context, and KV caching avoids recomputing that context on subsequent steps. The 4B model has no equivalent, so it pays the full cost each time.
The practical consequence is that model size is the wrong first question for editing workloads. If your application takes multiple reference images, the 9B KV checkpoint is the one the README recommends for quality-to-latency ratio. If you are doing plain text-to-image, klein 4B is the recommendation for consumer hardware, and the README says it fits in roughly 8GB of VRAM on an RTX 3090 or 4070 and up.
That 8GB figure is a claim from the project's own documentation, not an independently measured number, and it will depend on resolution and batch size. Treat it as a floor for the smallest configuration, not a guarantee for your workload.
Prompt upsampling, and the autoencoder change
The README states that FLUX.2 [dev] benefits significantly from prompt upsampling, and the inference script offers two routes: local upsampling with Mistral-Small-3.2-24B-Instruct-2506, the same model used for text encoding, or any model on OpenRouter via an API call. There is a guide at docs/flux2_with_prompt_upsampling.md with guidance on when to use it.
This is a real design decision with a cost. Local upsampling means loading a 24B language model alongside the image model, which is a substantial addition to VRAM and startup time. The OpenRouter route moves that cost to an API call and adds a network dependency and a third-party key to a pipeline that is otherwise fully local. Neither option is free, and the README does not present one as the default.
The autoencoder is separate and worth noting. The README says the FLUX.2 autoencoder improved considerably over the FLUX.1 autoencoder, is released under Apache 2.0, and lives in the FLUX.2-dev Hugging Face repository as ae.safetensors. A permissively licensed VAE does not change the license on the transformer weights around it, but it does mean the encode and decode stage is not the part of your stack you need to worry about.
Where FLUX.2 is the wrong tool
The clearest limitation is hardware. The README states the code was tested on GB200 with CUDA 12.9. FLUX.2 [dev] needs an H100-equivalent GPU unless you go through the quantized path, and that path involves a remote text encoder, which means part of your inference is no longer local. If your environment is CPU-only, or AMD, or macOS, nothing in the README describes a supported route.
The second limitation is licensing, and it is easy to get wrong because it is per-checkpoint. If you need commercial rights, you are restricted to the 4B models. The 9B models are the ones with the better editing latency story via KV caching, and they are non-commercial. So the technically attractive option and the legally usable option do not overlap. That is a genuine constraint, not a footnote.
The third is that this repository is inference only. There is no training script, no dataset tooling, and no evaluation harness in the top-level layout. Base checkpoints are offered for fine-tuning and LoRA training, but the training code is not here. You bring your own.
Finally, there are no releases listed for this repository, and the last push was on 2026-03-12. That is roughly six months before the date of this writing. The code is not archived, but anyone planning to build on it should check the commit history rather than assume a steady cadence.
FLUX.2 compared with diffusers
The obvious alternative for running these weights is Hugging Face diffusers, and the README itself points there for the quantized FLUX.2 [dev] path, linking docs/flux2_dev_hf.md and describing it as the diffusers quantization guide. The approaches differ in a way that matters beyond preference.
diffusers is a general library covering many model families, with quantization support, schedulers, and pipelines built around a common interface. That generality is the point: you can swap models, share code across projects, and use quantization configurations that this repository does not implement. FLUX.2's own repository is the opposite trade. It is pinned to exact dependency versions, it is minimal by design, and it is the reference implementation from the people who trained the models. When a new checkpoint lands, this is where it appears first.
If you want quantization on consumer hardware, or you want one codebase for several model families, diffusers is the better fit and the README acknowledges that by linking to it. If you want the reference path with the fewest layers between you and the weights, and you are on the hardware the project tested against, this repository is the shorter route. Note that the comparison is not either-or: the FLUX.2 [dev] quantization instructions live in the diffusers documentation.
Editorial conclusion
Adopt FLUX.2 if you need text-to-image and multi-reference editing in one model and can work inside its license terms: the klein 4B and 4B Base weights are Apache 2.0, while every 9B variant and FLUX.2 [dev] fall under the FLUX Non-Commercial License, which rules out most commercial deployment. Skip it if you have no NVIDIA GPU, since the README states the code was tested on GB200 with CUDA 12.9, or if you want a hosted endpoint rather than local weights. Before committing, read model_licenses/LICENSE-FLUX-NON-COMMERICAL and confirm which checkpoint your pipeline actually downloads, because the license follows the weights and not the repository.
Frequently asked questions
What is FLUX.2 used for?
It is the official inference repository for Black Forest Labs' open-weight FLUX.2 models, used for text-to-image generation, single-reference image editing, and multi-reference image editing. According to the model table, every listed checkpoint supports all three modes.
Is FLUX.2 free to use?
It depends on the checkpoint. FLUX.2 [klein] 4B and 4B Base are Apache 2.0, while the 9B variants and FLUX.2 [dev] use the FLUX Non-Commercial License stored in model_licenses/. The repository's own LICENSE.md is Apache-2.0, but that does not relicense the weights.
How do I install FLUX.2 locally?
The README gives a three-step install: create a Python 3.12 virtual environment, activate it, then run pip install -e . with the PyTorch cu129 extra index URL. The pyproject.toml restricts Python to >=3.10,<3.13 and pins torch to 2.8.0.
How do I use FLUX.2 [klein] 9B?
The README recommends klein 9B KV rather than plain 9B for image editing, because KV caching makes it faster than 4B for multi-reference work at equal quality. For text-to-image, it lists 9B as the high quality option. Both 9B checkpoints use the FLUX Non-Commercial License.
How do I use FLUX.2 [dev]?
FLUX.2 [dev] is a 32B parameter flow matching transformer, and the README warns its script needs an H100-equivalent GPU. For consumer hardware it points to quantized versions made with Hugging Face, with instructions for an RTX 4090 using a remote text encoder in docs/flux2_dev_hf.md.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/black-forest-labs-flux2)