# Qwen-VL-Series-Finetune, and the flag combinations that break it

> Qwen-VL-Series-Finetune is a HuggingFace training script for the Qwen vision-language line, covering SFT, DPO, GRPO, classification and video from one codebase across Qwen2-VL, Qwen2.5-VL, Qwen3-VL and Qwen3.5. Its real contribution is the compatibility list, which states in advance which flag pairs will fail, and a requirements file pinned hard enough to make that possible.

**2U1/Qwen-VL-Series-Finetune** — An open-source implementaion for fine-tuning Qwen-VL series by Alibaba Cloud.

- Repository: https://github.com/2U1/Qwen-VL-Series-Finetune
- Stars: 1,977 · Forks: 224
- Language: Python
- License: Apache-2.0
- Published: 2026-09-29 · Updated: 2026-09-29 · Language: en
- Canonical page: https://hysenlabs.com/projects/2u1-qwen-vl-series-finetune

## One training script across four Qwen vision generations

The scope is narrower than the name suggests and that is the point. This repository trains Qwen2-VL, Qwen2.5-VL, Qwen3-VL and Qwen3.5 using HuggingFace and Liger-Kernel, with no custom modelling layer in between. The layout shows how little is here: a scripts directory, a src directory, an environment.yaml and a requirements.txt. Maintenance is visible in the update log rather than in version tags, since the repository has no GitHub releases. The log runs from 2024/09/12, when the model was switched to Liger-Kernel, through multi-image and video training in September 2024, DPO in April 2025, GRPO in May 2025, non-MoE Qwen3-VL in October 2025, the MoE variant in November 2025, video with DPO and GRPO later that month, then a cluster of Qwen3.5 and reasoning-mode work on 2026/03/07, and a jump to liger_kernel==0.8.0 on 2026/05/18. The last push was 2026-09-09. Read that log as the compatibility contract: each line is a model generation that had to be added by hand.

## Training notes are the part worth reading twice

The README puts a warning above the training scripts telling you to read the training notes first, and the warning is justified because those notes are where the project's real knowledge sits. They read like a bug list written by someone who already lost the afternoon. For Qwen3.5, use --disable_flash_attn2 True, because in the author's local testing Flash Attention 2 raised CUDA errors while sdpa stayed stable, and this applies to SFT, classification, DPO and GRPO alike. For DeepSpeed, zero2 is described as usually faster and often more stable than zero3, at the cost of more memory. For video, do not set fps and nframes at the same time. On learning rates, the vision_model usually wants a rate about five to ten times smaller than the language_model. None of this is a tutorial. It is a set of boundaries, and a script that does not tell you its boundaries will cost you more time than one that does.

## QLoRA and vision training collide, by design of the flags

The most useful entry is the rule that quantization and vision training do not go together. The note says not to combine --bits 4 or --bits 8 with --vision_lora True, with --freeze_vision_tower False, or with --unfreeze_topk_vision above zero, and to pass --bits 16 instead when the vision side has to train. A second rule says to disable liger when using QLoRA, which is the kind of interaction that silently costs you memory savings you were counting on. A third says that if you use --unfreeze_topk_llm or --unfreeze_topk_vision, you must first freeze the corresponding base module with --freeze_llm True or --freeze_vision_tower True. The repository also offers a separate option to unfreeze only a few layers in the language model and the vision tower, plus an option for a two-layer MLP in classification training. These are not convenience switches. They are the difference between a run that fits in memory and a run that does not, and the notes tell you which is which before you start.

## Liger-Kernel is patched in, not just switched on

The optimisation story has two parts. The first is ordinary: an option to use Liger-Kernel, first enabled on 2024/09/12, with support added for Qwen2.5-VL in February 2025 and for Qwen3-VL in November 2025. The second is more invasive. On 2025/08/08 the entry reads that Qwen2.5-VL's window attention and forward are monkey patched to use less memory and gain speedups. A monkey patch in a training script is a maintenance liability, because it binds to the internals of a specific model revision and to the Liger version at the time it was written. The requirements file is the place where that risk is managed: it pins liger_kernel==0.8.0, transformers==5.3.0, peft==0.15.2, deepspeed==0.17.5, accelerate==1.10.1, bitsandbytes==0.49.2 and datasets==3.5.1, all with exact version operators. A model upgrade will therefore fail loudly rather than drift silently, which is the behaviour you want from a patch this deep.

## Install from requirements, conda, or a prebuilt image

Three documented paths, and the fastest one is the container. The README gives a Docker image with the conda environment already named train configured:

```bash
docker pull john119/vlm
docker run --gpus all -it -v /host/path:/docker/path --name vlm --ipc=host john119/vlm /bin/bash
```

The --ipc=host flag is not incidental, since shared memory sizing is a frequent cause of worker crashes in multimodal training. For a local install the stated platform is Ubuntu 22.04 with Nvidia-Driver 550.120 and CUDA 12.8, then either of two routes. With requirements.txt:

```bash
pip install -r requirements.txt -f https://download.pytorch.org/whl/cu128
pip install qwen-vl-utils
pip install flash-attn --no-build-isolation
```

Or with the conda file, which creates and activates the same train environment:

```bash
conda env create -f environment.yaml
conda activate train
pip install qwen-vl-utils
pip install flash-attn --no-build-isolation
```

Both routes carry the same warning about ordering: install flash-attn after everything else, and use --no-build-isolation so it can compile against the torch already present. The extra index for cu128 is what keeps the wheels matched to your CUDA version.

## LLaVA-format JSON is the only dataset contract

Data preparation is deliberately unopinionated. The script wants a JSON file in the LLaVA specification, where each entry carries the conversation and the images, and everything else follows from the feature list: multi-image, video, and mixed-modality datasets, the last of which got zero3 support in February 2025 and a fix in the same month for loading images properly. Video decoding is handled by av and decord in the requirements. If your corpus is already in ShareGPT format, or in a parquet dataset with an image column, converting it is your first task and the README does not describe a converter, so expect to write a small script. The payoff for doing that work is that one dataset file feeds SFT, DPO, GRPO and classification, since the same conversation structure is what every one of those objectives consumes.

## Evaluation, DPO and GRPO are edits you make yourself

The README's evaluation section is a four-step procedure rather than a feature: prepare an evaluation dataset, define a compute_metrics function, modify the training script, then add evaluation arguments. That is a description of integration work, and it is worth being blunt about what it means. There is no evaluation harness here, and you will be writing the metric yourself against whatever your task considers correct. DPO and GRPO are in the same category. Both are listed as supported, video support for DPO and GRPO was added on 2025/11/28, and the GRPO section carries a prerequisites list rather than a command, which suggests reward function construction is on you. The inference side is friendlier: a Gradio WebUI is provided, and the README has a section for libcudnn errors, which is a sign that the maintainer expects people to hit that wall. The honest summary is that this is a well-informed starting point, not a pipeline you can hand to someone else.

## What it is against, and when to choose the other route

The usual alternative for fine-tuning vision-language models is a platform such as LLaMA-Factory, which exposes many model families through one web interface and one set of YAML configuration files, with the training loop hidden behind the tool. The approach difference is the whole decision. Here you own the script: you read the training loop, you patch it, and you are responsible when the model internals move. There you trade that control for breadth and a UI, and you accept that a new model family needs the tool to support it first. Choose this repository when your work is specifically on a Qwen vision model, when you intend to modify the training code, and when you need a flag such as unfreezing only the top-k vision layers that a general platform will not expose. Choose the platform when you need to compare four model families in an afternoon, or when nobody on the team wants to own a training script. There is also a licensing difference that matters more than it looks: this repository is Apache-2.0, which is permissive enough for commercial work with no source disclosure attached.

## Conclusion

Adopt Qwen-VL-Series-Finetune if you are fine-tuning a Qwen vision-language model and want a script you can read and edit rather than a configuration layer, and if you can hold to the platform it names: Ubuntu 22.04, driver 550.120, CUDA 12.8 and a fully pinned requirements file. Do not adopt it for a model outside the Qwen family, and do not expect DPO, GRPO or in-training evaluation to run unconfigured, since the README describes them as steps you perform inside the script. Verify first that your Qwen3.5 run passes --disable_flash_attn2 True, that you are not combining --bits 4 or --bits 8 with any vision training flag, and that the liger-kernel and flash-attn versions in requirements.txt still match the transformers pin, since the project tracks upstream releases by editing that file and ships no tagged releases of its own.

## FAQ

### Which Qwen models can Qwen-VL-Series-Finetune train?

The README covers Qwen2-VL, Qwen2.5-VL, Qwen3-VL including the MoE variant, and Qwen3.5. Newer model generations are added by editing the scripts, and the update log records each one.

### What format does the dataset need to be in?

A JSON file following the LLaVA specification, where each entry holds the conversation and the images. Multi-image, video and mixed-modality datasets are supported, with av and decord handling video decoding.

### Why does Qwen3.5 training need Flash Attention disabled?

The training notes say to pass --disable_flash_attn2 True for the Qwen3.5 series, because in the author's local testing Flash Attention 2 raised CUDA errors while sdpa was stable. The note applies to SFT, classification, DPO and GRPO.

### Can I combine QLoRA with vision training in Qwen-VL-Series-Finetune?

Not as the flags are written. The notes say not to combine --bits 4 or --bits 8 with --vision_lora True, --freeze_vision_tower False or --unfreeze_topk_vision above zero, and to use --bits 16 when training vision modules. Liger should also be disabled under QLoRA.

### What is the fastest way to run a training job from this repository?

Use the prebuilt image, which has the conda environment named train already set up: docker pull john119/vlm, then run it with --gpus all and --ipc=host. A local install needs Ubuntu 22.04, driver 550.120 and CUDA 12.8.

### What licence is Qwen-VL-Series-Finetune under?

Apache-2.0, with the LICENSE file at the repository root. There are no GitHub releases, so the pinned requirements.txt and environment.yaml are the only way to reproduce an environment.

## Sources

- [2U1/Qwen-VL-Series-Finetune on GitHub](https://github.com/2U1/Qwen-VL-Series-Finetune)
- [Issues](https://github.com/2U1/Qwen-VL-Series-Finetune/issues)
- [License: Apache-2.0](https://github.com/2U1/Qwen-VL-Series-Finetune/blob/master/LICENSE)
- [README](https://github.com/2U1/Qwen-VL-Series-Finetune/blob/master/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/2u1-qwen-vl-series-finetune
