llamafactory defines your environment with dependency pins, and its support table drifts from its own release notes
Unified Efficient Fine-Tuning of 100+ LLMs & VLMs (ACL 2024). **Scalable resources**: 16-bit full-tuning, freeze-tuning, LoRA and 2/3/4/5/6/8-bit QLoRA via AQLM/AWQ/GPTQ/LLM.int8/HQQ/EETQ.
At a glance
- What is it?
- LlamaFactory wraps fine-tuning of 100+ LLMs and VLMs in one CLI and one Gradio interface, and supports everything from 16-bit full tuning down to 2-bit QLoRA. The cost of that breadth is that it dictates your dependency set, and the two places it documents its own coverage, the README support table and the release titles, do not agree.
- Who is it for?
- Adopt LlamaFactory when you want one interface across full tuning, LoRA and QLoRA and you are willing to give it a dedicated environment, because it pins your dependency set rather than fitting into yours.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 1 day ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
DEEP OPEN-SOURCE ANALYSIS
peft is capped at 0.18.1 and two transformers versions are excluded outright
The dependency list in pyproject.toml is where this project makes most of its real decisions, and two entries are tighter than the rest:
"transformers>=4.55.0,<=5.8.0,!=4.57.0,!=5.6.0",
"peft>=0.18.0,<=0.18.1",The peft window admits two releases. The transformers range is wide at both ends and then carves out two exact versions, 4.57.0 and 5.6.0, by number rather than by range.
The consequence is that your environment is defined by this project rather than by your own requirements. If anything else in the same environment needs a newer peft, resolution fails, and the practical answer is a dedicated virtual environment rather than a relaxed pin. The exclusions are the sharper edge: an unrelated package that happens to require transformers 4.57.0 breaks the install for a reason that has nothing to do with your model or your data. Read those two version numbers before you start, because they are the most likely cause of a resolution failure you will hit.
The test target disables the experiment logger and runs two suites at once
The Makefile has five targets, and the test one is the most revealing:
test:
WANDB_DISABLED=true $(RUN) pytest -vv --import-mode=importlib tests/ tests_v1/Two things are happening. The experiment logger is switched off explicitly through an environment variable, which tells you the Weights and Biases integration is treated as optional rather than required. And both tests/ and tests_v1/ are collected in the same run.
That second directory is not an accident. The repository also carries an examples/v1/ directory, so a v1 line is maintained alongside the current code. The consequence for a user is a good one: a pipeline written against v1 keeps a maintained path rather than rotting. The cost is a test surface roughly twice what a single-suite project of this size would have, and it means the two suites can disagree about behaviour, so an upgrade note that does not say which suite changed leaves you guessing which half of the repository a fix landed in.
The Makefile silently swaps uv for whatever is on your PATH
The build tooling is selected at parse time by probing for uv:
ruff_version := 0.15.5
RUFF := $(shell command -v uv >/dev/null 2>&1 && echo "uvx ruff@$(ruff_version)" || echo "ruff")The same pattern picks the build command and the run prefix, falling back to python -m build and to a bare ruff when uv is absent. The quality target then runs both a check and a format check:
quality:
$(RUFF) check $(check_dirs)
$(RUFF) format --check $(check_dirs)The consequence is that the lint result depends on what is installed on the machine. With uv present you get ruff 0.15.5 exactly, pinned by the variable above. Without it you get whatever ruff happens to be on your PATH, at whatever version, and a formatting difference between your laptop and continuous integration can have nothing to do with your change. Install uv before you rely on the quality target, or pin ruff yourself, otherwise you will spend an afternoon on a diff that is a version artefact.
The Day-N support table is behind the release that added those models
The README publishes a small table of how quickly new models get fine-tuning support. Day 0 covers Qwen3, Qwen2.5-VL, Gemma 3, GLM-4.1V, InternLM 3 and MiniCPM-o-2.6. Day 1 covers Llama 3, GLM-4, Mistral Small, PaliGemma2 and Llama 4.
Now compare that with the release titles. v0.9.5, dated 2026-05-30, is titled for Qwen3.5/3.6, Gemma 4 and Transformers v5. The table still says Qwen3 and Gemma 3.
So the project's own support-lag metric is maintained by hand in a README and has already drifted from what the releases added. The consequence is that you should not read the table as the current support matrix. Release titles are the better signal for what a given version handles, and for a specific model version you should verify the example configuration yourself rather than trusting either. The table is still worth reading, because the gap between Day 0 and Day 1 tells you that support is reactive by design rather than pre-planned.
ROCm and Ascend NPU instructions live on separate sites, and the main docs are marked WIP
The documentation links are split three ways. The primary one is llamafactory.readthedocs.io and it is labelled documentation, work in progress. The second is an AMD GPU page hosted on rocm.docs.amd.com. The third is an ASCEND NPU page hosted under the multibackend section of the same readthedocs domain.
The example tree confirms these are real paths rather than aspirations, with directories for accelerate, ascend, deepspeed, ktransformers, megatron, megatron_bridge, merge_lora, train_full, train_lora, train_qlora and inference, alongside a v1 directory.
The consequence is that if you are on ROCm or on an Ascend NPU, the main documentation is not the one describing your path, and even on the default path the primary docs are self-described as incomplete. The examples become your reference instead, which is workable because each backend has its own directory, but it means you are reading configuration files rather than prose when you hit an unfamiliar setting. Choose your backend before you read the docs, not after.
The README warns that most websites carrying the name are unauthorized
One note on the project page deserves more attention than it usually gets. It states that except for the links listed there, all other websites are unauthorized third-party websites, and asks readers to use them carefully.
That is the maintainers drawing a line around their own distribution channels, and it is worth taking at face value in a project whose whole purpose is executing training code that downloads models and datasets you name, on a machine that often holds credentials for exactly those hubs.
The practical consequence is about where you install from. The links the project does bless are the PyPI package, the project's own Docker Hub tags, the readthedocs documentation, the official blog, a Discord, a X account, a WeChat and NPU user group, and a community repository. The project also promotes a Colab notebook, a PAI-DSW free trial gallery and an AMD GPU Cloud notebook with free credits, all labelled as such. Source the install from the listed channels rather than from a page that merely mentions the name.
The example tree is the documentation, and training and serving are one tool
With the prose docs incomplete, the thirteen example directories do the explaining. Three of them map directly onto the resource levels the project advertises: train_full, train_lora and train_qlora, corresponding to 16-bit full-tuning, LoRA, and the 2, 3, 4, 5, 6 and 8-bit QLoRA levels reached through AQLM, AWQ, GPTQ, LLM.int8, HQQ and EETQ. A fourth, merge_lora, covers the step that follows adapter training, which is the one people forget until they have an adapter they cannot serve.
The same tool also serves. Faster inference is listed as an OpenAI-style API, a Gradio interface and a CLI, each with a vLLM worker or an SGLang worker, and the dependency list carries uvicorn, fastapi and sse-starlette, so streaming responses are part of the surface. The console entry point is llamafactory-cli.
The consequence is that the backend choice for inference is a decision you make at deploy time rather than at install time, and the training and serving halves share one dependency set with all the pinning that implies. If you only need to serve an existing model, pulling this dependency set to do it is a heavier choice than it looks, and the pages people ask about most, ollama and swift, are not comparisons this project makes anywhere in what it publishes.
Editorial conclusion
Adopt LlamaFactory when you want one interface across full tuning, LoRA and QLoRA and you are willing to give it a dedicated environment, because it pins your dependency set rather than fitting into yours. Check three things first: whether the peft window of 0.18.0 to 0.18.1 and the two excluded transformers versions collide with anything else you run, since resolution fails on your schedule and not the project's; whether the model you want appears in a release title rather than in the README support table, which is maintained by hand and already trails v0.9.5; and whether you are on the primary backend, because the ROCm and Ascend NPU instructions live on separate sites and the main documentation is labelled work in progress.
Frequently asked questions
What is LlamaFactory?
It is a Python tool for fine-tuning 100+ large language models and vision-language models, presented as a zero-code CLI and a Gradio interface called LlamaBoard. It covers continuous pre-training, supervised fine-tuning, reward modeling, PPO, DPO, KTO and ORPO, across 16-bit full tuning, freeze-tuning, LoRA and 2 to 8-bit QLoRA. The project is Apache-2.0 licensed and requires Python 3.11.0 or newer.
How do I install LlamaFactory?
The distribution channel the project points at is the PyPI package named llamafactory, and the console script it declares is llamafactory-cli. The package requires Python 3.11.0 or newer and the classifiers cover 3.11, 3.12 and 3.13. There is also a container route, with the project's own tags on Docker Hub, and a Build Docker section in the table of contents.
What does the llamafactory-cli command do?
It is the console entry point declared in pyproject.toml, and the README presents the tool as usable from a CLI with zero code as well as from the Gradio interface. The example directories show the shapes it drives, including train_full, train_lora, train_qlora, merge_lora and inference, and the related list also pairs the CLI with a webui form of the command.
How does LlamaFactory compare with TRL?
TRL is not an alternative to it here, it is a dependency. pyproject.toml requires trl>=0.18.0,<=0.24.0, and peft>=0.18.0,<=0.18.1 alongside it, so LlamaFactory is a layer built on those libraries rather than a replacement for them. The project's own material does not make a comparison against TRL, Axolotl, Swift or DeepSpeed anywhere in what it publishes.
Official sources
Where this project is recommended
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/hiyouga-llamafactory)
Community notes