All comparisons
Comparison

LlamaFactory vs unsloth: a fine-tuning framework against a local model studio

LlamaFactory is a Python training framework with a zero-code CLI and web UI for 100+ LLMs and VLMs, while unsloth is a desktop app and web UI for training and running models locally. They are complementary: LlamaFactory lists unsloth among its practical training tricks, so the real question is which layer you need.

Published September 20, 2026

At a glance

Projecthiyouga/LlamaFactoryunslothai/unsloth
LicenceApache-2.0Permissive: commercial use allowedApache-2.0Permissive: commercial use allowed
MaintenanceCommits in the last six monthsLast push September 28, 2026Commits in the last dayLast push September 29, 2026
LanguagePythonPython
GitHub stars75,18977,025
Read moreOur analysisGitHubOur analysisGitHub

Which one to choose

LlamaFactory

Choose LlamaFactory if you need one framework that covers pre-training, SFT, reward modeling, PPO, DPO, KTO and ORPO across more than 100 language and vision models, and you want to drive it from a CLI, a Gradio web UI or your own script.

unsloth

Choose unsloth if you want a desktop or web application to train and serve modern models on your own machine, with multi-GPU and multiple GPU vendor support, an OpenAI-compatible API, and built-in LAN or Cloudflare remote access.

What each project is actually built to do

LlamaFactory is a training framework. Its README describes integrated methods that run from continuous pre-training and multimodal supervised fine-tuning through reward modeling, PPO, DPO, KTO and ORPO, with 16-bit full-tuning, freeze-tuning, LoRA and 2 to 8-bit QLoRA via AQLM, AWQ, GPTQ, LLM.int8, HQQ and EETQ. It also lists advanced algorithms such as GaLore, BAdam, APOLLO, Adam-mini, Muon, OFT, DoRA, LongLoRA, LLaMA Pro, Mixture-of-Depths, LoRA+, LoftQ and PiSSA. The interface is a zero-code CLI plus the LLaMA Board web UI built on Gradio. unsloth is positioned differently: its README calls it the first desktop app to run and train models, and the description frames it as a local UI for training and running Gemma 4, Qwen3.6, DeepSeek, Kimi, GLM and other models. It ships native installers for Windows, macOS, Ubuntu deb and AppImage, plus a manual install script. The overlap is LoRA and QLoRA fine-tuning. The difference is that LlamaFactory is a library-shaped framework you configure, while unsloth is an application-shaped product you open.

The architecture difference: recipe configs against a bundled runtime

LlamaFactory separates the training recipe from the training loop. Model, dataset, quantization method and algorithm are declared, and the framework wires them to a training stack; the README points to Hugging Face Trainer internals as the layer underneath. That is why the same tool can cover full tuning, freeze tuning, QLoRA at several bit widths and a long list of optimizers without a new code path for each. unsloth bundles more of the stack into one executable. Its README lists training, an OpenAI-compatible API, a web UI called Unsloth Studio, dataset building from PDFs, CSVs and DOCX through Data Recipes, export to GGUF, NVFP4 and FP8, and agent connections for Claude Code, Codex and MCP. It also claims training 2x faster with 70% less VRAM and no accuracy loss, a claim made in its own README and not independently reproduced here. The practical consequence: LlamaFactory gives you composable pieces you can embed in a pipeline; unsloth gives you an opinionated environment where training, serving and export are already joined. If your workflow lives in a scheduler that calls a training command, the framework shape fits. If your workflow lives on a workstation where a person clicks through a UI, the application shape fits.

Getting each one running

LlamaFactory is installed as a Python package (the README links a PyPI project and a Docker image under hiyouga/llamafactory), then driven through its usage section, its CLI or LLaMA Board. The README also points to hosted starting points: a free Colab notebook, a PAI-DSW free trial, and AMD GPU Cloud free credits, plus documentation for AMD ROCm and Ascend NPU backends. That matters because the framework assumes you bring your own environment, and the supported model and quantization table is the thing to check before you commit. unsloth lowers the entry cost on the client side: download an installer for Windows, macOS, Ubuntu deb or AppImage, or run the curl install script on macOS, Linux and WSL, or the PowerShell script on Windows. It states support for Windows, Linux, WSL and macOS, and for Multi GPU, NVIDIA, AMD, Intel GPUs, CPUs and the Vulkan backend. The trade is control. A desktop installer is easier to start and harder to reproduce inside a container image or a CI job. LlamaFactory is harder to start and easier to pin, because the version you install is the version you configure.

Operations, serving and scaling

For serving, LlamaFactory documents deployment with an OpenAI-style API and vLLM, and it lists experiment monitors including LlamaBoard, TensorBoard, Wandb, MLflow and SwanLab. That is the operations story of a training framework: metrics go to a tracker, the trained adapter or model is served by a separate inference server you operate. unsloth folds serving into the same app. Its README describes an OpenAI-compatible API, LAN access for devices on the same network, and remote access through Cloudflare HTTPS, plus agent integrations that point local models at Claude Code, Codex and MCP. It also documents an auto-compaction feature that rolls the context window. The caution from our earlier reading stands: the default server tools demand care, and anyone exposing LAN or Cloudflare access should review those options before turning them on. Scaling also diverges. unsloth advertises multi-GPU setups in the README. LlamaFactory's scaling story is the breadth of methods and quantization formats, and its release cadence of roughly every six months means you should plan upgrades rather than expect continuous drift.

Where each one falls short

LlamaFactory's limitation is coverage lag. Its own documentation is marked work in progress in the README, and if a training method or a model architecture is not yet integrated, you are back to raw PyTorch and Hugging Face Trainer. The supported models table is the gate: a model or quantization format missing from it is a blocker, not a configuration detail. The v0.9.5 release also means recent API changes are worth checking before you build on a pinned older example. unsloth's limitation is the opposite shape. It is a beta-track product: the recent releases carry beta version numbers, and our earlier reading flagged that rapid beta releases and the auto-compaction feature warrant testing on your own hardware. A headless, script-only workflow is a poor fit, and the security implications of exposing local models are real rather than theoretical. Neither README documents rollback or a downgrade path, so version pinning is on you in both cases.

Licence and maintenance implications

Both projects are Apache-2.0, which permits commercial use and modification with the usual notice and patent terms. The maintenance picture differs in rhythm. LlamaFactory's last push was 2026-09-14, days before this writing, and its release history shows v0.9.5 on 2026-05-30, v0.9.4 on 2025-12-31 and v0.9.3 on 2025-06-16: active development with a roughly six-month release cadence, so an upgrade is a planned event with a changelog to read. unsloth's last push was 2026-09-16, the day after LlamaFactory's, with beta releases through late August 2026. That is a faster, smaller-grained release stream, which cuts both ways: fixes arrive quickly, and so does churn. For a regulated deployment, the Apache-2.0 grant is the same on both sides; the difference is how often your pinned version stops being the current one. Neither repository is archived, and neither README documents rollback, so treat the release history as your compatibility signal.

Choosing for a concrete scenario

A team fine-tuning several architectures for research, comparing DPO against ORPO against KTO on the same dataset, wants LlamaFactory: the method matrix and the CLI fit a pipeline, and the experiment monitors plug into existing tracking. A single engineer on a workstation with one or two GPUs who wants to train an adapter, chat with the result and expose it to Claude Code or Codex over the local network wants unsloth: the desktop app removes the environment work, and the OpenAI-compatible API plus LAN access is the point. A shop that must run training inside a container and serve on a separate inference cluster should start from LlamaFactory and its vLLM deployment path. A shop that must keep model weights and prompts on the machine, with no external inference provider, should look at unsloth's local serving before anything else. And because LlamaFactory lists unsloth among its practical training tricks, the two can be combined: use unsloth's kernels for a faster LoRA or QLoRA run inside a LlamaFactory recipe, and keep LlamaFactory as the orchestration layer. Verify the supported model and quantization entry for your exact target in LlamaFactory, and verify the exact model support and remote access settings in unsloth, before either choice becomes a commitment.

Bottom line

Pick LlamaFactory when breadth of training methods and a scriptable, container-friendly framework matter more than a one-click app; pick unsloth when a local desktop or web UI, multi-GPU and multi-vendor support, and built-in OpenAI-compatible serving are the priority, accepting beta-release churn and the need to review LAN and Cloudflare exposure. They also combine, since LlamaFactory names unsloth among its training tricks. Before committing, verify your exact model and quantization method in LlamaFactory's supported table, and verify unsloth's model support and remote access defaults on your own hardware.

Sources

  1. hiyouga/LlamaFactory repository
  2. hiyouga/LlamaFactory README
  3. unslothai/unsloth repository
  4. unslothai/unsloth README