ms-swift: one Python CLI for SFT, DPO and GRPO across 600+ LLMs
Use PEFT or Full-parameter to CPT/SFT/DPO/GRPO 600+ LLMs (Qwen3.6, DeepSeek-V4, GLM-5.1, InternLM3, Llama4, ...) and 300+ MLLMs (Qwen3-VL, Qwen3-Omni, InternVL3.5, Ovis2.5, GLM4.5v, Gemma4, Llava, Phi4, ...) (AAAI 2025).
At a glance
- What is it?
- ms-swift is the ModelScope community's fine-tuning and deployment framework, covering text and multimodal models with LoRA, full-parameter and Megatron paths. It is broad by design, which is also where its sharp edges are.
- Who is it for?
- Adopt ms-swift if you already live in the ModelScope or Hugging Face ecosystem and want one CLI to move from LoRA SFT to DPO or GRPO without rewriting your training stack, and if you accept that the documentation is spread across a readthedocs site and an examples/ tree rather than a single tutorial. Do not adopt it if you need a small, auditable training loop you can read end to end, or if your hardware is not covered by the listed backends.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository received new commits within the last day.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What ms-swift actually solves for a team with one GPU budget
Fine-tuning a large model in 2026 is not one problem, it is five that share a checkpoint. You need to load weights that may be sharded or quantized, pick a parameter-efficient method so the job fits in memory, feed it data in a format the trainer understands, run the optimization, and then get a servable artifact out. Most teams assemble that from Hugging Face Transformers, PEFT, TRL, DeepSpeed and a serving engine, and the glue code is where the weeks go.
ms-swift, maintained by the ModelScope community, is an attempt to make that one Python package. The README describes it as a framework for training, inference, evaluation, quantization and deployment of 600+ text-only large models and 400+ multimodal large models. The task list is deliberately wide: CPT and SFT, preference methods including DPO, KTO, RM, CPO, SimPO and ORPO, the GRPO family of reinforcement learning algorithms, plus Embedding, Reranker and sequence classification. It is for the engineer who has a model and a dataset and wants a working run before deciding whether to build something bespoke.
The honest framing is that ms-swift is a breadth play. It would rather support your model on day zero than be the smallest possible trainer. That choice shows up everywhere, including in the places where it costs you.
The mechanism: one config surface over several training backends
The repository layout tells you most of the architecture. There is a swift/ package holding the library, an examples/ tree split by intent (train, infer, eval, export, deploy, megatron, ascend, ray), a requirements/ directory that setup.py reads through requirements.txt, and a Makefile whose targets are whl, docs, linter, test and clean.
The design is a single entry point that dispatches to a backend. For ordinary fine-tuning the backends are DDP, device_map model parallelism, DeepSpeed ZeRO2/ZeRO3 and FSDP/FSDP2. For large MoE work there is a separate Megatron path with TP, PP, SP, CP, ETP, EP and VPP parallel strategies, exposed through the examples/megatron/ directory. Memory-side options are listed as GaLore, Q-Galore, UnSloth, Liger-Kernel, Flash-Attention 2/3, and Ulysses and Ring-Attention sequence parallelism.
The data flow is the part worth internalizing. The README states there are 150+ built-in datasets for pre-training, fine-tuning, alignment and multimodal tasks, and that users prepare a dataset once and train. A template layer sits between the dataset and the model, which is how the project claims one dataset can train different models, and how Agent templates work. Multimodal packing is offered as a speed feature, mixing text, images, video and audio in one batch.
Inference and evaluation are not bolted on afterwards: vLLM, SGLang and LMDeploy are named as acceleration engines for inference, deployment and evaluation, and EvalScope is the evaluation backend behind 100+ evaluation datasets. Quantization export covers AWQ, GPTQ, FP8 and BNB. The consequence is that a single configuration vocabulary has to describe all of this, and the examples/yaml/ directory exists because that vocabulary is large.
Installing ms-swift and running a first LoRA SFT
The README points at PyPI for the package and the readthedocs site for the full installation matrix. The Python badge in the README says 3.12, PyTorch is listed as >=2.0, and ModelScope as >=1.23, so check those three before anything else.
The straightforward install is the published wheel:
pip install ms-swiftThe repository also carries a requirements.txt whose only line is a reference to requirements/framework.txt, and setup.py parses that file. If you are installing from a checkout rather than the wheel, that indirection is where the dependency list actually lives. The Makefile builds distributions with python setup.py sdist bdist_wheel under the whl target.
For a first real run, the examples/ tree is the reference. It is organized by task, with examples/train/ for training and examples/yaml/ for configuration files. The project's own framing is that you prepare a dataset and then train, so the practical first step is to pick one of the built-in datasets rather than writing your own loader. Multimodal and Megatron runs have their own directories (examples/megatron/), and Ascend NPU users have examples/ascend/, which is a useful signal that the NPU path is not the same code path as the CUDA one.
Two things to verify rather than assume. First, the README's installation section is the authority on optional extras; the wheel alone will not give you every backend, and the requirements/ directory is where the split is recorded. Second, if you intend to use the Web-UI, note that the README lists it as a capability but the examples/ tree does not carry a directory named for it, so read the docs page before planning around it.
Where ms-swift gets in your way
Breadth has a cost, and the most visible one is documentation surface. The README is a feature catalogue, not a tutorial. It tells you that 600+ models, 150+ datasets and a dozen RL algorithms exist, and then hands you a readthedocs link. The examples/ tree is the real documentation, and it is organized for people who already know which task they are running. If you do not yet know whether your job is CPT or SFT, the repository will not decide for you.
The second constraint is hardware. The README lists A10/A100/H100, RTX series, T4/V100, AMD MI300 series, CPU, MPS and Ascend NPU. That is a long list, but the support is not uniform: examples/ascend/ existing as a separate directory implies the NPU path has its own setup, and quantized training (BNB, AWQ, GPTQ, AQLM, HQQ, EETQ) is a distinct mode from ordinary LoRA. The claim that a 7B model can train in 9GB is attached to quantized training specifically, not to every method in the list.
Third, the reinforcement learning surface is genuinely large. GRPO, DAPO, GSPO, SAPO, CISPO, CHORD, RLOO and Reinforce++ are all named. Choosing among them is a research decision, and the framework does not make it for you. If you wanted a tool that picks a sensible default and hides the rest, this is the wrong shape.
Finally, the version cadence is fast. Releases v4.5.0, v4.5.2 and v4.5.3 landed between 2026-08-14 and 2026-09-08, and the last push to main was 2026-09-09. That is healthy for support of new models and painful for reproducibility: pin your version and your requirements file together, or a re-run three months later will not match.
How ms-swift differs from LLaMA-Factory, Unsloth and verl
The three comparisons people actually search for each expose a different axis.
Against LLaMA-Factory, the closest overlap is the all-in-one fine-tuning surface: both wrap SFT and preference optimization behind configuration rather than code. The difference in approach is scope of the surrounding pipeline. ms-swift's README claims training, inference, evaluation, quantization and deployment as one workflow, with vLLM, SGLang and LMDeploy named as the serving engines and EvalScope as the evaluation backend. LLaMA-Factory's center of gravity is the training and the UI. If your need stops at producing an adapter, the distinction matters less than which one supports your exact model today.
Against Unsloth, the difference is a single-machine optimization versus a distributed framework. Unsloth appears inside ms-swift as one of several memory optimizations, alongside GaLore, Liger-Kernel and Flash-Attention. That placement is the argument: ms-swift treats kernel-level speed work as a component you can switch on, while Unsloth treats it as the product. For one GPU and one model, the focused tool is often the faster path to a result. For a fleet of models and a Megatron MoE job, the framework is the only one of the two that has an answer.
Against verl, the split is reinforcement learning. Both are named for the GRPO family, but ms-swift's RL support sits inside a framework whose primary job is supervised fine-tuning and deployment, with RL algorithms as one task among many. verl is built around the RL loop. If post-training with reward models is the whole project, evaluate both on the algorithm list you need rather than on the framework's other features.
One more comparison worth naming: TRL. It is the reference implementation most people meet first, and it is a library you import into your own script, whereas ms-swift is a CLI and configuration system you drive. That is a workflow preference, not a quality ranking.
Licence, upgrade cost and what the repository commits to
ms-swift is Apache-2.0, the same licence as many of the frameworks it competes with. For most teams that means permissive use with the usual obligations around notices and attribution; the LICENSE file at the repository root is the text that governs, and this is not legal advice. One thing worth checking separately: the licence covers ms-swift itself, not the model weights you download through it. Those carry their own terms, and the README's model list spans several vendors.
The upgrade cost is a function of the release cadence. Three releases in roughly three weeks (v4.5.0 on 2026-08-14, v4.5.2 on 2026-08-16, v4.5.3 on 2026-09-08) means new model support arrives quickly and so does churn. The repository gives you the tools to manage that: a requirements/ directory with a framework file, a setup.py that reads it, and a Makefile with a test target that runs .dev_scripts/citest.sh. Pinning the wheel version and the requirements file together is the difference between a reproducible run and a mystery.
Maintenance is not in question on the evidence available. The repository is not archived, and the last push was on 2026-09-09. What the repository does not document is a deprecation policy or a compatibility guarantee between minor versions, so treat any config you write as tied to the version you wrote it against.
Editorial conclusion
Adopt ms-swift if you already live in the ModelScope or Hugging Face ecosystem and want one CLI to move from LoRA SFT to DPO or GRPO without rewriting your training stack, and if you accept that the documentation is spread across a readthedocs site and an examples/ tree rather than a single tutorial. Do not adopt it if you need a small, auditable training loop you can read end to end, or if your hardware is not covered by the listed backends. Before committing, verify three things against your own setup: that your target model appears in the supported model list for the task you want (SFT, DPO, GRPO and Embedding are not the same list), that your GPU or NPU appears under hardware support, and that the pinned PyTorch and ModelScope versions in requirements/ match what your cluster already runs. The repository's last push was on 2026-09-09 and the newest release is v4.5.3 from 2026-09-08, so the moving target you are pinning against is that pair of dates.
Frequently asked questions
What is ms-swift?
It is a fine-tuning and deployment framework from the ModelScope community, covering training, inference, evaluation, quantization and deployment for 600+ text-only large models and 400+ multimodal large models. The README describes it as a scalable lightweight infrastructure for fine-tuning, with tasks spanning CPT, SFT, DPO, KTO, RM, the GRPO family, Embedding, Reranker and sequence classification.
How does ms-swift compare with Unsloth?
Unsloth appears inside ms-swift as one of several memory optimizations, listed alongside GaLore, Q-Galore, Liger-Kernel and Flash-Attention. That placement reflects the difference in approach: ms-swift is a framework that switches such components on, while Unsloth is a focused single-GPU optimization tool.
How does ms-swift compare with LLaMA-Factory?
Both wrap fine-tuning and preference optimization behind configuration rather than code. The distinction the README draws is pipeline scope: ms-swift claims training, inference, evaluation, quantization and deployment as one workflow, naming vLLM, SGLang and LMDeploy for serving and EvalScope for evaluation.
How does ms-swift compare with verl for reinforcement learning?
Both cover the GRPO algorithm family, but ms-swift positions RL as one task among many inside a framework whose primary job is fine-tuning and deployment. verl is built around the RL loop itself. If reward-based post-training is the entire project, compare the two on the specific algorithms you need.
Does ms-swift support multiple programming paradigms?
ms-swift is a Python framework with a command-line and configuration interface rather than a multi-paradigm language. The repository is written in Python, and training runs are driven through YAML configs under examples/yaml/ and the swift/ package.
How do I debug ms-swift runs in VS Code?
The repository does not document a VS Code debugging setup. What it does provide is a Makefile with a test target that runs .dev_scripts/citest.sh and a linter target that runs .dev_scripts/linter.sh, which are the entry points the project uses for its own checks.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/modelscope-ms-swift)