LMFlow: A Toolkit for Finetuning and Inferencing Large Language Models
An Extensible Toolkit for Finetuning and Inference of Large Foundation Models. Large Models for All.
At a glance
- What is it?
- LMFlow is an extensible Python toolkit for finetuning and running inference on large language models. It supports LoRA, LISA memory-efficient training, RAFT alignment, speculative decoding, and custom optimizers under a single interface, and is released under Apache-2.0.
- Who is it for?
- LMFlow suits researchers and ML engineers who need a configurable finetuning toolkit for large language models and want a single codebase that covers supervised finetuning, alignment, and inference acceleration. It is not the right choice for teams that need a graphical experiment tracking dashboard out of the box; that role belongs to tools like MLflow or Weights and Biases, which LMFlow integrates with as an optional dependency.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 52 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What LMFlow Solves and Who It Targets
LMFlow addresses the gap between having access to a pre-trained foundation model and having a production-ready, task-specific model. Training a large language model from scratch is resource-intensive and unnecessary for most applications. Finetuning an existing model on domain-specific data or task-specific examples is the practical alternative, and LMFlow provides the tooling to do that without writing low-level training loops.
The toolkit targets ML researchers and engineers working with models in the 1B to 70B parameter range. It was designed to be user-friendly in the sense that common workflows such as LoRA finetuning, reward modeling, and chat inference should run without deep familiarity with the underlying distributed training infrastructure. The README positions it as extensible and reliable, meaning teams can add custom pipelines without forking the core.
A paper accompanying the project was published at NAACL 2024 and received the Best Demo Paper Award at that conference, which indicates it has been peer-reviewed in an academic context.
Architecture: Pipelines, Datasets, and Supported Models
LMFlow is organized around the concept of pipelines. Different training objectives correspond to different pipeline modules: supervised finetuning, reward modeling, RAFT alignment, DPO, and inference each have a dedicated pipeline class in the src/lmflow directory. This structure means adding a new training approach involves adding a new pipeline class rather than modifying a monolithic training script.
The dataset handling follows the Hugging Face datasets format with version 3.6.0 pinned in requirements.txt. Conversation templates for popular models are built in, covering Llama-3, Phi-3, and the chatml format, among others. The `--conversation_template` flag in training scripts activates the appropriate template for the target model.
The toolkit integrates with Accelerate for distributed training (added in v1.0.0, released 2025-07-11), vLLM for high-throughput inference, SGLang for structured generation, and Ray for parallelism. These are declared as optional dependencies in setup.py under separate extras: `vllm`, `sglang`, `ray`, and others. A full-featured setup requiring all integrations will have a substantially larger dependency tree than a minimal finetuning-only installation.
Installing LMFlow and Running a Finetune
LMFlow is available on PyPI. The README documents the install command:
pip install lmflow-finetuneThe core dependencies, as listed in requirements.txt, include torch 2.0.1 or later, transformers 4.31.0 or later, peft 0.10.0 or later, and accelerate 0.27.2 or later. Optional integrations for vLLM, SGLang, Ray, Gradio, Flash Attention, DeepSpeed, and TRL are declared as separate extras in setup.py under the extra_require dictionary. Install only the extras your workflow requires to keep the environment lean.
Finetuning scripts are collected in the examples directory: examples/finetune.py for supervised finetuning, examples/dpo_train.py for DPO, and examples/reward_modeling.py for reward modeling, among others. Custom optimizer support is available through the script referenced in the README at scripts/run_finetune_with_custom_optim.sh. Training progress integrates with wandb when the dependency is present; it is listed in requirements.txt as an unconditional dependency.
LISA, LoRA, and Memory-Efficient Training
The most practically significant memory technique in LMFlow is LISA (Layerwise Importance Sampling for Alignment), added in March 2024 according to the README news section. LISA enables finetuning a 7B-parameter model within 24 GB of GPU memory without gradient checkpointing or CPU offloading. The underlying paper is linked as arxiv.org/abs/2403.17919 in the repository.
LoRA (Low-Rank Adaptation) is available through the peft integration. LoRA reduces the number of trainable parameters by adding small rank-decomposition matrices to existing weights instead of updating all parameters. The peft library handles this at the model level, and LMFlow wraps it as a first-class configuration option.
Flash Attention-2 support was added in August 2023, allowing models that support it to use the memory-efficient attention implementation for longer context windows and faster throughput. The flash_attn extra in setup.py pulls the flash-attn package at version 2.0.2 or later.
Alignment Beyond SFT: RAFT and DPO
LMFlow implements RAFT (Reward Ranked Finetuning), an alignment algorithm the team developed as an alternative to conventional RLHF with PPO. The paper is at arxiv.org/abs/2304.06767. RAFT works by generating candidate outputs from the model, ranking them with a reward model, and finetuning on the top-ranked outputs. The README describes it as more efficient than PPO-based RLHF, but the paper is the proper source for evaluating that claim.
DPO (Direct Preference Optimization) and its iterative variant are also supported, with dedicated example scripts at examples/dpo_train.py and examples/iterative_dpo_train.py. Reward modeling is a separate pipeline with its own script at examples/reward_modeling.py.
Speculative decoding was added in September 2023. It uses a smaller draft model to propose token sequences that a larger verifier model accepts or rejects in parallel, which can accelerate inference throughput for compatible model pairs.
Limitations and When LMFlow Is Not the Right Tool
LMFlow is not an experiment tracking tool. It does not provide a dashboard for comparing run metrics, visualizing training curves, or registering models. Those functions belong to tools like Weights and Biases (which LMFlow logs to when installed) or experiment tracking systems built on top of MLflow. The name similarity between LMFlow and MLflow is a persistent source of confusion: MLflow is a lifecycle management platform for ML experiments, not a finetuning toolkit.
The v1.0.0 release notes state that it added full Accelerate support and streamlined the codebase significantly, and recommend using `git checkout v0.0.10` for the previous version. This means code written against pre-v1.0.0 APIs may break on the current release.
The requirements.txt pins specific versions for some dependencies, such as datasets==3.6.0 and evaluate==0.4.0, which can create dependency conflicts in environments that also use other Hugging Face libraries at newer versions. Teams integrating LMFlow into a broader ML pipeline should audit requirements carefully.
LMFlow vs. Axolotl
Axolotl is another popular open-source LLM finetuning toolkit. Both support LoRA, full finetuning, and a range of model architectures. The differences lie in configuration style and scope.
Axolotl emphasizes a YAML configuration file as the primary interface, allowing users to define the entire training run without writing Python code. LMFlow's primary interface is shell scripts calling Python example files, which is more flexible for custom pipelines but less declarative for standard finetuning jobs.
LMFlow includes inference acceleration techniques (speculative decoding, vLLM integration) and alignment algorithms (RAFT, DPO) as first-class features. This wider scope makes it more appropriate for research workflows where training and inference are iterated together. Axolotl focuses more narrowly on the finetuning step itself. Teams that want a drop-in finetuning configuration tool will find Axolotl's YAML model more approachable. Teams that need to extend or customize the training pipeline in Python will find LMFlow's architecture more suited to that work.
Maintenance, License, and Python Support
LMFlow is released under the Apache-2.0 license, which permits use in commercial products without the source-disclosure requirement that AGPL carries. The pyproject.toml sets a target of Python 3.9, and the requirements.txt enforces torch 2.0.1 or later.
The last push to the repository was on 2026-08-10. v1.0.0 was released on 2025-07-11 and represents a major refactoring with full Accelerate support. v0.0.8, the previous stable release, was published in June 2024.
The README notes that the codebase is tested against GPU workloads. The pytest configuration in pyproject.toml marks GPU-dependent tests with a `gpu` marker, meaning a standard test run without GPU hardware will skip those tests. Developers who want to validate the full test suite need GPU access.
Editorial conclusion
LMFlow suits researchers and ML engineers who need a configurable finetuning toolkit for large language models and want a single codebase that covers supervised finetuning, alignment, and inference acceleration. It is not the right choice for teams that need a graphical experiment tracking dashboard out of the box; that role belongs to tools like MLflow or Weights and Biases, which LMFlow integrates with as an optional dependency. Before starting, verify that your GPU memory meets the requirements for your target model size: the README documents that LISA enables 7B model training within 24 GB of GPU memory without offloading, which is a concrete constraint to check against your hardware.
Frequently asked questions
How do I install LMFlow?
LMFlow is available on PyPI as lmflow-finetune. Run `pip install lmflow-finetune` for the base package. Optional integrations such as vLLM, SGLang, DeepSpeed, and Flash Attention are available as extras, for example `pip install lmflow-finetune[vllm]`.
What is the difference between LMFlow and MLflow?
LMFlow is a finetuning and inference toolkit for large language models. MLflow is an experiment tracking and lifecycle management platform for machine learning. They are unrelated projects that serve different functions, though LMFlow logs training metrics to Weights and Biases when that package is installed.
What GPU memory is required to finetune a 7B model with LMFlow?
The README documents that the LISA technique enables finetuning a 7B model within 24 GB of GPU memory without CPU offloading. Standard full finetuning without LISA will require significantly more memory.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/optimalscale-lmflow)