Hugging Face PEFT: Parameter-Efficient Fine-Tuning for Models You Cannot Fully Train
🤗 PEFT: State-of-the-art Parameter-Efficient Fine-Tuning.
At a glance
- What is it?
- PEFT wraps a pretrained model and trains only a small set of added parameters, so a 12B model can be adapted on a single 80GB GPU. Here is the mechanism, a first LoRA run, and where it stops being the right tool.
- Who is it for?
- Adopt PEFT when full fine-tuning of your base model does not fit the hardware you have, and when you want one frozen base plus swappable adapters per task. Skip it when you need to change the base weights themselves, when your task needs new vocabulary or a new architecture head that the adapter config does not cover, or when you cannot accept an extra inference hop through the adapter.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository received new commits within the last day.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The memory wall that PEFT was built to climb
Full fine-tuning updates every weight in a pretrained model. For a 3B parameter model that means optimizer state and gradients on top of the weights, and the README's own table puts full fine-tuning of bigscience/T0_3B at 47.14GB of GPU memory on an A100 80GB. The same table shows bigscience/mt0-xxl at 12B parameters running out of memory entirely under full fine-tuning. That is the problem PEFT addresses: adapt a large pretrained model to a downstream task while training only a small number of extra parameters, so the compute and storage cost of the run drops with it. The audience is anyone who has a model too large to fine-tune conventionally but still needs task-specific behaviour: a single GPU owner, a team that wants many task variants of one base model, or a shop that cannot justify one full checkpoint per task. The README frames the payoff as performance comparable to fully fine-tuned models, and its own comparison table backs that claim with lora-t0-3b at 0.863 accuracy against a 0.897 human baseline and 0.892 for Flan-T5.
How an adapter sits on top of a frozen base model
The core mechanism is a wrapper, not a rewrite. You keep the pretrained model as it is and attach a small set of trainable parameters whose configuration comes from a PEFT config object. In the LoRA path shown in the README, LoraConfig carries the rank r, a scaling factor lora_alpha, the task_type, and an optional target_modules list naming which internal modules get the adapter. get_peft_model then returns a model whose trainable parameter count is a fraction of the total, and model.print_trainable_parameters() reports that fraction directly. The README's example for Qwen/Qwen2.5-3B-Instruct with r=16 prints 3,686,400 trainable parameters out of 3,089,625,088 total, or 0.1193%. Because the base weights are untouched, the artifact you save is the adapter, not a copy of the model: the README states the final checkpoint for the T0_3B example is 19MB against 11GB for the full model. Loading reverses the step. PeftModel.from_pretrained takes a base model you have already loaded and a path to the adapter directory, and the result generates like the adapted model. This is why the library plugs into Transformers for training and inference, Diffusers for managing adapters, and Accelerate for distributed runs.
Install and run a first LoRA fine-tune
The README's quickstart begins with a single pip install. Nothing else is required to get the library into an environment that already has PyTorch and Transformers.
pip install peftThe next step loads a base model and wraps it with a LoRA config. The README uses Qwen/Qwen2.5-3B-Instruct and picks the accelerator through torch.accelerator when it exists, falling back to cuda. The target_modules line is commented out in the README, which means the default targeting applies unless you name modules yourself. After the call, print_trainable_parameters() is what tells you whether the configuration did what you expected.
import torch
from transformers import AutoModelForCausalLM
from peft import LoraConfig, TaskType, get_peft_model
device = torch.accelerator.current_accelerator().type if hasattr(torch, "accelerator") else "cuda"
model_id = "Qwen/Qwen2.5-3B-Instruct"
model = AutoModelForCausalLM.from_pretrained(model_id, device_map=device)
peft_config = LoraConfig(
r=16,
lora_alpha=32,
task_type=TaskType.CAUSAL_LM,
)
model = get_peft_model(model, peft_config)
model.print_trainable_parameters()Training itself is not part of the PEFT call. The README says to perform training on your dataset, for example using the Transformers Trainer, and then save. The save writes an adapter directory, here named qwen2.5-3b-lora.
model.save_pretrained("qwen2.5-3b-lora")For inference you load the base model again and attach the adapter. The README's generation example tokenizes a prompt, calls generate with max_new_tokens=50, and decodes the output with skip_special_tokens=True.
from transformers import AutoModelForCausalLM, AutoTokenizer
from peft import PeftModel
model_id = "Qwen/Qwen2.5-3B-Instruct"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id, device_map=device)
model = PeftModel.from_pretrained(model, "qwen2.5-3b-lora")
outputs = model.generate(**inputs.to(device), max_new_tokens=50)The thing to check after this sequence is the trainable percentage printed by the first script. If it is not a small fraction, the adapter is not attached where you think it is.
Quantized bases, offloading, and what the memory table actually shows
PEFT is designed to combine with quantization, which lowers the precision of the base model's weights and cuts memory further. The README points to a QLoRA walkthrough for meta-llama/Llama-2-7b-hf using the TRL library on a 16GB GPU, and to a notebook fine-tuning openai/whisper-large-v2 with LoRA plus 8-bit quantization. The second axis is offloading. The README's table lists a PEFT-LoRA DeepSpeed column with CPU offloading: bigscience/mt0-xxl drops from 56GB GPU to 22GB GPU at the cost of 52GB of CPU RAM, and bigscience/bloomz-7b1 drops from 32GB to 18.1GB GPU with 35GB CPU. Those are the real trade-offs. You are not removing memory pressure, you are moving it, and the CPU side can become the constraint. The README also carries a caution worth repeating: the T0_3B numbers are not optimized, and it says you can get more out of the model by adjusting instruction templates and LoRA hyperparameters. Treat the table as an order-of-magnitude guide to what fits, not as a target you are guaranteed to hit.
Where PEFT is the wrong tool
The adapter approach has boundaries. Because the base weights stay frozen, anything that requires changing the base representation itself is out of scope; PEFT is for adapting a model that already knows the domain, not for teaching it a fundamentally new one. The README's own T0_3B row shows the gap: 0.863 for lora-t0-3b against 0.892 for Flan-T5 and 0.897 for the human baseline. Comparable is not identical, and the README itself says that result is not optimized, which cuts both ways: better hyperparameters may close the gap, but the default configuration did not. There is also an operational cost that the quickstart hides. Inference goes through PeftModel, so every serving path needs to know about the adapter, either by merging it into the base or by holding it alongside. Teams with a fixed inference stack that only accepts a plain Transformers checkpoint will need an extra step. Finally, the repository ships a large number of method-specific example directories, from examples/lora-style runs through examples/boft_dreambooth and examples/delora_finetuning, and each method has its own config and its own hyperparameters. Choosing among them is a research decision, not a drop-in swap, and the README does not rank them for you.
The alternative: full fine-tuning, and when it wins
The direct alternative is full fine-tuning with the Transformers Trainer and no PEFT layer at all. The approach differs at the level of what is trainable: full fine-tuning updates every weight and produces a complete checkpoint, while PEFT updates an added parameter set and produces an adapter that must be paired with its base. The consequence is storage and memory. The README's table makes the storage side concrete: 19MB for the T0_3B adapter against 11GB for the full model. If you need many task variants of one base, that ratio decides the question, because you ship one base plus N small adapters instead of N full copies. Full fine-tuning wins in the opposite case: when the task genuinely requires the base weights to move, when you have the GPU budget for it, and when you want a single self-contained artifact with no adapter-loading step in the serving path. The README's own table shows the ceiling of the alternative, with mt0-xxl running out of memory under full fine-tuning on an 80GB GPU.
Licence, releases, and what maintenance costs you
The project is Apache-2.0, and the README carries the standard Apache header with its warranty disclaimer. For most users that means permissive use with attribution and a patent grant, but the licence covers the library, not the pretrained models you attach it to; those carry their own terms, and the README's examples reference models such as meta-llama/Llama-2-7b-hf whose licences are separate. The repository is not archived, and the last push was on 2026-09-09, so this is a codebase with recent activity. Releases are frequent enough to plan around: v0.20.0 landed on 2026-07-28, with v0.19.1 and v0.19.0 before it in April 2026. The upgrade cost sits in the config surface. Because LoraConfig and the other config classes carry the hyperparameters, and because the repository ships a method-specific example directory for each technique, a version bump can change defaults or add methods without changing your call site. The development version in setup.py is 0.20.1.dev0, which means the main branch is ahead of the latest tagged release. If you pin, pin to a tag and read the release notes for the method you use rather than assuming the adapter you saved last quarter loads unchanged.
Editorial conclusion
Adopt PEFT when full fine-tuning of your base model does not fit the hardware you have, and when you want one frozen base plus swappable adapters per task. Skip it when you need to change the base weights themselves, when your task needs new vocabulary or a new architecture head that the adapter config does not cover, or when you cannot accept an extra inference hop through the adapter. Before committing, verify three things: that the target modules you pass to LoraConfig match the module names in your specific checkpoint, that your saved adapter directory reloads through PeftModel.from_pretrained on a fresh base model, and that your serving stack can merge or hold the adapter. The README's own numbers are the benchmark to beat: 0.1193% trainable parameters for Qwen2.5-3B with r=16, and a 19MB checkpoint against an 11GB full model.
Frequently asked questions
What is PEFT in LLM?
PEFT stands for Parameter-Efficient Fine-Tuning. In the LLM context it means adapting a large pretrained model by training only a small number of added parameters instead of all of the model's parameters, which the README describes as significantly decreasing computational and storage costs.
What is a PEFT package?
It is the Python library published as peft, installed with pip install peft. It provides the config classes and wrappers, such as LoraConfig and get_peft_model, that attach parameter-efficient adapters to a pretrained model, and it integrates with Transformers, Diffusers and Accelerate.
Is LoRA a PEFT technique?
Yes. LoRA is one of the methods implemented in the library, and the README's quickstart uses LoraConfig with get_peft_model as the primary example. The PEFT Adapters API Reference lists the supported methods.
What does hugging face transformers do?
Transformers is the library that loads the pretrained models PEFT adapts, through classes such as AutoModelForCausalLM and AutoTokenizer. The README states that PEFT is integrated with Transformers for easy model training and inference.
What is huggingface peft?
It is the peft library from Hugging Face, described in the README as state-of-the-art Parameter-Efficient Fine-Tuning methods. It lets you adapt large pretrained models by training only a small number of extra parameters, and it is integrated with Transformers, Diffusers and Accelerate.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/huggingface-peft)