# NVlabs/DoRA: Weight-Decomposed Low-Rank Adaptation for PyTorch Fine-Tuning

> DoRA splits a pre-trained weight into magnitude and direction and trains LoRA only on the direction. The repository is the official ICML 2024 implementation, and the practical entry point for most teams is the use_dora flag in HuggingFace PEFT.

**NVlabs/DoRA** — [ICML2024 (Oral)] Official PyTorch implementation of DoRA: Weight-Decomposed Low-Rank Adaptation

- Repository: https://github.com/NVlabs/DoRA
- Website: https://nbasyl.github.io/DoRA-project-page/
- Stars: 1,000 · Forks: 70
- Language: Python
- License: NOASSERTION
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/nvlabs-dora

## What DoRA changes about LoRA fine-tuning

LoRA freezes the pre-trained weight and adds a low-rank update to it. DoRA keeps the low-rank branch but changes what that branch is allowed to move. The README describes the method as decomposing the pre-trained weight into two components, magnitude and direction, then applying LoRA to the directional update so the number of trainable parameters stays small. The stated payoff is higher learning capacity and better training stability than LoRA, with no additional inference overhead, because the decomposition is folded back into a single weight at the end.

The audience is narrow and identifiable. This is for engineers who already run parameter-efficient fine-tuning and have hit the point where LoRA rank is the knob they keep turning. The repository reports results on LLaMA, LLaVA and VL-BART across commonsense reasoning, visual instruction tuning, and image or video-text understanding. If none of those model families or task shapes resemble your workload, the reported comparisons will not tell you much about your case.

## Magnitude, direction, and where the extra cost lands

The mechanism is a reparameterization, not a new layer type. A weight matrix is expressed as a magnitude scalar times a direction matrix, and the trainable adapter updates the direction through a low-rank product while the magnitude is trained separately. That is why the README can claim no inference overhead: at merge time the two components recombine into the original weight shape, so serving looks the same as serving a LoRA-merged model.

The cost sits in training. There is an extra term to optimize compared with plain LoRA, and the README's own guidance reflects that: it suggests starting from a slightly lower learning rate than LoRA, experimenting with LoRA dropout ratios, and trying roughly half the LoRA rank, which it says often gives comparable or better accuracy. Those are adjustments, not drop-in defaults. The README is explicit that while a LoRA configuration usually already improves results under DoRA, reaching the best numbers still requires tuning.

The repository is organized as four directories rather than a single training script: commonsense_reasoning, instruction_tuning_dvora, image_video_text_understanding, and visual_instruction_tuning, plus QDoRA. Each carries its own lineage. The commonsense code is modified from LLM-Adapter, the instruction tuning code from VeRA, and the README points to the QDoRA directory for FSDP reproduction steps. Expect to read four separate setups if you want to reproduce more than one result.

## Installing PEFT and turning on use_dora

For most users the repository itself is a reference implementation, not the thing you install. The README states that DoRA is supported by the HuggingFace PEFT package and gives this install command, which pulls PEFT from git rather than from a pinned release:

```bash
pip install git+https://github.com/huggingface/peft.git -q
```

After that, the README says to set the use_dora argument of LoraConfig to True. The example shown in the README is abbreviated, so treat the structure below as the shape of the configuration rather than a complete script, and fill in the remaining LoraConfig arguments for your model:

```python
from peft import LoraConfig

# Initialize DoRA configuration
config = (
    use_dora=True, ...
)
```

What you should see is a training run whose adapter behaves like a LoRA run but whose loss curve and final accuracy come from the decomposed update. The README's tuning note applies from the first run: start with a slightly lower learning rate than your LoRA setting, and consider halving the rank. The README also points to the official PEFT documentation for the full argument list, which is where the remaining parameters belong.

If you want the paper's own numbers instead of a PEFT integration, the README directs you to the four directories in the repository and, for QDoRA and FSDP, to the step-by-step instructions under QDoRA. Those paths are reproduction code, so budget time for reading each directory's inherited setup.

## Where DoRA is the wrong tool

Diffusion fine-tuning is the clearest weak spot, and the README says so itself: DoRA finetuning on diffusion models is described as still experimental and likely to need different hyperparameter values than LoRA to perform best. Two specific differences are called out. LoRA appears to converge faster than DoRA, so a parameter set that overfits a LoRA run may work fine for DoRA. And the quality gap over LoRA is reported as larger at low ranks, with rank 8 showing a more significant difference than ranks 32 or 64.

That second point cuts both ways. If you are already training at rank 32 or 64 and your results are acceptable, the README's own framing suggests the gain narrows, and you would be paying the extra tuning effort for less. The same logic applies to any workflow where your LoRA configuration is already validated and you cannot afford a hyperparameter sweep: the README states that reusing a LoRA configuration usually helps but does not reach optimal performance, so a straight substitution is not a free upgrade.

The README also does not document rollback or a way to convert an existing trained LoRA adapter into a DoRA one. If you have adapters in production, DoRA is a new training run, not a migration path.

## DoRA against LoRA as a design choice

The obvious alternative is LoRA itself, and the difference is structural rather than a matter of degree. LoRA adds a low-rank delta to a frozen weight. DoRA instead splits the frozen weight into magnitude and direction, and spends its trainable parameters on the direction. That is the whole reason the README can claim no inference overhead while also claiming better learning capacity: the extra expressiveness lives in how the update is parameterized, not in extra parameters at serving time.

The practical consequence is where each method wins. LoRA is simpler to tune because its rank and alpha map directly onto the update it produces. DoRA adds a second quantity to optimize, which is why the README's advice is about learning rate, dropout and rank rather than a single flag. If your bottleneck is engineering time rather than final accuracy, LoRA remains the lower-friction choice. If your bottleneck is accuracy at a rank small enough to keep adapter size and memory down, the README's low-rank observation is the argument for DoRA.

A second option the README mentions is DVoRA, which combines DoRA with VeRA, with code under instruction_tuning_dvora for LLaMA-7B and LLaMA2-7B on the cleaned Alpaca dataset. That directory is derived from the VeRA supplementary material, so its structure will look different from the commonsense code.

## Maintenance, licence, and what an upgrade costs

The repository is not archived, and the last push was on 2026-03-24. That is roughly six months before today, which puts it at the edge of the window where a project can be described as recently touched. There have been no retrieved releases, so there is no version number to pin and no changelog to read. Upgrades therefore happen through the directories themselves or through PEFT, and the README's install command pulls PEFT from git rather than a fixed version. That means two teams installing on different days can end up with different PEFT code, which is a reproducibility problem you should decide about deliberately, for example by pinning a commit yourself.

The licence field is NOASSERTION, which means GitHub could not identify a standard licence from the repository contents. The README routes business inquiries to NVIDIA Research Licensing. This is not legal advice, but the practical reading is that you should not assume Apache-2.0 or MIT terms apply, and that commercial use deserves a look at the LICENSE file and, where the stakes justify it, at the licensing contact the README provides. The repository also carries no stated support commitment, and the README does not document a deprecation policy for any of the four directories.

## Conclusion

Adopt DoRA when you are already fine-tuning with LoRA and want better quality at the same or lower rank, especially in the low-rank regime the README singles out. Skip it if you need a supported recipe for diffusion training, since the README labels that path experimental and warns that hyperparameters tuned for LoRA will not transfer unchanged. Before committing, verify the licence terms for your intended use, since the repository reports NOASSERTION rather than a named licence, and confirm on a small run whether half the LoRA rank and a slightly lower learning rate hold up on your data.

## FAQ

### What is DoRA in machine learning?

DoRA is weight-decomposed low-rank adaptation, a fine-tuning method that splits a pre-trained weight into magnitude and direction components and applies LoRA to the directional update. The README states it improves learning capacity and training stability over LoRA without adding inference overhead.

### What are the key differences between LoRA and DoRA?

LoRA adds a low-rank update to a frozen weight, while DoRA decomposes that weight into magnitude and direction and trains the direction through a low-rank branch. The README reports that DoRA outperforms LoRA on LLaMA, LLaVA and VL-BART fine-tuning, with the quality gap described as larger at low ranks.

### Is there something better than LoRA for parameter-efficient fine-tuning?

The README positions DoRA as a higher-performing alternative to LoRA, claiming consistent gains on commonsense reasoning, visual instruction tuning, and image or video-text understanding. The README also notes that reaching optimal performance with DoRA requires adjusting hyperparameters rather than reusing a LoRA configuration unchanged.

## Sources

- [Issues](https://github.com/NVlabs/DoRA/issues)
- [NVlabs/DoRA on GitHub](https://github.com/NVlabs/DoRA)
- [Project website](https://nbasyl.github.io/DoRA-project-page/)
- [README](https://github.com/NVlabs/DoRA/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/nvlabs-dora
