Model or dataset
DaoyuanLi2816/can-i-finetune-this avatar
DaoyuanLi2816/can-i-finetune-this

canifinetune: find out if your GPU can fine-tune a model before downloading it

Estimate whether a Hugging Face model fits and fine-tunes on your local GPU.

793 stars107 forksPythonMIT

At a glance

What is it?
A Python package that estimates LoRA and QLoRA VRAM use from a memory model fitted to real hardware, runs local benchmarks to calibrate the numbers, and generates a ready-to-run training recipe.
Who is it for?
canifinetune is a small tool with a well-chosen scope: it does not train models, it tells you whether your training run will fit, and it backs that verdict with a memory model whose two hardest terms, the logits chain and the QLoRA fp32 upcast, are the ones generic estimators leave out.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 1 day ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 8, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The question it answers before you spend the disk space

You have one consumer NVIDIA card, a 12 GB to 24 GB model, and a Hugging Face repo promising an open-weight LLM you would like to fine-tune with LoRA or QLoRA. The traditional workflow is painful: download the weights, wait for the training script to crash with CUDA out of memory on step one, guess at a smaller batch size, and repeat. canifinetune, a Python package by DaoyuanLi2816, exists to answer the boring but expensive question first.

It answers six of them, in order: can this model be fine-tuned on this card at all, roughly how much VRAM will the run take, which batch size and sequence length and LoRA rank and quantization would work, how to downsize if the answer is no, whether there is local benchmark evidence behind that verdict, and whether you can walk away with a ready-to-run Hugging Face plus PEFT plus TRL training script for the winning configuration. All six come from a single CLI, and none of them require PyTorch to be installed, because estimation is pure arithmetic over a fitted memory model.

Two memory terms most estimators skip

The interesting engineering is in what the estimate includes. The obvious components are easy: model weights in fp32, fp16, bf16, int8, or NF4 with double quantization, gradients for trainable parameters only, and AdamW or 8-bit or paged AdamW optimizer states. Where canifinetune separates itself is the pair of terms that static tools routinely leave out.

The first is the logits and cross-entropy chain, which scales with sequence length, batch size, and vocabulary at roughly 14 bytes per element. For Qwen2.5, whose vocabulary runs to 152k tokens at a 2048 sequence length, that chain alone accounts for about 4.1 GB, and gradient checkpointing does nothing to remove it because logits are not activations you can recompute for free. The second is the fp32 upcast that QLoRA applies: prepare_model_for_kbit_training quantizes only the Linear layers, while embeddings, lm_head, and norms get upcast to fp32, which the README notes costs 4 GB on an untied 7B model. Skip those two terms and your estimator will promise a comfortable fit and hand you an OOM instead.

Grounded in measurement, not just arithmetic

Every coefficient in the activation model was fitted against measured torch.cuda peaks on real hardware, and the documented baselines live in docs/rtx4080_baselines.md. The README gives a concrete calibration example: for Qwen2.5-1.5B-Instruct with QLoRA at sequence length 2048 on a 16 GB card, the estimator reports a total of 8.42 GB with feasibility ratio 0.53, while the same configuration on a real RTX 4080 peaks at 7.10 GB reserved. The estimate lands about 1.3 GB above the measurement, on the safe side.

That bias toward overestimating is deliberate. An estimator that promises 3 GB and OOMs at 5 has wasted your evening; one that says 8.4 when the truth is 7.1 has cost you nothing. Because activation memory is the hardest component to predict statically, every estimate ships with an assumptions block and a confidence level, and two commands exist to replace the generic coefficients with numbers from your own machine.

The CLI surface

The package exposes one command with subcommands, and the README walks through them in a natural working order:

bash
canifinetune doctor
canifinetune estimate --model Qwen/Qwen2.5-1.5B-Instruct --method qlora --gpu-vram-gb 16 --seq-len 2048 --micro-batch-size 1 --lora-rank 16
canifinetune recommend --model Qwen/Qwen2.5-1.5B-Instruct --gpu-vram-gb 16
canifinetune bench    --model sshleifer/tiny-gpt2 --method lora --steps 3
canifinetune calibrate --benchmarks benchmarks/results
canifinetune recipe   --model Qwen/Qwen2.5-1.5B-Instruct --method qlora --output recipes/qwen2.5-1.5b-qlora-4080

doctor inspects the local machine, estimate prints a memory breakdown table, recommend searches for a feasible configuration, bench runs a tiny real training run (tiny-gpt2 is about 5 MB) to produce measured peaks, calibrate folds those measurements back into the estimator, and recipe emits the training script for the config that survived. A report command and a compare command turn a directory of benchmark results into markdown.

What the estimate output looks like

The estimate command prints a box with the feasibility verdict and a component-by-component memory table. For the Qwen2.5-1.5B QLoRA example the breakdown covers the static model at 1.496 GB, quantization overhead, trainable parameters of 4.4 MB, gradients, optimizer states, activations at 0.689 GB, the logits and loss chain at 4.057 GB, CUDA and fragmentation reserve at 1.280 GB, and a 0.800 GB safety margin, summing to 8.420 GB.

The value of the table is diagnostic rather than decorative. If the total is over your budget, the line that dominates tells you which knob to turn: shrink sequence length to cut the logits chain, drop LoRA rank to cut trainable parameters and gradients, or accept NF4 quantization to shrink the static weights. The tool also generates concrete degradation suggestions automatically when a configuration is not feasible, which turns the verdict from a no into a plan.

Install layers and MoE awareness

Installation is split so that estimation stays light:

bash
pip install canifinetune
pip install canifinetune[train]

The core install gets you every CLI command with no PyTorch dependency. The train extra adds torch, transformers, peft, bitsandbytes, trl, and datasets for actual benchmarks and fine-tuning runs. There are report and dev extras as well, and uv users can install the package in editable mode with the extras of their choice. The README recommends installing PyTorch with the CUDA wheel matching your driver and points to a troubleshooting document covering Windows, WSL, and bitsandbytes specifics.

On the modeling side, the estimator reads exact parameter counts from Hub safetensors rather than estimating from architecture names, and it handles MoE layouts with separate accounting for local experts versus experts per token, so a mixture-of-experts checkpoint does not silently get billed as a dense model of the same nominal size.

How it compares and who it is for

The README draws the comparison directly: accelerate estimate-memory tells you what loading a model costs, which is not the same as what training it costs. Training adds gradients, optimizer states, the logits chain, and the quantization upcasts, and those are exactly the components that decide whether a 12 GB card survives step one. Anyone choosing between the two tools should ask which question they are actually asking.

The audience is the owner of one consumer GPU deciding what to run on it: the student with a 4060, the researcher with a borrowed 4090, the hobbyist deciding whether that 7B fine-tune is a weekend project or a fantasy. For them, the combination of a conservative estimate, a calibration loop against their own hardware, and a generated training script collapses a frustrating trial-and-error cycle into an afternoon.

Editorial conclusion

canifinetune is a small tool with a well-chosen scope: it does not train models, it tells you whether your training run will fit, and it backs that verdict with a memory model whose two hardest terms, the logits chain and the QLoRA fp32 upcast, are the ones generic estimators leave out. With a benchmark and calibration loop to ground the numbers in your own card and a recipe generator to hand off the winning configuration, it replaces the download-and-hope workflow with an estimate you can act on.

Frequently asked questions

Does canifinetune need PyTorch installed to run estimates?

No. The core pip install canifinetune package runs every CLI command including estimate, recommend, and recipe without PyTorch. The train extra adds torch, transformers, peft, bitsandbytes, trl, and datasets, and is only needed when you want to run real benchmarks or the generated training recipes.

Why did my estimator say the model fits but training still OOMed?

The usual culprits are the two terms static estimators skip: the logits and cross-entropy chain, which scales with sequence length times vocabulary at about 14 bytes per element and survives gradient checkpointing, and the fp32 upcast of embeddings and norms that QLoRA applies to non-Linear layers. canifinetune models both, which is why its totals run higher and safer than naive estimates.

How accurate are the VRAM estimates?

The coefficients were fitted against measured torch.cuda peaks on real hardware, with baselines documented for an RTX 4080. The documented QLoRA example estimates 8.42 GB total against a measured 7.10 GB peak, so the estimate sits about 1.3 GB above reality on the safe side. You can tighten accuracy on your own machine with the bench and calibrate commands.

Official sources

  1. DaoyuanLi2816/can-i-finetune-this on GitHub
  2. License: MIT
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/daoyuanli2816-can-i-finetune-this.svg)](https://hysenlabs.com/projects/daoyuanli2816-can-i-finetune-this)