← Open Source
DaoyuanLi2816

can-i-finetune-this

Single-GPU LLM fine-tuning preflight: memory estimates, runnable LoRA/QLoRA recipes, and local measurements.

Model DevelopmentFine-tuningPython
Open on GitHub
Momentum
+0stars in 24 hours0.0%
793
Stars
107
Forks
+1
This week
1
Contributors
Created 2026-05-16 · Updated 2026-10-05 · #12982 today
Top developers
README

canifinetune — Can I fine-tune this LLM on my GPU? Estimate a memory budget, measure local training, and generate a runnable recipe.

Documentation · Quickstart · PyPI

CI PyPI 0.4.1 Docs Python License: MIT

Can I fine-tune this LLM on my GPU? Make a plan before loading the weights.

You have one consumer NVIDIA GPU and an open-weight model in mind. How much memory will training need? Which sequence length, batch size and LoRA rank should you start with? What should you change when the budget is tight?

canifinetune turns those questions into a memory breakdown, configuration suggestions and a runnable recipe. Then it helps you measure an actual local run and compare the observation with the plan. Core estimation needs no PyTorch and no model-weight download.

Your first result

Start in a virtual environment; you do not need to clone this repository.

python -m pip install canifinetune==0.4.1
canifinetune estimate --model Qwen/Qwen2.5-1.5B-Instruct --method qlora --gpu-vram-gb 16 --seq-len 2048 --offline
canifinetune demo

The estimate is 8.420 GiB, YES against a 16 GiB budget, with heuristic confidence medium. It includes the stock logits/loss allocation and a separate safety allowance. A planning result is not a guarantee that training will fit.

demo serves a local interactive entry at 127.0.0.1:8765. Choose a model, training settings and total/currently-free memory; inspect the breakdown and copy matching CLI commands. It uses the same Python estimator.

Local preflight demo: model/configuration form, 8.42 GiB memory breakdown and matching commands

Choose your path

You want to… Start here What you get
Check a memory budget Estimate and recommend A component breakdown, assumptions and candidate configurations
Run a real fine-tune PyPI-only quickstart A pinned 0.5B QLoRA recipe, real updates, an adapter and a reload check
Try training without Hub weights Offline CPU smoke A locally created tiny model and full update/save/reload execution
Understand or contribute evidence Measurements and limits Defined metrics, raw observations and opt-in redacted export

How it works

Model metadata, GPU memory and one shared configuration feed the estimator and recommender; installed recipes and benchmarks load weights, record execution and produce reviewable evidence.

One configuration contract connects the plan to execution. The estimate accounts for weights, quantization, trainable gradients, optimizer state, activations, logits/loss and overhead. Generated recipes use one Transformers/PEFT runtime, preserve system and multi-turn data, and record requested/effective settings. Benchmarks cover loading, first optimizer-state allocation and the bounded updates.

  • Inspect the budget. See where memory goes, including large-vocabulary loss buffers and the unquantized parts of QLoRA models.
  • Carry the configuration through. Dtype, attention, targets, optimizer, checkpointing, quantization and supervision have explicit meanings and errors.
  • Get usable artifacts. Recipes save a full model or adapter through a staged save, record the outcome and provide a reload/generation command.
  • Keep evidence traceable. Reports separate measured peaks, planning safety, historical fits, independent observations and unreviewed community submissions.

Architecture guide · Training/data contract

What has been measured

On a native Windows RTX 4080, four predeclared Qwen2.5-0.5B-Instruct cases covered LoRA/QLoRA at sequence 256/512, with three updates each. Reserved peaks ranged from 1.666 to 2.398 GiB. Old/new estimator predictions were identical and conservative: 46.2% MAPE, four overestimates, no demonstrated accuracy gain.

This is a small prospective cohort on one GPU/model, separate from historical measurements used during estimator development. It does not establish a general OOM probability. Candidate/public-package qualification additionally executes CPU full/LoRA and CUDA LoRA/QLoRA updates, saves and reloads. Tiny smoke proves the pipeline, not useful fine-tuned language quality.

Prospective validation and raw records · Historical baselines

Install and compatibility

Layer Install Scope
Core pip install canifinetune==0.4.1 Python 3.10–3.14; estimate, recommend, recipes, reports and local demo
Training Qualified Torch wheel, then pip install "canifinetune[train]==0.4.1" with constraints Python 3.12; real training and benchmarks
Reporting extras pip install "canifinetune[report]==0.4.1" Optional pandas/tabulate

Training uses Torch 2.6, minimum/recommended Transformers/PEFT/Accelerate stacks and bitsandbytes 0.49.2. The quickstart separates Windows and Linux/WSL commands. Native Windows CPU/CUDA and Linux CPU are qualified; WSL GPU, other GPUs, Flash Attention and Liger are not qualified. CPU requires fp32; pre-quantized bases, remote model code and distributed training are outside supported execution. Inference artifacts do not include full optimizer/RNG resume state.

Compatibility and migration · Troubleshooting

Contribute and develop

Improve a documented model family, reproduce a bounded measurement, or make the first-use workflow clearer. Measurements are manual and opt-in, redacted by default and unreviewed until checked; uploads do not automatically enter fits.

python -m pip install -e ".[dev]"
ruff check .
ruff format --check .
mypy src
pytest -q -m "not training" --cov=canifinetune --cov-fail-under=70
python scripts/check_generated.py

Contributing · Changelog · Release process · License

MIT. Maintainer: Daoyuan Li.