In the QLoRA vs LoRA choice, both methods freeze the original model and train small add-on matrices called adapters. The difference is how the frozen model is stored: LoRA keeps it in 16-bit, while QLoRA compresses it to 4-bit, cutting memory roughly three to four times at the cost of slightly slower training and a small potential quality loss. Choose QLoRA when memory is tight, LoRA when you have the VRAM and want the most faithful result, DoRA when you want a little more quality from low-rank training, and full fine-tuning only when you have a large dataset and serious hardware.
The four methods in one table
| Method | What is trained | Frozen base stored as | Memory need | Best for |
|---|---|---|---|---|
| Full fine-tuning | Every weight | Not frozen | Highest | Large datasets, major behaviour changes, labs |
| LoRA | Small low-rank adapters | 16-bit | Moderate | Most fine-tunes when VRAM allows |
| QLoRA | Small low-rank adapters | 4-bit | Lowest | Consumer GPUs, laptops, free notebooks |
| DoRA | Adapters plus a magnitude vector | Usually 16-bit | Slightly above LoRA | Squeezing more quality out of low ranks |
How LoRA works
LoRA, short for Low-Rank Adaptation, was introduced in the 2021 paper LoRA: Low-Rank Adaptation of Large Language Models. Instead of updating a huge weight matrix directly, LoRA freezes it and learns two thin matrices whose product is added to it. The "rank" sets how thin those matrices are.
The paper's headline results: compared with fine-tuning GPT-3 175B with Adam, LoRA cut the number of trainable parameters by 10,000 times and GPU memory by three times, while matching or beating full fine-tuning quality on the models it tested. Because the adapter can be merged into the base weights after training, LoRA adds no extra latency at inference time.
In practice, the adapter is tiny. In the PEFT library's own quickstart, a LoRA with rank 16 on Qwen2.5-3B trains about 3.7 million parameters, roughly 0.12 percent of the model.
How QLoRA works
QLoRA came from the 2023 paper QLoRA: Efficient Finetuning of Quantized LLMs. It stores the frozen base model in 4-bit and backpropagates through it into LoRA adapters that stay in higher precision. The authors reported fine-tuning a 65-billion-parameter model on a single 48 GB GPU while preserving full 16-bit fine-tuning task performance.
Three tricks make that work:
- 4-bit NormalFloat (NF4), a data type designed for the bell-shaped distribution of neural network weights.
- Double quantization, which also compresses the scaling constants used by the quantizer.
- Paged optimizers, which smooth out memory spikes during training.
The trade-offs are that de-quantizing on the fly makes each training step somewhat slower, and the 4-bit base can lose a little accuracy. Tool makers have worked on that second point; Unsloth's fine-tuning guide says that with its dynamic 4-bit quants, the accuracy loss of QLoRA relative to LoRA is now largely recovered.
How DoRA works
DoRA, or Weight-Decomposed Low-Rank Adaptation, was proposed in a 2024 paper. It splits each pretrained weight into a magnitude and a direction, uses LoRA to update the direction, and trains the magnitude separately. The authors report that it consistently outperformed LoRA on the language and vision-language tasks they tested, with no extra inference overhead once merged.
In PEFT you turn it on with use_dora=True. The PEFT documentation notes that DoRA adds more training overhead than plain LoRA, works on linear and Conv2D layers, and should be merged for inference. Apple's mlx-lm also supports it, with --fine-tune-type dora.
Memory: real numbers
Memory is usually what decides the method. Unsloth publishes minimum VRAM figures for its own optimised training stack, shown here for a few sizes. Other tools use more, and longer sequences or bigger batches add to these numbers.
| Model size | QLoRA, 4-bit | LoRA, 16-bit |
|---|---|---|
| 7B | 5 GB | 19 GB |
| 8B | 6 GB | 22 GB |
| 14B | 8.5 GB | 33 GB |
| 27B | 22 GB | 64 GB |
| 32B | 26 GB | 76 GB |
| 70B | 41 GB | 164 GB |
Source: Unsloth requirements page, checked October 5, 2026.
In words: an 8B model that needs about 22 GB for LoRA fits in about 6 GB with QLoRA, which is the difference between a data-centre card and a mid-range gaming GPU.
Full fine-tuning is far heavier. The ZeRO paper explains that mixed-precision training with the Adam optimizer needs about 16 bytes per parameter for weights, gradients and optimizer states, before counting activations. For an 8B model, that is roughly 128 GB.
When QLoRA is not the right choice
QLoRA is not always recommended, even when it fits:
- Some new architectures. Unsloth's Qwen3.5 fine-tuning guide recommends 16-bit LoRA rather than QLoRA for that family, and says 4-bit QLoRA for its mixture-of-experts models is not recommended because of bitsandbytes limitations.
- When you have the memory anyway. If LoRA fits, it trains faster per step and avoids any quantization error.
- When you will serve in 16-bit. Training on a 4-bit base and serving a 16-bit merged model introduces a small mismatch. Newer tools address this by training directly against quantized formats; Axolotl, for example, added merge-aware NVFP4 LoRA training in 2026.
Settings that matter more than the method
Whichever method you pick, a few settings drive results:
- Rank (r). Unsloth's hyperparameter guide suggests 16 or 32 as a starting point. Higher ranks add capacity and memory, and can overfit.
- Alpha. Set it equal to the rank, or twice the rank. The update is scaled by alpha divided by rank, so this keeps the ratio at 1 or 2. Note that PEFT's own default is rank 8 with alpha 8, so set both explicitly.
- Target modules. Apply adapters to all the main linear layers: the attention projections q, k, v and o, plus the MLP's gate, up and down projections.
- Learning rate and epochs. Unsloth recommends starting at 2e-4 for LoRA and QLoRA, and 1 to 3 epochs for most instruction datasets.
- rsLoRA. Rank-stabilised LoRA scales by alpha divided by the square root of the rank, which can help at high ranks.
Pros and cons
LoRA
- Pros: strong quality, close to full fine-tuning on many tasks; small adapter files; faster steps than QLoRA; no inference overhead after merging.
- Cons: needs enough VRAM for the 16-bit base model.
QLoRA
- Pros: the lowest memory; makes 7B to 14B fine-tunes practical on consumer GPUs and free notebooks; widely supported.
- Cons: slower steps; possible small quality loss; not recommended for some architectures.
DoRA
- Pros: reported quality gains over LoRA, especially at low ranks; drop-in flag in PEFT and mlx-lm.
- Cons: more training overhead; fewer layer types supported.
Full fine-tuning
- Pros: maximum flexibility; best for large datasets and deep behaviour changes.
- Cons: very high memory and compute; higher risk of forgetting what the model already knew; large checkpoints.
Which should you choose?
- Up to 24 GB of VRAM, or a free notebook: QLoRA, unless your model's guide says otherwise.
- A 40 to 80 GB GPU and a model that fits in 16-bit: LoRA.
- LoRA results are close but not quite there: try DoRA or a higher rank before jumping to full fine-tuning.
- A large, high-quality dataset, multiple big GPUs, and a goal like teaching a new language: consider full fine-tuning, with tools such as TRL or Axolotl and DeepSpeed or FSDP.
- A Mac: mlx-lm runs LoRA, QLoRA and DoRA on Apple Silicon; see our MLX vs GGUF guide.
To check whether a base model will even fit for inference afterwards, use our VRAM guide. Libraries to start with include Unsloth, PEFT with Transformers, and MLX on a Mac.
For the full workflow from dataset to GGUF, follow our guide to fine-tune an LLM locally. LoRA adapters also work with reinforcement learning; see GRPO vs PPO.
Who this is for
Developers and researchers planning their first or next fine-tune who need to match a method to their hardware, budget and quality bar.
FAQ
Is QLoRA worse than LoRA?
Slightly, in principle, because the frozen base is stored in 4-bit. The original QLoRA paper reported matching 16-bit fine-tuning task performance, and tools like Unsloth say their dynamic 4-bit quants largely close the remaining gap. If LoRA fits in memory, it is the safer default.
How much VRAM do I need for QLoRA?
Unsloth lists minimums of about 6 GB for an 8B model, 8.5 GB for 14B, 26 GB for 32B and 41 GB for 70B with its optimised stack. Longer sequences and larger batches need more, and other tools typically use more.
What is the difference between LoRA and full fine-tuning?
Full fine-tuning updates every weight and needs memory for gradients and optimizer states, about 16 bytes per parameter with Adam in mixed precision. LoRA freezes the model and trains small adapters, so it needs a fraction of the memory and produces small files.
What rank and alpha should I use for LoRA?
A common starting point is rank 16 or 32 with alpha equal to the rank or twice the rank, applied to all attention and MLP projection layers. Increase the rank only if the model underfits.
Is DoRA better than LoRA?
Its authors report consistent gains over LoRA on the tasks they tested, particularly at low ranks. It costs more training overhead, so it is worth trying when LoRA results fall short.
Can I merge a QLoRA adapter into the model?
Yes. Most tools merge the adapter into a 16-bit copy of the base model for export, and Unsloth and Axolotl can then convert the result to GGUF for local use.