To fine-tune an LLM locally, pick a small instruct model, put a few hundred to a few thousand high-quality examples into a chat-format JSONL file, and train a LoRA or QLoRA adapter with a tool like Unsloth on an NVIDIA GPU, or mlx-lm on an Apple Silicon Mac. Then test it on held-out examples and export it to GGUF so it runs in Ollama, LM Studio or llama.cpp. A 4B to 8B model with QLoRA fits on a single consumer GPU.
This guide walks through each step, explains which tool to use in late 2026, and flags the mistakes that waste the most time.
Step 0: decide whether you should fine-tune at all
Fine-tuning changes how a model behaves. It is good at teaching a consistent style, format, tone or narrow skill, such as always answering in your company's JSON schema, classifying support tickets, or writing in a house voice. It is a poor way to add frequently changing facts; retrieval, meaning feeding relevant documents into the prompt, handles that better.
Try a well-written prompt with a few examples first. If that gets you most of the way but results are inconsistent, too long, or too expensive in tokens, fine-tuning is worth it. If your real goal is to understand how models learn, our guide to training a small LLM from scratch is the better project.
Step 1: choose your hardware path
| Your machine | Recommended tool | Typical method |
|---|---|---|
| NVIDIA GPU on Linux, WSL or Windows | Unsloth, Axolotl or TRL | QLoRA or LoRA |
| AMD or Intel GPU | Unsloth or Axolotl, following their guides | LoRA or QLoRA |
| Apple Silicon Mac | mlx-lm, or Unsloth Studio | LoRA or QLoRA |
| No suitable GPU | A free Colab or Kaggle notebook | QLoRA |
For memory, Unsloth's requirements page lists minimums of about 6 GB of VRAM to QLoRA-train an 8B model and 22 GB for 16-bit LoRA. Our LoRA vs QLoRA guide explains the trade-off and lists more sizes.
Step 2: choose a tool
Unsloth
Unsloth is the most popular route for single-GPU fine-tuning. It is Apache 2.0 licensed and now comes in three forms, according to its GitHub README: a desktop app, a web UI called Unsloth Studio, and Unsloth Core, the original Python library. It supports LoRA, QLoRA, full fine-tuning, and preference and reinforcement learning methods like DPO and GRPO (see our GRPO vs PPO explainer), and it can export straight to GGUF. Its README claims about 2 times faster training with 70 percent less VRAM than standard setups; treat that as the vendor's own figure. Free Colab notebooks cover models such as Gemma 4, Qwen3.5, gpt-oss and Llama.
Axolotl
Axolotl is a free, Apache 2.0 fine-tuning framework driven by YAML config files. It shines when you want reproducible configs, multi-GPU training, or newer techniques: its README lists FSDP2, DeepSpeed, context and expert parallelism, NVFP4 LoRA, and, since September 2026, GGUF export through axolotl export. It needs an NVIDIA GPU, Ampere or newer for bf16 and Flash Attention, or an AMD GPU.
TRL
TRL, Hugging Face's post-training library, provides SFTTrainer, DPOTrainer, GRPOTrainer and more on top of Transformers, with PEFT integration for LoRA and QLoRA. It is the most flexible option if you are comfortable writing Python, and Unsloth itself plugs into TRL's trainers. There is also a CLI, for example trl sft --model_name_or_path Qwen/Qwen2.5-0.5B --dataset_name trl-lib/Capybara --output_dir out.
mlx-lm on a Mac
Apple's MLX ecosystem includes mlx_lm.lora, which supports LoRA, DoRA and full fine-tuning, and switches to QLoRA automatically for quantized models. See our MLX vs GGUF guide for the Mac picture.
What about torchtune?
Many older tutorials recommend torchtune. Its repository now states that it is no longer actively maintained and that development wound down in 2025, so choose one of the tools above for new projects.
Step 3: pick a base model
- Start with an instruct model, not a base model. Unsloth's guide notes that instruct models can be fine-tuned directly with chat templates and need less data.
- Start small. A 3B to 8B model trains quickly, so you can iterate on your data. Scale up once the pipeline works.
- Check the licence. Apache 2.0 and MIT models are simplest for commercial use. Our explainer on open weight vs open source has a licence table.
- Follow the model's own guide. Some families have specific advice. Unsloth's Qwen3.5 guide, for example, recommends 16-bit LoRA rather than QLoRA.
Qwen, Gemma and Llama models all have strong small instruct versions with wide tool support.
Step 4: prepare your data
Data quality decides the result more than any hyperparameter.
- Use the chat format. Store one example per line in JSONL, each with a list of messages, for example a system message, a user message and the ideal assistant reply. TRL, Unsloth and Axolotl all accept conversational data in this style and apply the model's chat template for you.
- Write the outputs you actually want. Every assistant message is a lesson. Fix inconsistent, wrong or rambling answers before training.
- Aim for coverage, not volume. A few hundred clean, varied examples often beat thousands of near-duplicates.
- Hold out a test set. Keep a slice of examples the model never trains on, so you can tell learning from memorising.
- Keep sensitive data in mind. Local training keeps data on your machine, but anything you later publish, including adapters, can leak training content.
You can load local files or browse public examples with Hugging Face Datasets.
Step 5: configure and train
Here is a minimal Unsloth Core script, adapted from Unsloth's Qwen3.5 fine-tuning guide. It trains a 16-bit LoRA on a 4B model.
from unsloth import FastLanguageModel
from datasets import load_dataset
from trl import SFTTrainer, SFTConfig
max_seq_length = 2048
dataset = load_dataset("json", data_files={"train": "train.jsonl"}, split="train")
model, tokenizer = FastLanguageModel.from_pretrained(
model_name="Qwen/Qwen3.5-4B",
max_seq_length=max_seq_length,
load_in_4bit=False,
load_in_16bit=True,
)
model = FastLanguageModel.get_peft_model(
model, r=16, lora_alpha=16, lora_dropout=0, bias="none",
target_modules=["q_proj", "k_proj", "v_proj", "o_proj",
"gate_proj", "up_proj", "down_proj"],
use_gradient_checkpointing="unsloth", random_state=3407,
)
trainer = SFTTrainer(
model=model, train_dataset=dataset, tokenizer=tokenizer,
args=SFTConfig(max_seq_length=max_seq_length,
per_device_train_batch_size=1, gradient_accumulation_steps=4,
warmup_steps=10, max_steps=100, learning_rate=2e-4,
logging_steps=1, optim="adamw_8bit", output_dir="outputs"),
)
trainer.train()
Unsloth's guide lists about 10 GB of VRAM for 16-bit LoRA on Qwen3.5-4B. Sensible starting settings from its hyperparameter guide are rank 16 or 32, alpha equal to the rank or double it, a learning rate of 2e-4, and 1 to 3 epochs.
With Axolotl, the equivalent is a YAML file and one command, for example axolotl train examples/llama-3/lora-1b.yml from its fetched examples.
Step 6: watch the loss and evaluate
Training loss should fall and then flatten. Unsloth's guide says a loss around 0.5 to 1.0 is often a good sign, depending on the task, and that a loss near zero can mean overfitting. Turn on evaluation against your held-out set so you see validation loss too; if training loss keeps falling while validation loss rises, stop earlier or add data.
Then judge outputs directly. Run the same 20 to 50 held-out prompts through the base model and your fine-tune, and compare them side by side. For a structured task, score them automatically, for example by checking whether the JSON parses.
Step 7: export and run it locally
You have three common outputs:
- The LoRA adapter on its own, a small file you can load on top of the base model.
- A merged 16-bit model, for serving with vLLM or Transformers.
- A GGUF file, for Ollama, LM Studio and llama.cpp.
In Unsloth, saving to GGUF is a single call, such as model.save_pretrained_gguf("my-model", tokenizer, quantization_method="q4_k_m"). To run the result in Ollama, write a Modelfile containing FROM ./my-model.Q4_K_M.gguf, then run ollama create -f Modelfile my-model followed by ollama run my-model, as described in Ollama's GGUF announcement. Our GGUF quantization guide helps you pick the quant type.
Common mistakes
- Wrong chat template. Training with one template and chatting with another produces odd output. Let the tool apply the model's own template.
- Too many epochs. More than three passes over a small dataset often leads to memorisation.
- Judging by training loss alone. Always test on held-out prompts.
- Starting too big. Debug your pipeline on a small model first.
Pros and cons of fine-tuning locally
- Pros: data never leaves your machine; no per-token training fees; full control over the model; the result runs offline.
- Cons: you need suitable hardware or a free notebook; you own the debugging; large models still need cloud GPUs.
Who this is for
Developers who need a model to follow a specific format or style reliably, teams with private data that cannot leave their network, and hobbyists learning how post-training works.
FAQ
Can I fine-tune an LLM on my own computer?
Yes. With QLoRA, Unsloth lists about 6 GB of VRAM as the minimum for an 8B model, so many consumer NVIDIA GPUs can do it, and Apple Silicon Macs can fine-tune with mlx-lm.
How much data do I need to fine-tune an LLM?
There is no fixed number. A few hundred clean, varied examples are often enough to teach a format or style, while new skills need more. Quality and consistency matter more than volume.
Unsloth or Axolotl: which should I use?
Unsloth is the quickest path on a single GPU and offers a desktop app and notebooks. Axolotl suits reproducible YAML configs, multi-GPU setups and newer research techniques. Both are free and open source.
Is fine-tuning better than RAG?
They solve different problems. Fine-tuning changes behaviour, such as style, format and narrow skills; retrieval supplies up-to-date facts. Many production systems use both.
Can I run my fine-tuned model in Ollama?
Yes. Export it to GGUF, create a Modelfile that points to the file with a FROM line, then run ollama create and ollama run.
Is torchtune still a good choice?
Its repository says it is no longer actively maintained, with development wound down in 2025. Use Unsloth, Axolotl, TRL or mlx-lm for new projects.