GGUF quantization is how most people shrink an open-weight model so it fits on a laptop or a single GPU. GGUF itself is a file format used by llama.cpp and the tools built on it; the quantization type in the filename, such as Q4_K_M or Q8_0, tells you how many bits each weight was compressed to. For most people, Q4_K_M is the sensible default, Q5_K_M or Q6_K are worth it if you have spare memory, and Q8_0 is close to the original quality at roughly half the size of 16-bit weights.
This guide decodes the names, shows real size numbers from the official llama.cpp documentation, and gives you a simple rule for picking a file.
What GGUF is (and is not)
GGUF is a binary file format from the ggml project. According to the GGUF specification, it is designed for single-file deployment, fast loading through memory mapping, and "full information": the architecture, tokenizer and other metadata live inside the file as key-value pairs, so you do not need extra config files. It replaced the older GGML, GGMF and GGJT formats.
GGUF is not a quantization method. A GGUF file can hold full 16-bit weights, 8-bit weights, or anything down to roughly 2 bits per weight. The quantization type is a separate choice, and it is usually written at the end of the filename.
GGUF files run in llama.cpp and everything built on it, including Ollama and LM Studio. If you have not chosen a runner yet, our comparison of Ollama vs llama.cpp vs LM Studio will help.
GGUF vs safetensors
Safetensors is the format you usually see in a model's original repository on the Hugging Face Hub. It stores tensors safely and is what Transformers, vLLM and most training tools load; our vLLM vs llama.cpp guide covers that serving side. It typically holds full-precision weights and relies on separate JSON files for the config and tokenizer.
| GGUF | Safetensors | |
|---|---|---|
| Main users | llama.cpp, Ollama, LM Studio | Transformers, vLLM, SGLang, training libraries |
| Typical contents | Quantized weights plus all metadata in one file | Weights only, with config and tokenizer files alongside |
| Typical use | Local inference on CPUs, Macs and consumer GPUs | Training, fine-tuning and GPU server inference |
| Quantization | Built-in types from about 1.5 to 8 bits | Usually 16-bit; quantized variants use other schemes |
Practical rule: train and fine-tune from safetensors, then convert to GGUF for local use.
How to read a GGUF quant name
Take Q4_K_M apart piece by piece.
- Q4 means about 4 bits per weight on average.
- K means a "k-quant", the block-based scheme introduced in llama.cpp in 2023 that mixes precisions inside the model.
- M means medium. K-quants come in S (small), M (medium) and L (large) variants that keep more or fewer of the sensitive tensors at higher precision. Larger means slightly bigger files and slightly better quality.
Other names you will see:
- Q8_0, Q4_0, Q5_1 are older, simpler "legacy" quants. Q8_0 is still popular because it is nearly lossless.
- IQ2_M, IQ3_XXS, IQ4_XS are "i-quants", designed for very low bit rates. They usually rely on an importance matrix.
- imatrix is an importance matrix: statistics gathered by running sample text through the model, used to decide which weights deserve more precision. The quantize README says a suitable imatrix file can reduce the accuracy loss from quantization.
- F16 and BF16 are unquantized 16-bit weights.
Some uploaders add their own labels, such as Unsloth's "UD" (dynamic) quants. Those are variations on the same idea; read the model card for details.
Real sizes: bits per weight by quant type
The llama.cpp quantize README publishes measurements for Llama 3.1 8B. Here is a subset, rounded.
| Quant type | Bits per weight | File size (GiB) |
|---|---|---|
| F16 | 16.0 | 14.96 |
| Q8_0 | 8.5 | 7.95 |
| Q6_K | 6.56 | 6.14 |
| Q5_K_M | 5.70 | 5.33 |
| Q4_K_M | 4.89 | 4.58 |
| Q4_K_S | 4.67 | 4.36 |
| IQ4_XS | 4.46 | 4.17 |
| Q3_K_M | 4.00 | 3.74 |
| IQ3_M | 3.76 | 3.52 |
| Q2_K | 3.16 | 2.95 |
| IQ2_M | 2.93 | 2.74 |
In words: Q8_0 is a little over half the size of 16-bit weights, Q4_K_M is a little under a third, and the 2-bit types are under a fifth. Notice that "Q4" is really closer to 4.9 bits, because some tensors are kept at higher precision.
The same README gives whole-model examples. Llama 3.1 8B goes from 32.1 GB in its original form to 4.9 GB at Q4_K_M, and Llama 3.1 70B goes from 280.9 GB to 43.1 GB. Those originals are 32-bit, which is why the reduction looks so dramatic.
Quality vs size: how to choose
Quantization trades accuracy for memory. The llama.cpp docs measure the loss with perplexity and KL divergence against the original model; lower bits mean more drift. Generation speed also usually rises as files shrink, because producing each token is limited mainly by how fast weights stream from memory.
A practical decision guide:
- Q8_0 if it fits comfortably. Quality is very close to 16-bit.
- Q6_K or Q5_K_M when you have some headroom. A good middle ground.
- Q4_K_M as the default for most laptops and single GPUs. This is the type the llama.cpp docs use in their own example.
- IQ4_XS or Q3_K_M when you are just short of memory.
- IQ2 and Q2 types only to squeeze a very large model onto limited hardware, and expect visible quality loss, especially on smaller models.
A useful rule of thumb from the community, worth testing for yourself: a bigger model at 4 bits often beats a smaller model at 8 bits for the same memory. Always leave room for the context window, which needs its own memory. Our VRAM guide shows how to add it up.
How to get a GGUF
Download a ready-made quant
Most popular models have GGUF repos on Hugging Face, from the model makers, from the ggml-org account, or from community quantizers. Each repo usually contains one file per quant type. You can run them directly:
# llama.cpp: download and run a specific quant
llama cli -hf ggml-org/gpt-oss-20b-GGUF
# Ollama: run any GGUF repo from Hugging Face, optionally picking a quant
ollama run hf.co/{username}/{repository}:Q4_K_M
The Ollama syntax comes from the Hugging Face Hub docs. LM Studio shows the available quant types in its download panel.
Make your own
The official route is two steps: convert the original model to a 16-bit GGUF, then quantize it.
python convert_hf_to_gguf.py --outfile model-bf16.gguf --outtype bf16 --remote google/gemma-4-E2B-it
./build/bin/llama-quantize model-bf16.gguf model-Q4_K_M.gguf Q4_K_M
Quantize from 16-bit or 32-bit weights, not from another quant; the README warns that re-quantizing can severely reduce quality. Multimodal models also need a separate "mmproj" file for the vision or audio encoder, which is usually kept at Q8_0 or 16-bit. If you prefer not to install anything, the GGUF-my-repo Space builds quants in the browser. Fine-tuning tools such as Unsloth can also export straight to GGUF.
Pros and cons of GGUF
- Pros: one self-contained file; runs on CPUs, Apple Silicon and most GPUs; wide choice of sizes; memory-mapped loading; supported by the most popular local tools.
- Cons: mainly an inference format, not for training; quality drops at very low bit rates; new architectures need llama.cpp support before GGUFs work; on Apple Silicon, MLX-format models are a strong alternative, compared in our MLX vs GGUF guide.
Who this is for
Anyone running models locally: hobbyists choosing a download, developers sizing a deployment, and researchers comparing quantization effects. If you plan to fine-tune, start from the original safetensors weights of a model such as Llama, and only convert to GGUF at the end.
FAQ
What does GGUF stand for?
The specification does not spell it out as an acronym; it describes GGUF as a file format for storing models for inference with GGML and GGML-based tools. It is the successor to the older GGML, GGMF and GGJT formats.
Is Q4_K_M or Q5_K_M better?
Q5_K_M keeps more precision, so quality is a little higher, but the file is about 16 percent larger for an 8B model (5.33 versus 4.58 GiB in the llama.cpp table). Choose Q5_K_M if it fits with room for your context window; otherwise Q4_K_M is the standard choice.
What is the difference between GGUF and safetensors?
Safetensors stores model tensors, usually at full precision, and is the standard for training and GPU serving with Transformers or vLLM. GGUF bundles quantized weights and all metadata in a single file for llama.cpp-based local inference.
Can I fine-tune a GGUF model?
GGUF is built for inference. Fine-tune from the original safetensors weights instead, for example with LoRA or QLoRA, then convert the merged result to GGUF.
What is an imatrix quant?
It is a quant made with an importance matrix, which records which weights matter most on sample text. The quantizer then spends more precision on those weights, which reduces quality loss, especially at 2 to 4 bits.
Does quantization make a model faster?
Usually yes for generation, because smaller weights move through memory faster. The exact effect depends on your hardware and backend, so compare speeds on your own machine.