KV cache quantization stores the attention cache, the keys and values a model keeps for every token in its context, at lower precision than the default 16 bits. Switching the cache to 8-bit (q8_0 in llama.cpp, Ollama and LM Studio, or fp8 in vLLM and SGLang) roughly halves its memory with little quality loss for most models. Going down to 4-bit (q4_0) cuts it to a little over a quarter, but the quality hit is more noticeable, especially in long conversations. It does not shrink the model weights at all; it only frees memory for longer context or more simultaneous users.
Why the KV cache matters
Every token in your context stores a key vector and a value vector in every attention layer, so the cache grows linearly with context length. Our guide to how much VRAM you need for an LLM walks through the formula:
2 × layers × KV heads × head dimension × bytes per value × tokens
For Llama 3.1 8B, with 32 layers, 8 key-value heads and a head dimension of 128, that is 128 KiB per token at 16-bit precision. A 32,768-token context costs 4 GiB, and the full 131,072-token context costs 16 GiB, which is more than the model's own weights at 8-bit. That is the memory KV cache quantization goes after.
How much memory each cache type saves
llama.cpp's q8_0 and q4_0 types store values in blocks of 32 with one 16-bit scale per block, according to the block definitions in ggml's source. That works out to 8.5 and 4.5 bits per value. Applied to the Llama 3.1 8B example:
| Cache type | Bits per value | 32k context | 128k context | Where you use it |
|---|---|---|---|---|
| f16 (default) | 16 | 4 GiB | 16 GiB | Everywhere |
| q8_0 | 8.5 | about 2.1 GiB | about 8.5 GiB | llama.cpp, Ollama, LM Studio |
| fp8 | 8 | 2 GiB | 8 GiB | vLLM, SGLang |
| q4_0 | 4.5 | about 1.1 GiB | about 4.5 GiB | llama.cpp, Ollama, LM Studio |
These numbers are for the cache alone; the weights come on top. Models with sliding-window or linear-attention layers, such as recent Gemma and Qwen releases, already have much smaller caches, so quantizing them saves fewer gigabytes.
How to turn it on, runtime by runtime
llama.cpp
llama.cpp sets the key and value caches separately. The server documentation lists -ctk / --cache-type-k and -ctv / --cache-type-v, with allowed values f32, f16, bf16, q8_0, q4_0, q4_1, iq4_nl, q5_0 and q5_1, and f16 as the default.
llama-server -m model.gguf -c 65536 -ctk q8_0 -ctv q8_0
Two details matter. First, a quantized value cache needs Flash Attention. The default --flash-attn auto turns it on for you when it can, and llama.cpp refuses to start if Flash Attention is forced off. Second, on NVIDIA GPUs, keep the key and value types the same. According to the build documentation, standard CUDA builds only compile Flash Attention kernels for matching pairs such as q8_0 with q8_0 and q4_0 with q4_0. Mixed pairs fall back to a slower path unless you build with GGML_CUDA_FA_QUANTS=all.
Ollama
Ollama has a single global setting, documented in its FAQ. Set the environment variable before starting the server:
OLLAMA_KV_CACHE_TYPE=q8_0 ollama serve
The options are f16, the default, q8_0 and q4_0. Ollama describes q8_0 as using about half the memory of f16 with a very small loss in precision that usually has no noticeable impact, and q4_0 as about a quarter of the memory with a small to medium loss that may be more noticeable at higher context sizes. Quantization only works when Flash Attention is active; Ollama enables it automatically on supported hardware, and you can force it with OLLAMA_FLASH_ATTENTION=1. Because the setting is global, every model you run uses it. Our Ollama vs llama.cpp vs LM Studio comparison covers the other differences between these tools.
LM Studio
LM Studio exposes K cache and V cache quantization among a model's load settings, and the same options exist in its SDK as llamaKCacheQuantizationType and llamaVCacheQuantizationType. Its API reference notes that value-cache quantization requires Flash Attention. Because LM Studio runs GGUF models on llama.cpp, the matching-pair advice above applies on NVIDIA cards too.
vLLM and SGLang
The production servers use 8-bit floating point instead of llama.cpp's block formats. In vLLM, pass --kv-cache-dtype fp8; variants are fp8_e4m3 and fp8_e5m2. The vLLM documentation warns that without calibration every scale defaults to 1.0, and recommends calibrating scales with llm-compressor for the best accuracy. It also lets you leave sensitive layers, such as sliding-window layers, at full precision with --kv-cache-dtype-skip-layers. SGLang uses the same --kv-cache-dtype flag with fp8 options, plus FP4 options that need CUDA 12.8 or newer. For when to choose these servers at all, see vLLM vs llama.cpp.
MLX on a Mac
MLX's mlx-lm package has a --kv-bits option for generation and its server, with a default group size of 64 and a --quantized-kv-start setting that keeps the first part of the cache at full precision. If you are choosing between formats on Apple Silicon, read our MLX vs GGUF guide.
Does KV cache quantization hurt quality?
It can, and the effect depends on the model. A few points are well supported:
- 8-bit is the safe default. Ollama calls q8_0 the recommended choice if you are not using f16, and halving the cache rarely changes answers noticeably.
- 4-bit is a trade. Errors in the cache add up over long contexts, so q4_0 is most noticeable exactly where you need the memory most.
- Keys and values behave differently. The KIVI paper found that keys have outlier channels and are best quantized per channel, while values suit per-token quantization. llama.cpp's simple block formats treat both the same, which is one reason people often keep keys at higher precision. On NVIDIA, though, remember the matching-pair limitation.
- Model layout matters. Ollama's FAQ notes that models with a high grouped-query attention ratio may lose more precision from cache quantization; it gives Qwen2 as an example.
The honest approach is to test your own workload: run the same long prompt at f16 and q8_0, compare, then try q4_0 only if you still need memory.
Pros and cons
- Pros: frees gigabytes for longer context or more parallel users; costs nothing to try; no re-download needed; 8-bit is close to lossless for most models.
- Cons: 4-bit can degrade long-context answers; requires Flash Attention for the value cache in llama.cpp-based tools; mixed key and value types can be slow on NVIDIA; does nothing for the weights.
When to use it
- You hit out-of-memory errors only at long context: switch both caches to q8_0 first.
- You need even more room: compare a smaller weight quant against a q4_0 cache. Our GGUF quantization guide explains the weight side.
- You serve many users with vLLM or SGLang: use fp8 with calibrated scales.
- Your model already uses sliding-window or linear attention: check the actual cache size first; the gain may be small.
Who this is for
Anyone running open-weight models locally who wants longer context on the same GPU or Mac, and engineers squeezing more concurrent requests out of a serving box.
FAQ
What is KV cache quantization?
It stores the attention keys and values that a model caches for each token at lower precision, usually 8-bit or 4-bit instead of 16-bit, to cut memory use during long conversations.
Should I use q8_0 or q4_0 for the KV cache?
Start with q8_0. It roughly halves cache memory with very little quality loss for most models. Use q4_0 only if you still need memory, and test long prompts, because errors are more noticeable at higher context sizes.
How do I set the KV cache type in Ollama?
Set the environment variable OLLAMA_KV_CACHE_TYPE to q8_0 or q4_0 before starting the server. It applies to all models and requires Flash Attention, which Ollama enables automatically on supported hardware.
Why does llama.cpp say a quantized V cache requires Flash Attention?
Quantized value caches are only supported through the Flash Attention code path. Leave --flash-attn on its default of auto, or set it to on, and llama.cpp can use a quantized value cache.
Does KV cache quantization make inference faster?
Not directly. Its main benefit is memory; for raw speed, look at speculative decoding. It can help speed indirectly by letting a model stay entirely on the GPU instead of spilling to system RAM, but mixed key and value types on NVIDIA can be much slower.
Is FP8 in vLLM the same as q8_0 in llama.cpp?
No. Both use about 8 bits per value, but FP8 is a floating-point format with per-tensor or per-head scales, while q8_0 is an integer format with a scale for every block of 32 values.