How much VRAM you need for an LLM comes down to two numbers: the size of the model weights at your chosen precision, plus the KV cache for your context window, plus a little overhead. A quick estimate: at 4-bit quantization, budget about 0.6 GB per billion parameters for weights; at 8-bit about 1.1 GB; at 16-bit about 2 GB. Then add the KV cache, which for a typical 8B model is about 4 GiB for every 32,000 tokens of context at 16-bit precision.

So an 8B model at 4-bit with a 32k context needs roughly 9 to 10 GB, while the same model at 16-bit needs around 20 GB or more. The rest of this guide shows where those numbers come from, so you can size any model yourself.

The formula

Total memory is approximately:

weights + KV cache + runtime overhead

  • Weights = number of parameters × bits per weight ÷ 8. Bits per weight depends on the quantization.
  • KV cache = 2 × layers × KV heads × head dimension × bytes per value × context tokens. The "2" is for keys and values.
  • Overhead covers activations, compute buffers and the runtime itself. It varies by tool, so keep at least a gigabyte or two spare.

Step 1: weights

Bits per weight are not exactly what the name says. The llama.cpp quantize README measures Q4_K_M at about 4.89 bits per weight, Q8_0 at 8.5, and F16 at 16. That gives these rough multipliers:

PrecisionBits per weightGB per billion parameters
16-bit (F16 or BF16)162.0
8-bit (Q8_0)8.5about 1.06
6-bit (Q6_K)6.56about 0.82
4-bit (Q4_K_M)4.89about 0.61
3-bit (Q3_K_M)4.0about 0.50

The README's own measurement agrees: Llama 3.1 8B at Q4_K_M is 4.58 GiB. If quant names are new to you, read our guide to GGUF quantization.

Step 2: the KV cache

Every token in your context stores a key and a value vector in every layer. You can read the numbers you need from a model's config.json on Hugging Face: num_hidden_layers, num_key_value_heads and head_dim.

Worked example with Llama 3.1 8B, which has 32 layers, 8 key-value heads and a head dimension of 128:

  • 2 × 32 × 8 × 128 × 2 bytes = 131,072 bytes, which is 128 KiB per token at 16-bit.
  • 32,768 tokens × 128 KiB = 4 GiB.
  • At its full 131,072-token context, that becomes 16 GiB, more than three times the 4-bit weights.

A bigger model costs more per token. Qwen3-32B has 64 layers with the same head layout, so it needs 256 KiB per token, or 8 GiB for 32k tokens. Qwen3-8B, with 36 layers, needs 144 KiB per token.

Newer architectures cut this sharply. gpt-oss-20b alternates full-attention layers with sliding-window layers that only look back 128 tokens, and uses a head dimension of 64. Only its 12 full-attention layers grow with context, at about 24 KiB per token, so a full 131,072-token context costs roughly 3 GiB. Models with sliding-window or linear-attention layers, such as recent Gemma and Qwen releases, have similar savings, so always check the config rather than assuming.

Step 3: shrink the cache if you need to

  • Quantize the KV cache. In llama.cpp, --cache-type-k q8_0 --cache-type-v q8_0 roughly halves cache memory compared with the 16-bit default. MLX's server offers --kv-bits. Our KV cache quantization guide covers every runtime and the quality trade-offs.
  • Use a shorter context. Do not allocate 128k tokens if your prompts are 8k.
  • Pick a model with an efficient attention design if you need very long contexts.

What fits in 8, 16, 24 and 32 GB?

Here are real download sizes from the Ollama model library pages for Gemma 4 and gpt-oss, checked on October 5, 2026, with memory guidance from the gpt-oss model card. These are file sizes; add your context and overhead on top.

Model and quantDownload sizeComfortable on
Gemma 4 E4B, Q4_K_M6.6 GB8 to 12 GB of VRAM, with short contexts on 8 GB
Gemma 4 12B, Q4_K_M8.0 GB12 to 16 GB
gpt-oss-20b, native MXFP414 GB16 GB (OpenAI says it runs within 16 GB of memory)
Gemma 4 26B A4B, Q4_K_M18 GB24 GB
Gemma 4 31B, Q4_K_M20 GB24 GB with a moderate context, 32 GB with room to spare
gpt-oss-120b, native MXFP465 GBa single 80 GB GPU, per OpenAI, or a Mac with a lot of unified memory

As a rough guide by tier:

  • 8 GB: small models of a few billion parameters at 4-bit, with modest context.
  • 12 to 16 GB: models up to about 12 to 14 billion parameters at 4-bit, or gpt-oss-20b.
  • 24 GB, for example an RTX 3090 or RTX 4090: models around 27 to 32 billion parameters at 4-bit with a reasonable context.
  • 32 GB, for example an RTX 5090: the same models with longer contexts or higher-precision quants.
  • 48 GB and up: 70B-class dense models at 4-bit, or large mixture-of-experts models with offloading.

Mixture-of-experts models: total vs active parameters

Mixture-of-experts (MoE) models only use a fraction of their weights for each token, but you still need memory for all of them. gpt-oss-120b has 117 billion parameters with 5.1 billion active per token, and still needs a single 80 GB GPU according to its model card. Gemma 4 26B A4B has about 4 billion active parameters but downloads at 18 GB in Q4_K_M.

The upside is speed: an MoE model generates tokens much faster than a dense model of the same total size, closer to a model the size of its active parameters, as long as everything stays in fast memory. That also makes MoE models the best candidates for partial CPU offload, because only a small part of the weights is touched for each token.

When the model does not fit

You have three options, from best to worst:

  1. Use a smaller quant or a smaller model. Usually the best trade.
  2. Offload some layers to system RAM. llama.cpp supports CPU plus GPU hybrid inference, and Ollama and LM Studio do this automatically; our Ollama vs llama.cpp vs LM Studio comparison explains the differences. It works, but every offloaded layer slows generation, sometimes dramatically. In Ollama, ollama ps shows the split under PROCESSOR.
  3. Run entirely on CPU. Possible for small models, slow for large ones.

Watch the context default too. Ollama sets its default context from your VRAM: 4k tokens under 24 GiB, 32k from 24 to 48 GiB, and 256k above that, according to its context-length docs. Raising it costs memory exactly as calculated above.

Apple Silicon and unified memory

Macs share one pool of memory between CPU and GPU, so a 64 GB Mac can hold models no consumer graphics card can. macOS limits how much of that memory the GPU can wire by default. The mlx-lm README explains that if a model fits in RAM but runs slowly, you can raise the limit with sudo sysctl iogpu.wired_limit_mb=N, where N is larger than the model size in megabytes but smaller than total memory. Leave headroom for macOS and your apps. Mac users should also look at MLX, Apple's framework, which Ollama and LM Studio can use; our MLX vs GGUF guide compares the two formats.

Pros and cons of each way to get more memory

  • Bigger GPU: fastest option and best for long contexts, but the most expensive per gigabyte.
  • Two GPUs: llama.cpp and vLLM can split a model across cards (see our vLLM vs llama.cpp comparison), but you lose some speed to communication and need a board and power supply that can handle it.
  • Mac with lots of unified memory: very large capacity in a quiet box, but prompt processing is typically slower than on high-end NVIDIA cards.
  • CPU offload: free if you already have RAM, but much slower generation.

Who this is for

Anyone choosing a model for existing hardware, or choosing hardware for a model: hobbyists, developers sizing a local coding assistant, and teams planning a small on-premises server. If you are fine-tuning rather than running models, memory needs are higher; tools like Unsloth publish their own tables for that, and our LoRA vs QLoRA guide explains how to cut them.

FAQ

How much VRAM do I need to run a 7B or 8B model?

At 4-bit quantization the weights take about 4.5 to 5 GB, so 8 GB of VRAM works with a short context and 12 GB is comfortable. At 16-bit you need about 16 GB for the weights alone, plus the KV cache.

Is 16 GB of VRAM enough for local LLMs?

Yes for most models up to about 14 billion parameters at 4-bit, and for gpt-oss-20b, which OpenAI says runs within 16 GB of memory. Larger models need offloading or more memory.

How do I calculate KV cache size?

Multiply 2 × layers × key-value heads × head dimension × bytes per value × number of tokens. For Llama 3.1 8B at 16-bit that is 128 KiB per token, or 4 GiB for 32,768 tokens.

Does context length affect VRAM?

Yes, directly. The KV cache grows linearly with context, and for long contexts it can exceed the size of the weights. Quantizing the cache or choosing a model with sliding-window attention reduces it.

Can I use system RAM instead of VRAM?

Yes. llama.cpp, Ollama and LM Studio can offload layers to system RAM, but generation slows down for every layer that leaves the GPU. Mixture-of-experts models tolerate offloading better than dense models.

Do MoE models need less VRAM?

Not for storage. You need memory for all experts, so a 117B-parameter MoE still needs room for 117B parameters. MoE models are faster per token, not smaller.