To run Gemma 4 locally, install Ollama and type ollama run gemma4 for the default E4B build, or pick a size tag such as gemma4:12b, gemma4:26b or gemma4:31b. The family is Apache 2.0 on the published Hugging Face cards, multimodal (text and image on every size; audio on E2B, E4B and 12B), and also runs in llama.cpp, LM Studio and vLLM. Memory stretches from roughly 4–6 GB for E2B at 4-bit up to about 17–20 GB for 31B at 4-bit.

This guide covers which Gemma 4 size fits your machine, the setup commands that actually work, thinking-mode controls, vision caveats, and when to prefer the MoE 26B-A4B over the dense 31B.

What Gemma 4 is

Gemma 4 is Google DeepMind's open-weight multimodal Gemma family. According to the google/gemma-4-26B-A4B-it model card, the release includes five instruction-tuned sizes: E2B, E4B, 12B Unified, 26B A4B (mixture-of-experts) and 31B dense. Context is 128K tokens on the edge models and 256K on 12B, 26B-A4B and 31B. Weights are published under Apache 2.0, which is a clearer commercial grant than older Gemma Terms of Use builds — see open weight vs open source for the difference.

Architecturally, Gemma 4 interleaves local sliding-window attention with global attention, and the MoE variant activates only about 3.8B of its 25.2B parameters per token. That is why the 26B-A4B often feels closer to a mid-size model in speed while retaining a larger quality ceiling than E4B or 12B.

VariantStyleActive / totalContextModalities (card)
E2BDense + PLE~2.3B effective128KText, image, audio
E4BDense + PLE~4.5B effective128KText, image, audio
12B UnifiedDense~12B256KText, image, audio
26B A4BMoE3.8B / 25.2B256KText, image
31BDense~30.7B256KText, image

Vendor benchmarks on the same card put 31B ahead on LiveCodeBench v6 (80.0%) and MMLU Pro (85.2%), with 26B-A4B close behind (77.1% / 82.6%). Treat those as Google's reported numbers, not independent replications.

How much VRAM Gemma 4 needs

Unsloth's Gemma 4 local guide publishes recommended total memory (RAM plus VRAM, or unified memory on Apple Silicon). Use it as a planning floor; long context and vision projectors add more. For the general method behind these tables, see how much VRAM you need for an LLM.

Variant4-bit8-bitBF16 / FP16
E2B~4 GB5–8 GB~10 GB
E4B5.5–6 GB9–12 GB~16 GB
12B Unified7–8 GB13–14 GB~25 GB
26B A4B16–18 GB28–30 GB~52 GB
31B17–20 GB34–38 GB~62 GB

Ollama's gemma4 library page lists download/runtime bands that line up with those tiers: E2B about 4.6–7.5 GB, E4B about 6.6–9.5 GB, 12B about 7.7–8.0 GB, 26B about 16–19 GB, and 31B about 19–20 GB. If quant names like Q4_K_M are unfamiliar, read GGUF quantization explained.

Which size for which machine

  • Phone / edge or 8 GB laptop: E2B or E4B. Fast multimodal helpers; audio works on these tags.
  • 16 GB GPU or 16–24 GB Mac: 12B Unified for vision and audio on one machine, or a 4-bit 26B-A4B if you want the MoE quality jump and can accept less headroom for context.
  • 24 GB GPU (RTX 3090/4090) or 32 GB+ Mac: 26B-A4B at 4-bit is the usual sweet spot — MoE speed with strong coding and reasoning scores on the vendor table.
  • 32 GB+ VRAM or multi-GPU / large Mac: 31B at 4-bit when you want the densest Gemma 4 quality and can live with slower decode than 26B-A4B.

For a broader shortlist that also includes Qwen and gpt-oss, see open source LLM models to run locally.

How to run Gemma 4 locally with Ollama

Ollama is the lowest-friction path. After installing Ollama:

ollama run gemma4          # default E4B
ollama run gemma4:e2b
ollama run gemma4:12b
ollama run gemma4:26b      # 26B-A4B MoE
ollama run gemma4:31b

Apple Silicon users can pull MLX tags such as gemma4:26b-mlx from the same library page when they want the Metal path. Two settings still catch people:

  1. Context length. Ollama's context-length docs default to only 4K tokens on GPUs under 24 GiB of VRAM. Raise it for coding agents (OLLAMA_CONTEXT_LENGTH=64000 ollama serve is a common pattern Ollama itself documents).
  2. Thinking mode. Gemma 4 thinking is controlled by chat-template / system-prompt conventions; Ollama handles the template, but long "thought" traces still burn tokens and latency. Prefer shorter tasks with thinking off when you only need a direct answer.

Run Gemma 4 with llama.cpp

For full GGUF control, Unsloth publishes Dynamic quants on Hugging Face (for example unsloth/gemma-4-26B-A4B-it-GGUF). With a recent llama.cpp build:

llama-cli -hf unsloth/gemma-4-26B-A4B-it-GGUF:UD-Q4_K_XL \
  --temp 1.0 --top-p 0.95 --top-k 64

Swap the repo tag for E2B, E4B, 12B or 31B. For vision, download the matching mmproj file and pass --mmproj to llama-cli or llama-server. To turn thinking off in llama-server, Unsloth documents --chat-template-kwargs '{"enable_thinking":false}' (escape the quotes differently in PowerShell).

You can also fetch files with the Hugging Face Hub CLI first if you want offline copies:

hf download unsloth/gemma-4-26B-A4B-it-GGUF \
  --include "*UD-Q4_K_XL*" --include "*mmproj-BF16*" \
  --local-dir ./gemma4-26b

Run Gemma 4 on Mac (LM Studio / MLX)

On Apple Silicon, LM Studio and mlx-lm / MLX vision stacks often beat generic GGUF for decode speed. Ollama's *-mlx tags and Unsloth's MLX Dynamic quants both target this path. If you are deciding between engines, our MLX vs GGUF on Mac guide explains the tradeoffs. Practical rule: use MLX when the model is always on one Mac; use GGUF when you need the same files on Windows, Linux and Mac.

Sampling, thinking and multimodal tips

Google's card and Unsloth both recommend the same defaults for Gemma 4:

  • temperature=1.0
  • top_p=0.95
  • top_k=64

Thinking. Enable by putting <|think|> at the start of the system prompt (libraries that expose enable_thinking do the equivalent). When thinking is on, the model emits an internal thought channel before the final answer. For multi-turn chats, keep only the final visible answer in history — do not feed prior thought blocks back in (except where tool-call turns require them).

Images and audio. Put images before the text instruction for best results. Audio is available on E2B, E4B and 12B only, with a roughly 30-second limit called out on the model card. Vision token budgets (70 through 1120) trade detail for speed: higher budgets help OCR and documents; lower budgets help captioning and video-as-frames.

26B-A4B vs 31B. Prefer 26B-A4B when latency and memory matter; prefer 31B when you want the strongest Gemma 4 scores on the vendor table and have the VRAM. That choice mirrors the MoE-vs-dense pattern discussed for other local stacks in run Qwen 3.8 locally.

Who Gemma 4 local is for

Good fit

  • Developers who want a private multimodal assistant on a laptop or single GPU
  • Teams that need Apache 2.0 weights for redistribution or fine-tuning
  • People comparing open models for coding, OCR, charts and tool-calling demos

Weaker fit

  • Pure text coding on a 24 GB card where a dense Qwen 3.8 27B is already tuned in — compare both on your own prompts
  • Edge devices that only need embeddings (use EmbeddingGemma 2 locally instead)
  • Multi-tenant serving where vLLM with NVFP4 / MTP is the real goal — Gemma 4 works there, but start with the desktop path above first

Pros and cons

ProsCons
Clear size ladder from edge to workstation31B at high precision is heavy for one consumer GPU
Official Ollama tags plus strong GGUF/MLX ecosystemThinking traces can inflate latency if left on for every chat
Apache 2.0 on published Gemma 4 cardsAudio only on the smaller three sizes
Native multimodal + system roleCommunity GGUFs with separate mmproj files can confuse older Ollama builds — prefer official tags or llama.cpp for vision

FAQ

How do I run Gemma 4 locally on 16 GB VRAM?

Use ollama run gemma4:12b for a comfortable multimodal laptop build, or a 4-bit 26B-A4B GGUF if you accept tighter headroom for context. Avoid BF16 26B/31B on 16 GB — those need roughly 50 GB+ of total memory per Unsloth's table.

What is Gemma 4 26B A4B?

It is the mixture-of-experts member of the family: about 25.2B total parameters with roughly 3.8B active per token, 256K context, and text plus image inputs per the official model card. The "A4B" label means about 4B active parameters.

Does Ollama support Gemma 4 vision?

Yes on the official gemma4 library tags. If a third-party GGUF with a separate mmproj fails to load in Ollama, update Ollama or run that file in llama.cpp with --mmproj instead.

Gemma 4 26B-A4B or 31B for coding?

Vendor LiveCodeBench and Codeforces numbers favour 31B slightly, but 26B-A4B is usually the better local default because MoE decode is faster at similar quality for many prompts. Benchmark both on your own repository tasks before committing disk space.

Is Gemma 4 open source or only open weight?

The published Gemma 4 Hugging Face cards list Apache 2.0, which is a true open-source licence for the weights and accompanying materials as released. Always read the exact card for the tag you download; older Gemma generations used different terms.