The practical shortlist of open source LLM models to run locally in late 2026 is Qwen3 / Qwen3.8, Gemma 4, gpt-oss, and DeepSeek-R1 distill tags, chosen by how much memory you have. Install Ollama, pick a tag that fits your machine, and run it with one command. File sizes below come from the Ollama library pages; licences come from the official model cards.

Most people searching this phrase want downloadable weights they can run offline, not a lecture on licences. Still, "open source" and "open weight" are not the same thing. Weights you can download are usually open-weight; only a few releases (for example Ai2's OLMo) also publish training data and training code under OSI-style terms. Our open weight vs open source guide covers that spectrum. Here we focus on models you can actually pull and chat with today.

Pick by memory first

Memory, not a leaderboard score, decides whether a model feels usable. Weights need room; context needs more. For the general method, see how much VRAM you need for an LLM. As a starting point, match the Ollama download size to your GPU VRAM or Mac unified memory, then leave a few gigabytes free for the KV cache.

Memory availableSensible starter (Ollama tag)Download size (library)Why start here
About 8 GBqwen3:8b or gemma3:4b5.2 GB / 3.4 GBFits with headroom for short context
About 16 GBgpt-oss:20b or gemma3:12b14 GB / 8.2 GBStrong reasoning or multimodal laptop pick
About 24 GBqwen3.8 (27B) or gemma3:27b18 GB / 17 GBDense quality on a single consumer GPU
64 GB+ / 80 GB GPUgpt-oss:120b65 GBLarge MoE for harder agent and reasoning work

Those sizes are the published Ollama builds: qwen3, qwen3.8, gemma3, and gpt-oss. Raising context length adds memory on top; Ollama's own context-length docs default to only 4K tokens on GPUs under 24 GiB of VRAM.

Best all-rounders: Qwen3 and Qwen3.8

Qwen is the default recommendation for most local users who want Apache 2.0 weights.

  • Qwen3-8B is a dense 8.2B model with thinking and non-thinking modes, multilingual support, and an Apache 2.0 licence on its model card. Ollama's qwen3:8b tag is a 5.2 GB download.
  • Qwen3.8-27B is the denser, vision-capable follow-up: 27B parameters, native 262,144-token context (extendable with YaRN), thinking on by default, Apache 2.0, per the Qwen3.8-27B card. Ollama's default qwen3.8 / qwen3.8:27b tag is an 18 GB Q4_K_M build with image input and a 256K context window.
ollama run qwen3:8b
ollama run qwen3.8

For VRAM tiers, quant choices and thinking controls on the 27B, use our dedicated run Qwen 3.8 locally guide. Qwen is strong when you want one model for coding, chat and multilingual work without a custom licence.

Pros: permissive Apache 2.0 licence; excellent Ollama and Hugging Face support; 27B adds vision. Cons: thinking mode can be slow if you leave it at full effort; the 27B wants roughly 16–19 GB of total memory at 4-bit.

Best multimodal laptop pick: Gemma 3 and Gemma 4

Google DeepMind's Gemma line is the other high-quality single-GPU family.

  • Gemma 3 ships 270M through 27B sizes. Ollama lists gemma3:4b at 3.4 GB, gemma3:12b at 8.2 GB (text + image), and gemma3:27b at 17 GB, with a 128K context window on the multimodal tags (gemma3 library page). Older Gemma 3 weights use the Gemma Terms of Use, not a plain Apache grant — read that page before commercial redistribution.
  • Gemma 4 moves to Apache 2.0 on the Hugging Face cards (for example google/gemma-4-12B-it). Ollama's gemma4 page lists tags such as 12b (about 8 GB) and 26b / 31b (about 19–20 GB), aimed at laptop-to-server local runs with multimodal inputs.
ollama run gemma3:12b
ollama run gemma4:12b

Pick Gemma when you care about image (and, on Gemma 4 smaller tags, audio/video) understanding on one machine. For embedding-only Google models, see EmbeddingGemma 2 locally.

Pros: strong small and mid sizes; Gemma 4 is Apache 2.0 on the published cards. Cons: Gemma 3 still carries custom terms; always check the exact card for the tag you pull.

Best Apache reasoning models: gpt-oss

gpt-oss is OpenAI's open-weight reasoning pair. The gpt-oss-20b card and Ollama's gpt-oss page agree on the essentials:

ModelParameters (active)Ollama sizeOfficial memory guidance
gpt-oss-20b~21B total / 3.6B active14 GBRuns within about 16 GB of memory
gpt-oss-120b~117B total / 5.1B active65 GBFits a single 80 GB GPU

Both use Apache 2.0, configurable reasoning effort, and MXFP4 MoE weights. They are text-only with a 128K context window.

ollama run gpt-oss:20b
ollama run gpt-oss:120b

Full setup across Ollama, LM Studio, llama.cpp and vLLM is in how to run gpt-oss locally.

Pros: permissive licence; strong tool-use and reasoning story; 20b fits many 16 GB machines. Cons: no vision; 120b is workstation-class hardware.

Best local reasoning distillations: DeepSeek-R1

The full DeepSeek-R1 checkpoint is a 671B MoE (37B active) under MIT on its model card — powerful, but the Ollama 671b tag is a 404 GB download, so it is not a laptop default. What most local users actually run are the distill sizes Ollama publishes under deepseek-r1:

Ollama tagDownload sizeTypical role
deepseek-r1:8b5.2 GBEntry reasoning on 8–16 GB machines
deepseek-r1:14b9.0 GBMid-size distill
deepseek-r1:32b20 GBStrong dense distill on ~24 GB
ollama run deepseek-r1:14b

Licence nuance matters: the R1 card says the main R1 weights are MIT, but the Qwen-based distills inherit Apache 2.0 from Qwen, and the Llama-based distills inherit the Llama community licence. Check the specific distill card before shipping a product.

Pros: excellent math and coding reasoning for the size; MIT on the flagship weights. Cons: full 671B is impractical locally; distill licences follow their base models.

Other solid local options

  • Llama 3.2 small tags (llama3.2:1b / 3b, about 1.3–2.0 GB on Ollama) are handy on phones and tiny boxes, but Meta's Llama licence is not Apache — acceptable for many personal uses, stricter for large commercial deployments.
  • Mistral classic 7B (mistral, 4.4 GB on Ollama) remains a clean, well-supported chat baseline.
  • SmolLM 2 (smollm2:1.7b, 1.8 GB) is Apache 2.0 on its model card and useful for on-device experiments.
  • Extremely compressed specialty builds such as ternary Bonsai 2 27B can fit in roughly 6–7 GB, but need Prism's fork rather than stock Ollama.

How to run them (three common tools)

Ollama (fastest start)

  1. Install Ollama from ollama.com.
  2. Run a tag from the tables above.
  3. Point any OpenAI-compatible client at http://localhost:11434.

Compare runners in Ollama vs llama.cpp vs LM Studio.

LM Studio (GUI)

LM Studio browses Hugging Face GGUF and MLX builds, estimates fit for your machine, and exposes a local API. Prefer it when you want a desktop UI and Mac MLX paths; see MLX vs GGUF on Mac.

llama.cpp / GGUF (full control)

llama.cpp is the engine behind most GGUF apps. Download a quant from Hugging Face, then serve with llama-server or your preferred frontend. Quant names (Q4_K_M and friends) are explained in GGUF quantization explained. For multi-user GPU serving, compare vLLM vs llama.cpp.

Who this shortlist is for

You are…Start with
On an 8 GB laptop GPUqwen3:8b or gemma3:4b
On 16 GB, want reasoninggpt-oss:20b
On 16–24 GB, want visiongemma3:12b / gemma4:12b or qwen3.8
Building a private coding assistantQwen3.8 or a DeepSeek-R1 distill
Need OSI-style full opennessLook at OLMo, then accept a capability trade-off

Skip giant MoE flags (DeepSeek-R1 671B, huge Qwen MoE tags) unless you have server memory. Skip custom ternary forks until stock runners already feel too big — start with a standard GGUF or Ollama tag first.

FAQ

What are the best open source LLM models to run locally right now?

For most people: Qwen3-8B on small machines, gpt-oss-20b or Gemma 3/4 12B on 16 GB, and Qwen3.8-27B on about 24 GB. Those picks balance download size, licence clarity and tooling support (Ollama, LM Studio, llama.cpp).

How do I use open source LLM models without a GPU?

CPU-only inference works with llama.cpp and Ollama, but a dense 8B+ model will be slow. Prefer tiny tags (smollm2, gemma3:1b, qwen3:1.7b) or a Mac with enough unified memory. Treat anything above about 14 GB of weights as "possible but painful" on CPU alone.

Are Llama and Gemma really open source?

They are downloadable open-weight models. Llama uses Meta's community licence. Gemma 3 uses Google's Gemma Terms of Use; Gemma 4's published Hugging Face cards mark Apache 2.0. If you need a classic permissive grant, prefer Qwen, gpt-oss, SmolLM2 or OLMo, and still read the exact card.

How do I deploy an open source LLM locally for an app?

Run Ollama or llama.cpp as a local HTTP server and call it with an OpenAI-compatible client. For several concurrent users on one GPU, prefer vLLM or SGLang. Keep context length explicit — defaults are often much shorter than the model's maximum.

Which free open source LLM models are safe for commercial products?

Apache 2.0 and MIT weights (Qwen3/3.8, gpt-oss, SmolLM2, DeepSeek-R1's MIT card, Gemma 4 per its HF cards) are the usual starting point. Llama and Gemma 3 add extra conditions. This is not legal advice: read the licence file shipped with the exact revision you ship.

Should I fine-tune these models or just prompt them?

Prompt and tool-call first. Fine-tune when you have a narrow style or domain that prompting cannot fix. Local fine-tuning options are covered in how to fine-tune an LLM locally and LoRA vs QLoRA.