To run gpt-oss locally, the quickest route is Ollama: install it and type ollama run gpt-oss:20b. The 20b model needs about 16 GB of memory, so it fits on a 16 GB GPU or a Mac with 16 GB or more of unified memory. The 120b model is designed to fit on a single 80 GB GPU, or a Mac with a large amount of unified memory. Both are free to download under the Apache 2.0 licence, and you can also run them with LM Studio, llama.cpp, vLLM or Hugging Face Transformers.

This guide covers what the two models are, the hardware each needs, setup commands for every major runtime, and the settings that trip people up.

What gpt-oss is

gpt-oss is OpenAI's family of open-weight reasoning models, released on August 5, 2025. It was OpenAI's first open-weight language model release since GPT-2. As of October 2026, OpenAI has not released a general-purpose successor; the only follow-up is gpt-oss-safeguard, a research-preview pair of safety-classification models fine-tuned from gpt-oss.

Key facts from the official repository and model card:

  • Two sizes. gpt-oss-120b has 36 layers, about 117 billion total parameters and 5.1 billion active per token. gpt-oss-20b has 24 layers, about 21 billion total and 3.6 billion active.
  • Mixture of experts. Only a few experts run for each token, which keeps generation fast.
  • Native MXFP4. The expert weights were post-trained in the 4-bit MXFP4 format, so the official download is already compact.
  • Text only, with a context window of about 128,000 tokens.
  • Licence: Apache 2.0, plus OpenAI's gpt-oss usage policy. Our explainer on open weight vs open source covers what that means.
  • Not on the OpenAI API. You run it yourself or through a hosting provider.

Hardware: what you need

gpt-oss-20bgpt-oss-120b
Official memory guidanceRuns within 16 GB of memoryFits a single 80 GB GPU, such as an H100 or MI300X
Ollama download size14 GB65 GB
Realistic local hardware16 GB GPU, or a Mac with 16 GB or more80 GB GPU, multiple GPUs, or a Mac with enough unified memory for 65 GB plus context
Best forLaptops, desktops, low-latency toolsWorkstations and servers, harder reasoning tasks

These figures cover the weights. Your context window needs extra memory, although gpt-oss is efficient here: half its layers use a 128-token sliding window, so long contexts cost much less than in a typical dense model. Our VRAM guide works through the numbers. If the model does not fit, llama.cpp-based tools can offload some layers to system RAM, but generation slows down.

Option 1: Ollama (easiest)

Install Ollama, then:

ollama run gpt-oss:20b
# or, with enough memory
ollama run gpt-oss:120b

Ollama supports MXFP4 natively, so it runs the official weights without re-quantizing. Your apps can reach it at http://localhost:11434/v1 using any OpenAI-compatible client.

To set reasoning effort through the API, pass a think value of low, medium or high. Ollama's thinking docs show that gpt-oss defaults to medium. If you use gpt-oss with coding tools or long documents, raise Ollama's context length, because the default is only 4k tokens on GPUs under 24 GiB.

Option 2: LM Studio (desktop app)

In LM Studio, search for gpt-oss and download it, or use the CLI:

lms get openai/gpt-oss-20b
lms server start

LM Studio gives you a chat window and a local server on port 1234. On Apple Silicon it can use either its llama.cpp or MLX runtime; see MLX vs GGUF on Mac. See our comparison of Ollama vs llama.cpp vs LM Studio if you are undecided.

Option 3: llama.cpp (most control)

llama.cpp can download the official GGUF conversion straight from the Hugging Face Hub:

llama serve -hf ggml-org/gpt-oss-20b-GGUF --reasoning-effort high

That starts an OpenAI-compatible server and web UI on http://127.0.0.1:8080. The --reasoning-effort flag passes your chosen level to the chat template. Add -c to set the context size, and use the GPU-offload flags if you need to split the 120b model between GPU and CPU.

Option 4: vLLM (serving many users)

On an NVIDIA or AMD GPU server, vLLM gives the best throughput when several people or processes share the model; our vLLM vs llama.cpp guide explains why. Current vLLM releases list GPT-OSS among supported mixture-of-experts architectures:

uv pip install vllm
vllm serve openai/gpt-oss-20b

The server listens on port 8000 by default. OpenAI's own README pins a special early build from launch week; with a current release, follow the vLLM docs instead.

Option 5: Transformers (for Python and research)

Transformers loads gpt-oss with a standard pipeline, and its chat template applies the harmony format for you:

from transformers import pipeline

pipe = pipeline("text-generation", model="openai/gpt-oss-20b",
                torch_dtype="auto", device_map="auto")
messages = [{"role": "user", "content": "Explain MXFP4 in two sentences."}]
print(pipe(messages, max_new_tokens=256)[0]["generated_text"][-1])

This route is best for experiments, interpretability work and fine-tuning prep rather than daily chat.

Settings that matter

Reasoning effort

gpt-oss supports three reasoning levels: low for fast replies, medium for balance, and high for deeper analysis. The model card on Hugging Face says you can set it in the system prompt, for example "Reasoning: high". Runtimes expose the same setting as a flag or API field. Higher effort means more thinking tokens, which means slower answers and more context used.

The harmony format

Both models were trained on OpenAI's harmony response format and, per OpenAI, will not work correctly without it. Ollama, LM Studio, llama.cpp's built-in template and the Transformers chat template handle this automatically. Problems usually appear only when you build raw prompts by hand or use an outdated GGUF with a broken template; if output looks garbled, update your runtime and re-download the model.

Chain of thought

gpt-oss returns its full reasoning trace separately from the final answer. OpenAI says this trace is for debugging and is not intended to be shown to end users, so hide it in customer-facing apps.

Pros and cons

  • Pros: permissive Apache 2.0 licence; 20b runs on common hardware; fast thanks to mixture of experts; adjustable reasoning effort; built-in support for function calling, web browsing and Python tools; supported by every major runtime.
  • Cons: text only, with no image input; over a year old, with no general-purpose successor; newer open-weight families such as Qwen and Gemma have released since; the 120b model still needs data-centre-class memory.

Fine-tuning gpt-oss

OpenAI says gpt-oss-20b can be fine-tuned on consumer hardware and gpt-oss-120b on a single H100 node. Unsloth publishes free Colab notebooks for gpt-oss-20b fine-tuning and GRPO reinforcement learning. Fine-tune with the harmony format, then export back to GGUF or serve with vLLM.

Who should run gpt-oss locally?

  • Developers who want a capable local reasoning model with a clean licence for commercial work.
  • Privacy-sensitive teams who cannot send data to a hosted API.
  • Researchers studying reasoning traces, tool use or safety behaviour in an open model.
  • Anyone with a 16 GB GPU or Mac who wants a strong all-rounder without heavy quantization.

FAQ

Can I run gpt-oss-20b on a 16 GB GPU?

Yes. OpenAI says gpt-oss-20b runs within 16 GB of memory thanks to its native MXFP4 weights, and Ollama's download is 14 GB. Keep the context window modest so the KV cache still fits.

Can I run gpt-oss on a Mac?

Yes. Ollama, LM Studio and llama.cpp all support Apple Silicon. The 20b model suits Macs with 16 GB or more of unified memory; the 120b model needs enough unified memory for its 65 GB download plus context.

Is gpt-oss free for commercial use?

It is released under the Apache 2.0 licence, which allows commercial use, modification and redistribution, subject to OpenAI's gpt-oss usage policy.

How do I change gpt-oss reasoning effort?

Set "Reasoning: low", "Reasoning: medium" or "Reasoning: high" in the system prompt, or use your runtime's control: the think field in Ollama's API or the --reasoning-effort flag in llama.cpp.

Is gpt-oss available through the OpenAI API?

No. OpenAI's help centre says the gpt-oss models are not available through the OpenAI API. You run them on your own hardware or through a third-party host.

Does gpt-oss support images?

No. Both gpt-oss models are text-only. For local image input, look at multimodal open-weight families instead.