To run Qwen 3.8 27B locally, the fastest route is Ollama: install it and type ollama run qwen3.8, which pulls an 18 GB 4-bit build. A 4-bit quant needs roughly 16 to 19 GB of total memory, so it fits a 24 GB GPU, a 16 GB GPU with some spill-over into system RAM, or a Mac with 24 GB of unified memory. The model is free under the Apache 2.0 licence, handles text, images and video, and also runs in llama.cpp, LM Studio, Unsloth Desktop and vLLM.
This guide covers what the model is, which file to download for your hardware, setup commands for each runtime, and the settings that matter most: thinking mode, sampling and context length.
What Qwen 3.8 27B is
Qwen3.8-27B is the dense, deployment-friendly member of the Qwen team's Qwen 3.8 family. According to the official model card, it is a 27-billion-parameter vision-language model with thinking mode on by default, a native context window of 262,144 tokens that can be stretched to about one million with YaRN scaling, and built-in multi-token prediction (MTP) layers. The weights are published under Apache 2.0, so commercial use and fine-tuning are allowed. For background on licences, see open weight vs open source.
The architecture is unusual, and it matters for memory. Of its 64 layers, only 16 use standard attention; the other 48 use Gated DeltaNet, a linear-attention layer that keeps a fixed-size state instead of a growing cache. The vLLM recipe describes it the same way: linear attention on 48 of 64 layers, a vision tower and a built-in MTP draft head.
| Fact | Qwen3.8-27B |
|---|---|
| Parameters | 27B, dense (all active) |
| Inputs | Text, images, video |
| Native context | 262,144 tokens (up to about 1M with YaRN) |
| Attention | 16 full-attention layers, 48 Gated DeltaNet layers |
| Licence | Apache 2.0 |
| Default mode | Thinking on, reasoning effort xhigh |
The rest of the family is a different story. Qwen3.8-Flash-Next is an experimental mixture-of-experts model with 125B parameters (6B active) plus a large n-gram embedding table, released under a separate Qwen community licence rather than Apache 2.0, per its model card. The 2.4T-A95B flagship is server-class hardware territory. For a desk-side machine, the 27B is the one to run.
How much memory Qwen 3.8 27B needs
Unsloth's Qwen 3.8 run guide publishes a requirements table in total memory (RAM plus VRAM, or unified memory on a Mac). Its rule of thumb: RAM plus VRAM should roughly equal the quant size, or it still works but much more slowly.
| Precision | Total memory (Unsloth) | Example file size |
|---|---|---|
| 2-bit | 9-11 GB | UD-Q2_K_XL, 9.8 GB |
| 3-bit | 12-14 GB | UD-Q3_K_XL, 13.2 GB |
| 4-bit | 16-19 GB | UD-Q4_K_XL, 17.6 GB |
| 6-bit | 23-26 GB | UD-Q6_K, 22.0 GB |
| 8-bit | 31 GB | Q8_0, 29.1 GB |
| BF16 | 56 GB | Full precision |
File sizes come from the unsloth/Qwen3.8-27B-GGUF repository. Add about 0.9 GB for the vision projector (the mmproj file) if you want image input, and Unsloth advises 1 to 2 GB of extra headroom if you turn on MTP. If quant names like Q4_K_XL are new to you, our GGUF quantization guide decodes them.
Context memory is small for its size
Because only 16 layers keep a KV cache, long context is cheaper than on a typical 27B model. From the model's config.json (4 KV heads, head size 256, 16 full-attention layers), the cache works out to about 64 KiB per token at 16-bit precision. That is our arithmetic, not a vendor figure:
| Context | KV cache at f16 | At q8_0 |
|---|---|---|
| 8K tokens | about 0.5 GiB | about 0.27 GiB |
| 32K tokens | about 2 GiB | about 1.1 GiB |
| 128K tokens | about 8 GiB | about 4.3 GiB |
| 262K tokens | about 16 GiB | about 8.5 GiB |
The linear-attention layers add a small fixed state on top that does not grow with context. To shrink the cache further, see KV cache quantization; for the general method, see how much VRAM you need for an LLM.
Which file for which machine
- 24 GB GPU (RTX 3090, 4090) or 32 GB Mac: 4-bit (UD-Q4_K_XL or Ollama's default) with room for 32K or more of context.
- 16 GB GPU: 3-bit (UD-Q3_K_XL) fits on the card; 4-bit works with a few layers in system RAM, at lower speed.
- 8 GB GPU plus 16 GB or more of RAM: a 2-bit or 3-bit file split across GPU and RAM. It runs, but a dense 27B model slows sharply when most layers sit in system memory, so expect a big drop in speed.
- Two 24 GB GPUs, or a 64 GB Mac: 8-bit or 6-bit for the best quality short of BF16.
Run it with Ollama
The Ollama library page lists twelve tags. The default qwen3.8 (same as qwen3.8:27b) is an 18 GB Q4_K_M build with a 256K context window and image input. Other tags include 27b-q8_0 (30 GB), 27b-bf16 (56 GB), MLX builds for Apple Silicon and -mtp variants that keep the MTP layers.
ollama run qwen3.8
# or a specific build
ollama run qwen3.8:27b-q8_0
Two settings catch people out. First, Ollama's context-length docs say it defaults to only 4K tokens of context on GPUs under 24 GiB of VRAM, and 32K between 24 and 48 GiB. For coding tools or agents, Ollama itself recommends at least 64,000 tokens, which you can set with OLLAMA_CONTEXT_LENGTH=64000 ollama serve. Second, thinking is on by default. In API calls, set "think": false to request a direct answer, as described in Ollama's thinking docs.
Run it with llama.cpp
The llama.cpp project publishes its own conversion at ggml-org/Qwen3.8-27B-GGUF, with Q4_K_M, Q8_0 and BF16 files plus separate MTP and DFlash draft files. Its model card gives a one-line start:
llama serve -hf ggml-org/Qwen3.8-27B-GGUF
For Unsloth's dynamic quants, download one file and point llama-cli or llama-server at it, with the recommended thinking-mode sampling:
hf download unsloth/Qwen3.8-27B-GGUF --local-dir qwen38 --include "*UD-Q4_K_XL*"
llama-server -m qwen38/Qwen3.8-27B-UD-Q4_K_XL.gguf \
--temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.0 \
--chat-template-kwargs '{"reasoning_effort":"medium"}'
The --chat-template-kwargs flag, from Unsloth's guide, lowers reasoning depth for faster answers. llama.cpp also lists a draft-mtp speculative decoding type that uses a model's own MTP heads; see our guide to speculative decoding vs MTP before you enable it.
Run it in LM Studio, Unsloth Desktop or MLX
LM Studio users can search for Qwen3.8-27B in the app. The lmstudio-community account on Hugging Face publishes both a GGUF build and MLX builds at 4, 5, 6 and 8 bits. On Apple Silicon, the MLX files run through Apple's MLX engine, also available as mlx-lm if you prefer a command line; our MLX vs GGUF on Mac guide explains which to pick.
Unsloth Desktop, a free app for macOS, Windows and Linux, can search, download and run the same GGUF and MLX files, with thinking toggles built in. It also fine-tunes, which is handy if you later want to fine-tune an LLM locally.
Run it with vLLM or SGLang
For a GPU server or several users at once, the model card recommends dedicated serving engines such as vLLM, SGLang or TokenSpeed. The vLLM recipe's single-GPU command for an NVFP4 build on a Blackwell card is:
vllm serve Inferact/Qwen3.8-27B-NVFP4 \
--max-model-len 262144 --kv-cache-dtype fp8 \
--reasoning-parser qwen3 \
--enable-auto-tool-choice --tool-call-parser qwen3_xml
Add --speculative-config '{"method":"mtp","num_speculative_tokens":3}' to use the built-in MTP head. NVFP4 needs an NVIDIA Blackwell GPU; on older cards, use the official Qwen/Qwen3.8-27B-FP8 checkpoint or a GGUF runtime. If you are choosing between these engines, read vLLM vs llama.cpp.
Settings that matter
The official model card recommends two sampling presets:
| Setting | Thinking mode | Non-thinking mode |
|---|---|---|
| temperature | 1.0 | 0.7 |
| top_p | 0.95 | 0.80 |
| top_k | 20 | 20 |
| presence_penalty | 0.0 | 1.5 |
Thinking is controlled per request. Pass enable_thinking: false inside chat_template_kwargs for a direct answer, or set reasoning_effort to xhigh (the default), medium or low. The card adds a caution worth repeating: in multi-step agent work, lower effort can mean more failed attempts, so it does not always save time overall. A related default, preserved thinking, keeps earlier reasoning in the conversation; it uses more tokens and can be turned off with preserve_thinking: false.
Pros, cons and who it's for
- Pros: Apache 2.0 licence; fits one consumer GPU at 4-bit; images and video as well as text; long context is cheap thanks to the hybrid attention design; already packaged for Ollama, llama.cpp, LM Studio, MLX and vLLM.
- Cons: thinking mode makes answers long and slow unless you lower the effort; a dense 27B model is slower than small-active MoE models on the same hardware; 16 GB cards need a 3-bit file or partial offload.
- Who it's for: developers and tinkerers with a 24 GB GPU or a 24 GB-plus Mac who want a capable private model for coding, documents and image questions. On 8 GB of VRAM, a smaller model will feel far more responsive.
FAQ
Can Qwen 3.8 27B run on 16 GB of VRAM?
Yes. A 3-bit file such as UD-Q3_K_XL (13.2 GB) fits on a 16 GB card with room for modest context. A 4-bit file needs 16 to 19 GB in total, so on a 16 GB card some layers move to system RAM and generation slows.
Can I run Qwen 3.8 27B on 8 GB of VRAM?
Only with offloading. Pick a 2-bit or 3-bit file and make sure GPU memory plus system RAM at least matches the file size. It will work, but most of the model runs from system memory, so it is much slower than on a 24 GB card.
How do I turn off thinking in Qwen 3.8?
In an OpenAI-compatible API call, pass enable_thinking: false inside chat_template_kwargs. In Ollama's API, set "think": false. You can also keep thinking on but set reasoning_effort to medium or low for shorter traces.
Is Qwen 3.8 27B free for commercial use?
Yes. The 27B model is published under Apache 2.0, which permits commercial use, modification and redistribution. Qwen3.8-Flash-Next uses a separate Qwen community licence, so read that one before shipping a product on it.
Can I run Qwen 3.8 on a Mac?
Yes. A Mac with 24 GB of unified memory can run a 4-bit build. Use Ollama's MLX tags, LM Studio's MLX builds or mlx-lm for the best speed on Apple Silicon, or any GGUF build through llama.cpp.
Can I run the Qwen 3.8 2.4T model locally?
Not on ordinary hardware. Unsloth's smallest 1-bit build of Qwen3.8-2.4T-A95B is 397 GB and, by their guidance, needs at least 450 GB of RAM. For a single workstation, the 27B model is the practical choice.



