vLLM vs llama.cpp comes down to who you are serving. vLLM is built to serve many simultaneous requests on data-centre GPUs with the highest total throughput. llama.cpp is built to run one model efficiently on almost any hardware, from a laptop CPU to a single consumer GPU or a Mac, usually with quantized GGUF files. If you are one person running a model locally, llama.cpp, or a wrapper like Ollama, is the better fit. If you are putting a model behind an API for a team or a product on NVIDIA or AMD GPUs, start with vLLM, and evaluate SGLang alongside it.
vLLM vs llama.cpp vs SGLang vs Ollama at a glance
| vLLM | SGLang | llama.cpp | Ollama | |
|---|---|---|---|---|
| Written in | Python and PyTorch, with custom kernels | Python and PyTorch, with custom kernels | C and C++ | Go, with llama.cpp and MLX engines |
| Licence | Apache 2.0 | Apache 2.0 | MIT | MIT |
| Sweet spot | High-throughput multi-user GPU serving | High-throughput serving, agent workloads, RL rollouts | Local and edge inference on any hardware | Easy local model server |
| Typical model format | Hugging Face safetensors, plus FP8, AWQ, GPTQ, GGUF and more | Hugging Face safetensors and quantized variants | GGUF | Its own library, GGUF imports, MLX on Macs |
| Start command | vllm serve <model> | sglang serve <model> | llama serve -hf <repo> | ollama run <model> |
| Default port | 8000 | 30000 in the official examples | 8080 | 11434 |
| OpenAI-compatible API | Yes | Yes | Yes | Yes |
| CPU-only use | Supported, but not its focus | Supported on some CPUs | A core strength | Yes |
How vLLM works and why it is fast
vLLM is "a fast and easy-to-use library for LLM inference and serving." Its signature idea is PagedAttention, introduced in the 2023 paper Efficient Memory Management for Large Language Model Serving with PagedAttention. The KV cache, the per-request memory that grows with every token, is stored in fixed-size blocks, much like an operating system pages memory. The paper reports near-zero wasted KV cache memory and two to four times higher throughput than the systems it compared against at the same latency.
On top of that, vLLM's README lists continuous batching of incoming requests, chunked prefill, prefix caching, speculative decoding, tensor, pipeline, data and expert parallelism, structured outputs, tool calling and efficient multi-LoRA serving. It loads models directly from the Hugging Face Hub, supports more than 200 architectures, and runs on NVIDIA, AMD and Intel GPUs and on x86, ARM and PowerPC CPUs, with plugins for TPUs, Gaudi, Ascend and others.
uv pip install vllm
vllm serve Qwen/Qwen3-8B
# OpenAI-compatible server on http://localhost:8000
One quirk to know: according to the vLLM quickstart, its OpenAI-compatible server hosts one model at a time, so multiple models usually means multiple processes behind a router.
How llama.cpp works and why it is everywhere
llama.cpp is plain C and C++ with no required dependencies. It targets almost every backend: CUDA, HIP for AMD, Metal for Apple Silicon, Vulkan, SYCL for Intel, many CPUs and more. Its strength is running quantized GGUF models, from roughly 1.5-bit to 8-bit, with CPU plus GPU hybrid inference when a model does not fit in VRAM.
The built-in server is more capable than its "local tool" reputation suggests. Its README lists parallel decoding with multiple server slots, continuous batching, speculative decoding, function calling, OpenAI and Anthropic-compatible routes, a router mode that can load several models, and a web UI.
llama serve -hf ggml-org/gpt-oss-20b-GGUF
# OpenAI-compatible server and web UI on http://127.0.0.1:8080
Where SGLang fits
SGLang is the other major open-source serving engine, released under Apache 2.0 and positioned for agentic workloads, RL rollouts and large-scale serving. Its original paper introduced RadixAttention, which reuses KV cache across requests that share a prefix, a common pattern when agents resend long system prompts and tool definitions. SGLang's README lists NVIDIA, AMD, Google TPU, Intel, Apple Silicon, Huawei Ascend and Moore Threads hardware, and several RL post-training frameworks use it to generate rollouts.
For a team choosing a GPU server engine, vLLM and SGLang are the two to benchmark against each other on your own model, traffic and hardware. Both move fast, and published comparisons go stale within months.
Where Ollama fits
Ollama is a convenience layer, not a high-throughput engine. It is ideal for a developer laptop or a small internal box, and it can handle a few concurrent requests, but it is not designed to squeeze maximum tokens per second out of a multi-GPU server. Our comparison of Ollama vs llama.cpp vs LM Studio covers the local side in depth.
Throughput vs latency: the distinction that decides it
Most "which is faster" arguments mix up two different measurements.
- Single-user speed is how quickly one person sees tokens appear. On one consumer GPU with a quantized model, llama.cpp is very competitive, and it can run models that would not fit in vLLM's default 16-bit format at all.
- Aggregate throughput is the total tokens per second across many users. Here, vLLM and SGLang's batching and KV cache management let one GPU serve far more requests at once.
So a benchmark of one prompt at a time on a gaming PC tells you little about serving 50 users on an H100, and the reverse is just as true.
Hardware and model format considerations
- Memory headroom. vLLM pre-allocates most of the GPU's memory for weights and KV cache by default, so plan for it to own the GPU. llama.cpp allocates what the model and context need.
- Quantization. vLLM supports FP8, MXFP4, NVFP4, INT8, INT4, GPTQ, AWQ, GGUF and more, but its fastest paths are the GPU-native formats. llama.cpp's sweet spot is GGUF k-quants and i-quants; see our GGUF quantization guide.
- Apple Silicon. llama.cpp and MLX-based tools are the natural fit on a Mac. vLLM and SGLang list Apple Silicon support, but the Mac is not their main target.
- Model support. vLLM and SGLang usually support new Hugging Face architectures quickly, because they build on Transformers-style model definitions in PyTorch. llama.cpp needs each new architecture implemented in C++, which often happens within days for popular models.
Pros and cons
vLLM
- Pros: top-tier multi-user throughput; PagedAttention and prefix caching; broad GPU and quantization support; multi-LoRA serving; OpenAI and Anthropic-compatible APIs; Apache 2.0.
- Cons: heavier install with PyTorch and CUDA dependencies; aimed at GPUs; one model per server process; overkill for a single user.
SGLang
- Pros: RadixAttention prefix reuse suits agents and multi-turn chat; strong at scale; wide hardware list; popular for RL rollouts.
- Cons: similar heavy dependencies to vLLM; fast-moving, so pin versions.
llama.cpp
- Pros: runs nearly anywhere; tiny footprint; excellent quantization; hybrid CPU and GPU offload; capable server with batching and router mode; MIT licence.
- Cons: lower peak throughput than vLLM or SGLang on large GPU servers; GGUF conversion needed; fewer distributed-serving features.
Which should you choose?
- Personal use on a laptop, desktop or Mac: llama.cpp, or Ollama or LM Studio on top of it.
- A small team sharing one consumer GPU: llama.cpp's server with several parallel slots is often enough and much simpler to run.
- A product or internal API on data-centre GPUs: vLLM, and benchmark SGLang against it.
- Agent-heavy traffic with long shared prompts, or RL training rollouts: give SGLang a serious look.
- Serving many fine-tuned LoRA adapters on one base model: vLLM's multi-LoRA support.
For a concrete example, our guide on how to run gpt-oss locally shows the same model served with Ollama, llama.cpp and vLLM, and gpt-oss or Qwen models make good test subjects. Before buying hardware for either engine, size the memory with our VRAM guide.
Who this is for
Developers and ML engineers deciding how to serve open-weight models, from a single workstation to a GPU cluster.
FAQ
Is vLLM faster than llama.cpp?
For many concurrent users on data-centre GPUs, vLLM typically delivers much higher total throughput thanks to batching and PagedAttention. For a single user on consumer hardware with a quantized model, llama.cpp is often just as fast or faster, and it runs on hardware vLLM does not target.
Should I use vLLM or Ollama?
Use Ollama for local development and personal use, where easy setup matters most. Use vLLM when you need to serve many requests on GPUs in production.
What is the difference between vLLM and SGLang?
Both are Apache 2.0 GPU serving engines with OpenAI-compatible APIs. vLLM is known for PagedAttention and broad hardware support; SGLang is known for RadixAttention prefix reuse and is popular for agent workloads and RL rollouts. Benchmark both on your own workload.
Can vLLM run GGUF models?
vLLM lists GGUF among its supported quantization formats, but its fastest paths are GPU-native formats such as FP8, AWQ and GPTQ. For GGUF, llama.cpp remains the reference engine.
Can llama.cpp serve multiple users?
Yes. Its server supports parallel decoding with multiple slots and continuous batching, which is enough for a small team. It is not designed to match vLLM's peak throughput on large GPU servers.
Does vLLM work without a GPU?
vLLM supports x86, ARM and PowerPC CPUs, but it is designed for GPUs. For CPU-only machines, llama.cpp is usually the more practical choice.