Speculative decoding vs MTP is not really a contest, because multi-token prediction is one way to do speculative decoding. Speculative decoding is the general trick: something fast guesses the next few tokens, and the big model checks all of the guesses in a single pass, keeping the ones it agrees with. MTP (multi-token prediction) is a training technique that gives a model extra heads to predict tokens further ahead. When a model ships with those heads, they can serve as a built-in drafter, so you get speculative decoding without downloading a separate draft model. Either way, done correctly, the output is the same as the big model would have produced on its own; only the speed changes.

Why generation is slow, and why guessing helps

A language model writes one token at a time, and each token needs a full pass through the model. On a single GPU or a Mac, that pass is limited mostly by how fast the weights can be read from memory, not by raw compute. Checking several tokens at once costs about the same as generating one, because the weights only have to be read once. The llama.cpp speculative decoding docs put it simply: computing several tokens in a batch is more efficient than computing them one after another.

So if a cheap drafter can guess the next four tokens and the big model agrees with three of them, you get three or four tokens for roughly the price of one pass. If the guesses are wrong, the big model discards them and you lose a little time.

Is speculative decoding lossless?

Yes, when implemented as designed. The original paper, Fast Inference from Transformers via Speculative Decoding from 2022, describes it as sampling from large models faster "without any changes to the outputs," and reports a 2 to 3 times speed-up on T5-XXL. DeepMind's parallel work on speculative sampling uses a modified rejection-sampling step that preserves the target model's distribution, and reports a 2 to 2.5 times speed-up on a 70-billion-parameter model. With greedy decoding, the final text is identical; with sampling, the statistics are the same.

Some servers offer optional "relaxed" acceptance thresholds that trade exactness for speed. Leave those at their defaults if you want truly lossless output.

The drafting methods compared

MethodWhere the guesses come fromExtra downloadNotes
Draft modelA smaller model with the same vocabulary, such as a 1B siblingYesSimplest to understand; widely supported
EAGLE-3A one-layer drafter that reads the big model's hidden statesYes, trained per target modelHigher acceptance than a plain draft model of the same size
MTPExtra prediction heads trained into the model itselfNo, if the model ships themOnly for models trained with MTP
DFlash and DSparkA small block-diffusion drafter that proposes a whole block at onceYes, trained per target modelNewest approach; drafting is parallel rather than token by token
N-gram or suffix matchingRepeated patterns already in the prompt or outputNoGreat for code edits and summarising documents you pasted in

How MTP works

Meta's 2024 paper Better and Faster Large Language Models via Multi-token Prediction trained models to predict several future tokens at once, using independent output heads on a shared trunk. The authors found it improved sample efficiency, especially for code, and noted that models trained with four-token prediction were up to three times faster at inference when the extra heads were used for self-speculative decoding.

DeepSeek-V3 made the idea mainstream. Its technical report describes MTP modules used mainly to improve training, which can be discarded at inference or "repurposed" for speculative decoding. DeepSeek reported that the second predicted token was accepted 85 to 90 percent of the time, giving about 1.8 times the tokens per second.

Several open-weight families now ship MTP layers. You can see them in each model's config.json on Hugging Face: DeepSeek-V3, GLM-4.5 and Xiaomi's MiMo list num_nextn_predict_layers, and recent Qwen 3.5 models list mtp_num_hidden_layers. Note that GGUF conversions and quantized uploads do not always keep these layers, so check the model card.

How DFlash fits in

The rising query "mtp vs dflash" comes from a 2026 paper, DFlash: Block Diffusion for Flash Speculative Decoding, from UC San Diego researchers. Instead of drafting one token at a time, a small diffusion model proposes a whole block in a single pass, conditioned on the big model's hidden features. The authors report over six times lossless acceleration and up to 2.5 times the speed-up of EAGLE-3 on their tests; treat those as the authors' figures. The trade-off is that you need a DFlash drafter trained for your exact model, while MTP needs nothing extra if your model already has it.

How to turn it on

llama.cpp

llama.cpp's server selects methods with --spec-type, and you can combine a model-based method with an n-gram one:

# Use the model's own MTP heads
llama-server -m model.gguf --spec-type draft-mtp

# Use a separate small draft model
llama-server -m big.gguf -md small.gguf --spec-type draft-simple

# No extra model: n-gram matching, useful for code refactoring
llama-server -m model.gguf --spec-type ngram-simple

Other types include draft-eagle3, draft-dflash and draft-dspark. The number of drafted tokens is set with --spec-draft-n-max, which defaults to 3.

vLLM and SGLang

vLLM takes a JSON --speculative-config. For a model with native MTP, its speculative decoding docs show:

vllm serve XiaomiMiMo/MiMo-7B-Base --speculative-config '{"method":"mtp","num_speculative_tokens":1}'

vLLM's documentation rates EAGLE and MTP as high-gain methods at low traffic, and suggests starting with one speculative token for MTP. SGLang uses --speculative-algorithm with options including EAGLE, EAGLE3, NEXTN (its name for MTP), STANDALONE and NGRAM.

LM Studio

LM Studio supports the draft-model approach in its chat sidebar. Its documentation explains that the draft model must share the main model's vocabulary, and suggests pairings such as Llama 3.1 8B with Llama 3.2 1B.

When it helps and when it does not

  • Helps most: one user at a time, memory-bound hardware, predictable text such as code, structured output, or rewriting a document you supplied.
  • Helps less: creative writing at high temperature, where guesses are rejected more often, and busy servers with large batches, where the GPU is already doing useful parallel work.
  • Costs memory: a separate drafter needs its own weights and KV cache; see our VRAM guide and KV cache quantization for ways to make room.

Pros and cons

  • Pros: faster output with identical results; no retraining needed for draft-model or n-gram methods; MTP adds speed with no extra download.
  • Cons: gains vary a lot by prompt and model; drafters must match the target's vocabulary or be trained for it; extra memory; less useful at high batch sizes.

Which should you use?

  1. Your model ships MTP layers and your runtime supports them: start with MTP.
  2. An EAGLE-3 or DFlash drafter exists for your exact model: try it, and measure.
  3. Your model has a small sibling with the same tokenizer: use the draft-model approach.
  4. You mostly edit code or long documents: add n-gram drafting; it is nearly free.

For choosing the runtime itself, see Ollama vs llama.cpp vs LM Studio and vLLM vs llama.cpp.

Who this is for

People running open-weight models locally who want faster responses, and engineers tuning a serving stack for latency.

FAQ

What is the difference between speculative decoding and MTP?

Speculative decoding is the general method of drafting tokens cheaply and verifying them with the main model. MTP is a way of training a model with extra heads that predict tokens further ahead; those heads can act as the drafter, making MTP one form of speculative decoding.

Does speculative decoding change the model's answers?

No, not with standard settings. The verification step keeps only tokens the main model would have produced, so greedy output is identical and sampled output has the same distribution.

How much faster is speculative decoding?

It depends on how often the guesses are accepted. Papers report around 2 to 3 times on their test setups, and DeepSeek reported about 1.8 times the tokens per second from MTP on DeepSeek-V3. Measure on your own prompts.

Which models support MTP?

Models trained with MTP layers, including DeepSeek-V3, GLM-4.5, Xiaomi MiMo and recent Qwen 3.5 releases. Check the model's config file and your runtime's documentation, because support varies.

What is DFlash?

DFlash is a 2026 speculative decoding method that uses a small block-diffusion model to draft a whole block of tokens in one pass. It needs a drafter trained for the specific target model, and it is supported in llama.cpp and SGLang.

Do I need a GPU for speculative decoding?

No. llama.cpp supports it on CPUs and Macs as well, and memory-bound hardware is exactly where it tends to help.