GRPO vs PPO is mostly a question of what each method uses as a baseline. PPO (Proximal Policy Optimization) trains a separate value model, often called the critic, to judge how good each response is expected to be. GRPO (Group Relative Policy Optimization) drops that critic. It samples a group of answers to the same prompt and scores each one against the group's average. That makes GRPO simpler and lighter on memory, which is why it became the default for training reasoning models on tasks with checkable answers, such as maths and code. PPO is still a strong choice when rewards come from a learned reward model and you can afford the extra model.

GRPO vs PPO at a glance

PPOGRPO
Introduced2017, for general reinforcement learning2024, in the DeepSeekMath paper
Baseline for advantageA learned value model (critic)The average reward of a group of sampled answers
Models kept in memoryPolicy, value model, usually a reference model and a reward modelPolicy, plus an optional reference model; rewards can be plain functions
Samples per promptOften oneSeveral, eight by default in TRL
Best suited toLearned reward models, dense or subjective rewardsVerifiable rewards: maths, code, formats, tool outcomes
Main costsMemory and tuning of the criticGeneration time, since every prompt needs a group of answers

How PPO works for language models

PPO was introduced in the 2017 paper Proximal Policy Optimization Algorithms and powered early reinforcement learning from human feedback, including OpenAI's InstructGPT work. The loop looks like this:

  1. The model, called the policy, writes a response to a prompt.
  2. A reward model scores the response.
  3. A value model estimates how much reward the policy should have expected at each token. The difference between what happened and what was expected is the "advantage."
  4. The policy is nudged toward responses with positive advantage, with a clipping rule that stops any single update from moving it too far.
  5. A penalty keeps the policy close to a frozen reference copy, so it does not drift into gibberish that happens to score well.

The catch is the value model. The DeepSeekMath paper points out that in PPO the value function is typically another model of comparable size to the policy, which brings a substantial memory and compute burden. A critic is also one more network to tune, and for language the reward usually arrives only at the end of a response.

How GRPO works

GRPO was introduced in DeepSeekMath in February 2024 as a variant of PPO. In the authors' words, it "foregoes the critic model, instead estimating the baseline from group scores." DeepSeek later used it to train its R1 reasoning model; you can find DeepSeek-R1 in our directory.

The TRL documentation breaks each training step into four parts:

  1. Generate a group. For each prompt, sample several completions.
  2. Score and compare. Compute a reward for each completion, then subtract the group's mean reward and, by default, divide by the group's standard deviation. A completion that beats its siblings gets a positive advantage; a worse one gets a negative advantage.
  3. Optionally penalise drift. Estimate the KL divergence from a reference model.
  4. Update. Increase the probability of tokens in above-average completions and decrease it for below-average ones.

Because the comparison happens inside the group, GRPO needs no critic at all, and the reward can be a simple Python function: did the final answer match, did the code pass the tests, did the output parse as JSON.

What changed since the original GRPO

Most explainers stop at the 2024 formula, but practice has moved on. TRL's defaults reflect that:

  • No KL penalty by default. TRL sets the KL coefficient, called beta, to zero by default, so the reference model is not even loaded, which saves memory. Its docs cite studies such as Open-Reasoner-Zero finding the KL term is not essential for GRPO, and note that DAPO and Dr. GRPO leave it out too. The DeepSeek-R1 paper used a small value of 0.001.
  • Length-bias fixes. The original loss averaged per response, which subtly favoured short correct answers and long wrong ones. The DAPO paper switched to token-level averaging, and Dr. GRPO went further by dividing by a constant. TRL now uses the DAPO-style loss by default and offers Dr. GRPO and others as options.
  • Optional reward scaling. The Dr. GRPO authors argue that dividing by the group's standard deviation biases training by question difficulty; TRL lets you turn that scaling off.

Other GRPO defaults in TRL's GRPOConfig at the time of writing: eight generations per prompt, a learning rate of 1e-6, a clipping range of 0.2, and a maximum completion length of 512 tokens. Check the config source for your installed version, as these evolve.

Where DPO fits

DPO, or Direct Preference Optimization, is not online reinforcement learning at all. It learns directly from a fixed dataset of preferred and rejected answer pairs, with no sampling during training and no reward model. It is cheaper and more stable than either PPO or GRPO, and TRL notes it was used to post-train Llama 3. The trade-off is that the model only learns from the examples you collected, rather than exploring and being graded on its own attempts.

A common recipe is supervised fine-tuning first, then DPO for tone and preferences, and GRPO where there is a verifiable target.

A minimal GRPO run

The TRL README shows how little code a GRPO run needs:

from datasets import load_dataset
from trl import GRPOTrainer
from trl.rewards import accuracy_reward

dataset = load_dataset("trl-lib/DeepMath-103K", split="train")
trainer = GRPOTrainer(
    model="Qwen/Qwen2.5-0.5B-Instruct",
    reward_funcs=accuracy_reward,
    train_dataset=dataset,
)
trainer.train()

A custom reward function simply takes the completions and returns a list of floats, one per completion. On a single GPU, Unsloth publishes free GRPO notebooks, including ones for gpt-oss-20b and Qwen3, and Axolotl added asynchronous GRPO in 2026. For fast generation during training, TRL can use vLLM inside the trainer process.

Pros and cons

PPO

  • Pros: well studied; works with learned reward models and dense rewards; the critic can give finer-grained credit to individual tokens.
  • Cons: an extra model of similar size to train and store; more hyperparameters; harder to get stable.

GRPO

  • Pros: no critic, so less memory and simpler code; pairs naturally with programmatic rewards; strong track record on reasoning tasks; widely supported in TRL, Unsloth and Axolotl.
  • Cons: each prompt needs several generations, so rollouts often dominate training time; if every answer in a group gets the same reward, that prompt teaches nothing; vulnerable to reward hacking if the reward function has loopholes.

Which should you choose?

  1. You can check answers automatically, for maths, code, structured output or tool success: use GRPO.
  2. Your reward comes from a learned reward model of human preferences and you have the compute: PPO remains a solid option.
  3. You have pairs of good and bad answers but no way to score new ones: use DPO.
  4. You are just starting: do supervised fine-tuning first, then add GRPO on a small model. Our guide on how to fine-tune an LLM locally covers the first step, and LoRA vs QLoRA helps you fit it in memory.

Who this is for

ML engineers and researchers moving beyond supervised fine-tuning, and anyone trying to understand how reasoning models like DeepSeek-R1 were trained. For the underlying papers, arXiv hosts all of them.

FAQ

What is the main difference between GRPO and PPO?

PPO uses a learned value model, the critic, to estimate a baseline for each response. GRPO removes the critic and instead compares each response with the average reward of a group of responses to the same prompt.

Why is GRPO more memory-efficient than PPO?

Because it does not train a value model, which in PPO is typically about as large as the policy. With the KL penalty turned off, as in TRL's default, it also skips loading a reference model.

Is GRPO better than PPO?

Not universally. GRPO is simpler and has worked very well for reasoning tasks with verifiable rewards. PPO can still be the better fit for learned or dense rewards, where a critic helps assign credit.

What is the difference between GRPO and DPO?

GRPO is online reinforcement learning: the model generates answers during training and is scored by a reward function. DPO is offline: it learns from a fixed dataset of preferred and rejected answers, with no generation during training.

How many generations per prompt does GRPO use?

TRL's default is eight. More generations give a better baseline but cost more time; the effective batch size must be divisible by the number of generations.

Can I run GRPO on a consumer GPU?

Yes, for small models. Unsloth publishes free notebooks for GRPO on models such as Qwen3 4B and gpt-oss-20b, and using LoRA keeps memory down.