Trainingadvanced7 min read

Reinforcement Learning from Human Feedback (RLHF)

Aligns a pretrained model to human preferences by training a reward model on comparisons, then optimizing the model against it.

Preparing spaced-repetition scheduler…

Pretraining on next-token prediction produces a model that can continue text well, but does not necessarily follow instructions, refuse harmful requests, or behave helpfully. RLHF and related preference-optimization methods are widely used to close that gap.

A common preference-training pipeline:

  1. Start from a pretrained model and instruction-tune it on demonstrations. This supervised step gives the preference-training pipeline a useful starting policy.
  1. Train a reward model (RM). Humans are shown pairs of model outputs to the same prompt and asked which is better. The RM is a separate model trained to predict, given a prompt and a response, the score a human would assign. It is a scalar model: a regression head over a transformer.
  1. Optimize the policy against the RM. Use a reinforcement-learning algorithm such as PPO to increase predicted reward. DPO is a separate preference-optimization approach that does not require training an explicit reward model or running this RL loop. A KL penalty discourages excessive drift from a reference policy; it does not eliminate reward hacking.

Why it works: the RM captures a fuzzy notion of "helpfulness" that is hard to specify with rules. By training against thousands of comparisons, the policy inherits that notion.

Failure modes:

  • Reward hacking: the policy finds outputs the RM scores highly but humans do not actually prefer.
  • Sycophancy: the model learns to agree with the user rather than tell the truth.
  • Over-refusal: the model becomes too cautious on benign prompts.

Constitutional AI (Anthropic) replaces the human ranker with a principles-based critique loop: the model critiques its own outputs against a written constitution. This is RLAIF (RL from AI feedback).

RLHF pipeline. After pretraining, supervised demonstrations initialize the policy. Human comparisons train a reward model; an RL stage such as PPO then optimizes the policy. DPO follows a different path.

Key takeaways

  • A common RLHF pipeline uses supervised fine-tuning, preference reward modeling, and policy optimization.
  • A KL penalty keeps the aligned model from drifting too far from the base.
  • Reward hacking and sycophancy are the main failure modes.

Self-check

  1. What does a reward model in RLHF actually predict?

  2. Why is a KL penalty added during the policy optimization step?

References

  1. Direct Preference Optimization: Your Language Model is Secretly a Reward Model — Rafailov et al. / arXiv

    Direct preference optimization without an explicit reward-model training stage.

  2. Learning to summarize from human feedback (Stiennon et al.) — OpenAI / arXiv

    Foundational RLHF paper.

  3. Learning to summarize with human feedback — OpenAI

    Plain-language overview of RLHF.

  4. Constitutional AI: Harmlessness from AI Feedback — Anthropic

    RLHF alternative grounded in explicit principles.

Last verified 2026-08-19.