Reinforcement Learning from Human Feedback (RLHF)
Aligns a pretrained model to human preferences by training a reward model on comparisons, then optimizing the model against it.
Preparing spaced-repetition scheduler…
Pretraining on next-token prediction produces a model that can continue text well, but does not necessarily follow instructions, refuse harmful requests, or behave helpfully. RLHF and related preference-optimization methods are widely used to close that gap.
A common preference-training pipeline:
- Start from a pretrained model and instruction-tune it on demonstrations. This supervised step gives the preference-training pipeline a useful starting policy.
- Train a reward model (RM). Humans are shown pairs of model outputs to the same prompt and asked which is better. The RM is a separate model trained to predict, given a prompt and a response, the score a human would assign. It is a scalar model: a regression head over a transformer.
- Optimize the policy against the RM. Use a reinforcement-learning algorithm such as PPO to increase predicted reward. DPO is a separate preference-optimization approach that does not require training an explicit reward model or running this RL loop. A KL penalty discourages excessive drift from a reference policy; it does not eliminate reward hacking.
Why it works: the RM captures a fuzzy notion of "helpfulness" that is hard to specify with rules. By training against thousands of comparisons, the policy inherits that notion.
Failure modes:
- Reward hacking: the policy finds outputs the RM scores highly but humans do not actually prefer.
- Sycophancy: the model learns to agree with the user rather than tell the truth.
- Over-refusal: the model becomes too cautious on benign prompts.
Constitutional AI (Anthropic) replaces the human ranker with a principles-based critique loop: the model critiques its own outputs against a written constitution. This is RLAIF (RL from AI feedback).
Key takeaways
- A common RLHF pipeline uses supervised fine-tuning, preference reward modeling, and policy optimization.
- A KL penalty keeps the aligned model from drifting too far from the base.
- Reward hacking and sycophancy are the main failure modes.
Self-check
What does a reward model in RLHF actually predict?
Why is a KL penalty added during the policy optimization step?
References
- Direct Preference Optimization: Your Language Model is Secretly a Reward Model
Direct preference optimization without an explicit reward-model training stage.
- Learning to summarize from human feedback (Stiennon et al.)
Foundational RLHF paper.
- Learning to summarize with human feedback
Plain-language overview of RLHF.
- Constitutional AI: Harmlessness from AI Feedback
RLHF alternative grounded in explicit principles.
Last verified 2026-08-19.