📋 KEY INSIGHTS
- RLHF (Reinforcement Learning from Human Feedback) is the technique that transformed pre-trained language models into instruction-following assistants — it is the key step that made ChatGPT, Claude, and Gemini behave helpfully rather than just continuing text patterns from their training corpora.
- RLHF has three stages: supervised fine-tuning (SFT) on demonstration data, reward model training on human preference rankings, and RL optimisation (typically PPO) that steers the LLM to generate outputs the reward model rates highly.
- The reward model is the weakest link in the RLHF pipeline: it is trained on a finite set of human preferences and can be exploited by the LLM through reward hacking — generating outputs that score highly on the reward model without actually being helpful or safe.
- Direct Preference Optimisation (DPO) is a mathematically equivalent but simpler alternative to RLHF that eliminates the explicit reward model and the RL training loop, achieving alignment by directly optimising a classification objective on preference pairs.
- Constitutional AI (Anthropic’s approach) extends RLHF by having the AI critique and revise its own outputs against a written constitution of principles — reducing dependence on expensive human feedback collection for the harmlessness component of alignment.
- RLHF is not a solved problem: reward hacking, distributional shift between RLHF training and deployment, and the difficulty of specifying human values completely remain active research challenges at every major AI lab.
Every large language model that you interact with today — ChatGPT, Claude, Gemini, Llama-3-Instruct — has been transformed from a next-token predictor into a useful assistant by a technique called Reinforcement Learning from Human Feedback. Without RLHF, these models behave like sophisticated autocomplete: they continue text patterns from their training data, sometimes helpfully and often not. With RLHF, they follow instructions, refuse harmful requests, produce structured outputs, and adapt their tone to context. Understanding how RLHF works — and why it was necessary — is increasingly important for data scientists working with LLMs, building AI products, or studying alignment. This guide covers the full RLHF pipeline, its successors (DPO, Constitutional AI, RLAIF), and the unresolved challenges that define current alignment research.
Why Pre-training Alone Is Not Enough
A language model trained purely on next-token prediction learns to imitate the statistical patterns in its training corpus. The internet contains enormous amounts of text, but only a small fraction of it is the kind of helpful, honest, structured response you want from an assistant. Much more of it is casual conversation, opinionated journalism, persuasive marketing, factually incorrect claims, and content that would be harmful to reproduce on request. A raw pre-trained model reflects all of these patterns simultaneously — it has no preference for being helpful over being engaging, no preference for accuracy over plausibility, and no preference for safety over compliance.
The pre-training objective — predict the next token — also provides no guidance about what to do when a prompt is ambiguous (should I interpret this charitably or literally?), when a request is harmful (should I refuse or comply?), or when the model is uncertain (should I acknowledge uncertainty or generate a confident-sounding answer?). These are value judgements that cannot be resolved from the text distribution alone. RLHF is the mechanism by which human preferences on these judgements are translated into model behaviour.
The Three-Stage RLHF Pipeline
Stage 1 — Supervised Fine-Tuning (SFT): Human labellers write example (prompt, ideal response) pairs covering a wide range of tasks — answering questions, writing code, summarising documents, following multi-step instructions, refusing harmful requests with explanation. The pre-trained model is then fine-tuned on these demonstrations using standard supervised learning (cross-entropy loss on the ideal responses). SFT teaches the model the format and style of helpful responses and provides a much better starting point for RL than the raw pre-trained model. However, SFT is limited by the volume and quality of human-written demonstrations — it is expensive to produce at scale, and humans are better at judging quality than producing it (they can identify a great response when they see one, but writing it from scratch is harder).
Stage 2 — Reward Model Training: Rather than asking humans to write ideal responses (which is expensive and slow), human labellers are instead asked to rank or compare pairs of model outputs for the same prompt. Which of these two responses is more helpful? More accurate? Safer? Ranking is significantly faster than writing — a labeller can evaluate 10–20 response pairs in the time it takes to write one ideal response. A reward model (a separate, usually smaller LM with a scalar output head) is trained on these preference labels: given a (prompt, response) pair, it learns to predict the human preference score. After training, the reward model can evaluate any model output instantly without human involvement, enabling automated feedback at the scale required for RL training.
Stage 3 — RL Optimisation with PPO: The SFT model is used as the starting point for RL training. The RL agent (the LLM being optimised) generates responses to prompts, and the reward model scores each response. PPO (Proximal Policy Optimisation) — a policy gradient algorithm — updates the LLM’s weights to increase the probability of generating responses that receive high reward model scores. A critical component is a KL divergence penalty that constrains how far the RL-optimised model drifts from the SFT model — without this constraint, the model quickly learns to exploit the reward model in ways that maximise its score while producing incoherent or useless text (reward hacking). The KL penalty ensures the model stays within the distribution of sensible text while being steered toward higher-quality outputs.
| Stage | Input | Output | Human Involvement | Key Challenge |
|---|---|---|---|---|
| Pre-training | Trillion-token text corpus | Base LM (next-token predictor) | None (automated) | Scale and compute cost |
| SFT | Human-written (prompt, response) pairs | SFT model (instruction-following baseline) | High (writing demonstrations) | Volume and quality of demos |
| Reward Model Training | Human preference rankings of response pairs | Reward model (scalar quality scorer) | Medium (ranking pairs) | Reward model generalisation |
| PPO / RL Optimisation | SFT model + reward model | RLHF model (aligned assistant) | None (automated) | Reward hacking, training stability |
DPO — Direct Preference Optimisation
Direct Preference Optimisation (Rafailov et al., 2023) showed that the RLHF objective — maximise reward while staying close to the SFT policy — can be rewritten as a simple binary classification objective on preference data, eliminating both the explicit reward model and the RL training loop. DPO directly fine-tunes the LLM on (prompt, chosen response, rejected response) triples, optimising to increase the log-probability of the chosen response relative to the rejected response in a way that implicitly encodes the preference.
The advantages of DPO over PPO-based RLHF are substantial. Training is simpler — no separate reward model to train and maintain, no RL hyperparameters (learning rate, KL coefficient, clipping ratio) to tune, and no complex multi-model infrastructure. Training is also more stable — PPO RLHF is notoriously sensitive to hyperparameter settings, and getting the KL penalty right requires careful calibration. DPO achieves competitive or superior performance on alignment benchmarks while being significantly cheaper and easier to implement. Most open-source LLM fine-tuning pipelines (Axolotl, TRL by Hugging Face) now support DPO as a first-class training mode alongside SFT.
The limitation of DPO is that it still requires preference data — (chosen, rejected) pairs annotated by humans or by a stronger AI judge. The data collection bottleneck is the same as in RLHF. RLAIF (RL from AI Feedback) addresses this by using a stronger LLM (GPT-4, Claude Opus) to generate the preference labels instead of human annotators — dramatically reducing the cost of alignment data while achieving quality that approaches human labelling on many tasks.
Constitutional AI and Harmlessness Training
Anthropic’s Constitutional AI (CAI) approach, introduced in 2022, addresses a specific limitation of standard RLHF: the difficulty and expense of collecting human feedback for harmlessness. Training a model to be helpful is relatively straightforward — humans can easily judge whether a coding answer is correct or a summary is accurate. Training a model to refuse harmful requests, handle edge cases appropriately, and be consistently honest is much harder: the space of possible harmful requests is enormous, and human labellers may disagree, have inconsistent standards, or be exposed to distressing content during labelling.
Constitutional AI has two phases. In the Supervised AI (SAI) phase, the model is asked to critique and revise its own outputs against a written constitution — a set of principles like “choose the response that is least likely to contain harmful or unethical content” or “prefer the response that is most honest, even if it is less immediately flattering.” The model generates an initial response, is given the constitution and asked to identify which principle its response violates, and then generates a revised response. This creates (original, revised) pairs that are used for SFT. In the RL from AI Feedback (RLAIF) phase, the model is used to generate preference labels on (chosen, rejected) pairs by asking which response better satisfies the constitutional principles — replacing human preference labellers with the model itself. This makes harmlessness training scalable without exposing human labellers to harmful content.
| Technique | Feedback Source | Reward Model? | RL Loop? | Main Advantage |
|---|---|---|---|---|
| RLHF (standard) | Human labellers | Yes (explicit) | Yes (PPO) | Human preferences directly encoded |
| DPO | Human labellers | No (implicit) | No | Simpler, more stable training |
| RLAIF | AI judge (stronger LLM) | Yes (AI-scored) | Yes | Scalable without human labelling cost |
| Constitutional AI | AI self-critique + constitution | Yes (AI-scored) | Yes | Scalable harmlessness training |
| SPIN (Self-Play) | Comparisons to earlier self | No | No | No external feedback required |
| KTO | Binary thumbs up/down | No | No | Works with unpaired feedback data |
✦ SUMMARIZE THIS ARTICLE WITH AI
The foundational RL concepts — policy gradient, reward functions, Markov Decision Processes — that underpin PPO and RLHF are explained in our Reinforcement Learning guide. The Transformer architecture and pre-training process that produces the base models that RLHF is applied to are covered in our Transformers and Attention guide. The broader generative AI landscape — RAG, fine-tuning, prompt engineering — that data scientists use alongside RLHF-aligned models is in our Generative AI guide. NLP interview questions covering BERT, fine-tuning, and alignment are in our NLP Interview Q&A.



