RLHF — Reinforcement Learning from Human Feedback
As of 2026-10-04
What is RLHF?
RLHF fixes what SFT can't — it trains a separate reward model on human preference rankings, then uses PPO to optimize the policy against it, generalizing quality judgments beyond the demonstrations.
What it is
Reinforcement Learning from Human Feedback, RLHF, is the training phase that took models capable of following instructions and made them consistently prefer better answers over merely adequate ones — the difference between a model that produces an acceptable response and one that reliably produces the best response it is capable of. It sits after pre-training and Supervised Fine-Tuning in the canonical training stack, and it is the phase most directly responsible for the jump in usefulness that separated early instruction-tuned models from the assistants that followed.
Why it matters
- The three-stage pipeline (SFT → reward model on 6B-13B params → PPO with KL penalty) is why GPT-3.5/4, Claude, and Llama 2-Chat generalize preferences to prompts labellers never saw.
- Real failure modes are catalogued, not theoretical: reward hacking produces sycophancy and hedging, mode collapse narrows output diversity, and collecting tens of thousands of comparison labels is expensive and slow.
- PPO's four-model orchestration (policy, value, reward, reference) is notoriously unstable — small hyperparameter changes destabilize training, the direct motivation for DPO (C256).
Key points
- Three-step pipeline — SFT → train reward model from human pairwise preferences → PPO-optimise policy against reward model + KL penalty to SFT.
- Reward model — separate 6B-13B network initialised from SFT, predicts which of two outputs a human would prefer; the scalable proxy for human judgement.
- Key result — 1.3B RLHF-tuned model preferred over 175B raw GPT-3 (InstructGPT 2022); the proof alignment outperforms scale on usefulness.
- Failure modes — reward hacking (sycophancy, verbosity, over-refusal), mode collapse (reduced diversity), PPO instability, four-model orchestration cost.
- Successors — Constitutional AI / RLAIF (C255) replaces human raters with AI critic; DPO (C256) eliminates the reward model entirely.
- RLHF happens entirely upstream at the foundation-model vendor — an SAP consultant's actual task is due diligence on it, not implementation; ask where the human-in-the-loop checkpoint sits for a specific agent, not just whether the base model was RLHF-trained.
- SAP's runtime governance (NVIDIA OpenShell, Joule Studio's runtime governance layer) constrains what an already-RLHF-trained model is allowed to DO inside an agent loop — it does not retrain or re-align the model itself; the two are separate, complementary control layers, not substitutes for each other.
Terms used on this page
- RLHF
- Reinforcement Learning from Human Feedback — the three-step alignment recipe (SFT, reward modelling, PPO optimisation) introduced operationally by InstructGPT 2022; the technique behind ChatGPT, GPT-4, original Claude, Llama 2-Chat.
- Reward model
- A separate neural network (typically 6B-13B parameters, initialised from the SFT model) trained on human pairwise preference labels to predict which of two candidate outputs a human would prefer; the scalable proxy for human judgement in RLHF.
- Proximal Policy Optimisation (PPO)
- The reinforcement-learning algorithm (Schulman et al. 2017, OpenAI) used in the canonical RLHF pipeline; clips policy-gradient updates to prevent destabilising large policy changes; orchestrates four models in memory simultaneously.
- Reward hacking
- The pathology where an RL-trained policy finds ways to score highly on the reward model that humans would actually dislike (sycophancy, excessive verbosity, hedging, over-refusal); the dominant failure mode of RLHF in practice.
- Human-in-the-loop (SAP)
- SAP's runtime analogue to RLHF's core insight — a structured checkpoint where a human accepts, adjusts or rejects an AI-proposed action before it takes effect; the Accounting Accruals Agent (journal-entry proposals reviewed by an accountant) is the shipped example. Distinct from RLHF: no reward model is trained from these decisions in the SAP product today.
- Runtime guardrails vs training-time alignment
- Two separate control layers that are often conflated: RLHF (or Constitutional AI) shapes a model's weights at training time, upstream at the vendor; SAP's NVIDIA OpenShell and Joule Studio's runtime governance layer constrain what an already-trained model is allowed to DO inside a governed agent loop. Neither substitutes for the other.
Sources
- Ouyang et al. — Training language models to follow instructions with human feedback (InstructGPT, OpenAI 2022)
- Christiano et al. — Deep Reinforcement Learning from Human Preferences (foundational RLHF, 2017)
- Schulman et al. — Proximal Policy Optimization Algorithms (PPO, OpenAI 2017)
- HuggingFace — Illustrating Reinforcement Learning from Human Feedback (technical blog)
- SAP Help Portal — generative AI hub orchestration service (content filtering, masking — runtime governance distinct from training-time RLHF)
- SAP Community — Knowledge graphs for LLM grounding and avoiding hallucination (contrast with RLHF as a hallucination mitigation)
- SAP Help Portal — SAP Accounting Accruals Agent, human-in-the-loop review step
- SAP Community — Automating month-end accruals with the SAP Accounting Accruals Agent
- SAP News Center — NVIDIA OpenShell, secure AI agent execution co-developed with SAP (2026-05-12)
- arXiv — Direct Preference Optimization (2305.18290), cited as RLHF's simplification successor for cross-reference
- Learning to summarize from human feedback — Stiennon et al., arXiv
- Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback — Bai et al., arXiv
- Scaling Laws for Reward Model Overoptimization — Gao et al., arXiv
- Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback — Casper et al., arXiv
- Llama 2: Open Foundation and Fine-Tuned Chat Models — Touvron et al., arXiv
- RewardBench: Evaluating Reward Models for Language Modeling — Lambert et al., arXiv
Full card available to members. What the full card adds: the full decision framework · the SAP vs Snowflake / Databricks / Fabric comparison · the common pitfalls and their fix · the cheat sheet · the architecture schemas · the code blocks · the facts worth quoting.