Constitutional AI + RLAIF — How Anthropic / Claude Removed the Human-in-the-Loop
As of 2026-10-04
What is Constitutional AI + RLAIF?
Constitutional AI replaces RLHF's human labellers with an AI critic scoring against a published, written constitution — solving both the labelling bottleneck and the opacity of unwritten human preferences.
Constitutional AI, introduced by Anthropic in 2022 and refined through every subsequent Claude generation, answers two specific weaknesses in RLHF: the human-labelling bottleneck, and the opacity of preferences that live only in labellers' heads rather than being written down anywhere. It does both by making the standard explicit — a written "constitution," a set of principles the model is trained to follow — and by replacing human comparison-raters with an AI critic, a technique known as Reinforcement Learning from AI Feedback, or RLAIF.
How the two-phase pipeline works
Phase one is supervised constitutional learning. The pre-trained model is shown harmful or borderline prompts, generates an initial response, then is asked to critique that response against the constitution's principles and rewrite it accordingly — identify what is harmful, deceptive, or unfair about this response, then rewrite it to remove the problem while staying as helpful as possible. The resulting prompt-and-revised-response pairs become an SFT dataset, producing a model already shaped by the constitution before any reinforcement learning begins.
Phase two is RLAIF itself. Rather than collecting human rankings of paired responses, an AI critic — often the constitutionally-tuned model itself, or a stronger one — ranks response pairs according to the same constitution. A preference model trains on those AI-generated rankings, and the policy is then optimized against it using the same Proximal Policy Optimization step RLHF uses; the only structural difference from RLHF is that AI judgment replaces human judgment at the comparison stage. Everything downstream of that swap is mechanically identical to Reinforcement Learning from Human Feedback.
Why it matters
- Two-phase pipeline: the model first critiques and revises its own harmful responses against explicit principles, then an AI critic ranks outputs for RLAIF instead of human raters.
- The constitution is public and auditable — drawn from the UN Declaration of Human Rights, Apple's ToS, and Anthropic's own safety principles — a transparency step-change over RLHF's implicit labeller preferences.
- It directly targets RLHF's labelling bottleneck: tens of thousands of comparisons take months to collect at frontier-model quality.
Key points
- Two-phase pipeline — Phase 1 supervised constitutional learning (model self-critiques and revises against principles); Phase 2 RLAIF (AI critic ranks pairs, preference model trained, PPO optimises policy).
- Constitution — written, public, auditable principles (UN UDHR, Apple ToS, non-Western perspectives, Anthropic safety); explicit replaces implicit labeller preferences.
- Scales beyond RLHF — AI critic labels millions of examples for API-call cost, eliminating months-long human-labelling bottleneck.
- Risks — constitution drift (biased principles amplified at scale), AI-critic miscalibration inherited; mitigated by published, iterated constitution.
- Adoption — Anthropic Claude (2022 onward); Meta Llama 3 self-critique steps; OpenAI Deliberative Alignment for o1; now mainstream technique.
- SAP+Anthropic partnership makes this concrete, not hypothetical, for an SAP consultant — Claude is positioned as a primary reasoning/agentic capability across Joule and Joule agents (HR, procurement, supply chain), connected via MCP, reading SAP's Knowledge Graph for business context.
- Anthropic's published constitution is general-purpose and says nothing about SAP-specific process norms (clean-core, HR data masking, ABAP conventions) — it complements but never substitutes for SAP's own runtime governance (OpenShell, Joule Studio governance layer) and human-in-the-loop review (C254).
Terms used on this page
- Constitutional AI (CAI)
- Anthropic's 2022 alignment technique (Bai et al.) that aligns LLMs against an explicit, written set of principles ('constitution') via two phases: supervised self-critique-and-revision, then RLAIF; the technique behind every Claude generation.
- RLAIF
- Reinforcement Learning from AI Feedback — the RLHF variant where an AI critic (not human labellers) ranks output pairs according to documented principles; eliminates the human-labelling bottleneck and enables scale to millions of preference labels.
- Constitution (in CAI)
- The written, public, auditable set of principles that govern an LLM's behaviour under Constitutional AI; Anthropic's draws from UN UDHR, Apple ToS, non-Western perspectives, and Anthropic safety principles; iterated openly across Claude generations.
- Self-critique step
- The supervised constitutional learning move where the model is prompted to identify problems with its own initial response, then revise it according to constitutional principles; the (prompt, revised response) pair becomes SFT data.
- SAP+Anthropic partnership
- SAP's announced partnership positioning Claude as a primary reasoning and agentic capability across Joule and Joule agents (HR, procurement, supply chain), connected via MCP, with Claude reading the SAP Knowledge Graph for business context — announced 2026-05-12, not a blanket GA claim for every use case.
- Constitution scope limitation
- The general-purpose scope of a published constitution (e.g. Anthropic's, drawing on the UN UDHR and general safety principles) does not extend to domain-specific business norms — clean-core principles, HR data-masking expectations, ABAP conventions — which remain SAP's own runtime-governance and human-in-the-loop responsibility, not something a constitution rewrite can substitute for.
Sources
- Bai et al. — Constitutional AI: Harmlessness from AI Feedback (Anthropic 2022)
- Anthropic — Claude's Constitution (public published principles)
- Lee et al. — RLAIF: Scaling Reinforcement Learning from Human Feedback with AI Feedback (Google 2023)
- Anthropic — Collective Constitutional AI (public deliberation on principles)
- SAP Generative AI — official product page
- Model Context Protocol — specification 2026-07-28 (the MCP connection referenced in the SAP+Anthropic Claude-Joule partnership)
- SAP Help Portal — connect to the MCP server for SAP BTP administration (MCP as the integration path for Claude-based agents)
- SAP News Center — SAP + Anthropic, bringing Claude to the SAP Business AI Platform (2026-05-12)
- SAP Community — Why SAP needs a Knowledge Graph, grounding enterprise AI reasoning (Claude/Joule context)
- SAP News Center — Secure AI agents: how SAP and NVIDIA co-define enterprise-grade agent execution (OpenShell, runtime governance distinct from training-time alignment)
- Red Teaming Language Models to Reduce Harms — Ganguli et al., arXiv
- Specific versus General Principles for Constitutional AI — Kundu et al., arXiv
- Constitutional Classifiers: Defending against universal jailbreaks — Anthropic Research
- Self-critiquing models for assisting human evaluators — Saunders et al., arXiv
- Deliberative Alignment: Reasoning Enables Safer Language Models — OpenAI, arXiv
- Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations — Meta, arXiv
Full card available to members. What the full card adds: the full decision framework · the SAP vs Snowflake / Databricks / Fabric comparison · the common pitfalls and their fix · the cheat sheet · the architecture schemas · the code blocks · the facts worth quoting.