AI & Analytics Legends The knowledge platform for SAP Analytics
Academy module

Post-Training and Fine-Tuning — SFT, RLHF, Constitutional AI, DPO and LoRA, and When an SAP Team Should Fine-Tune at All

Post-Training and Fine-Tuning — SFT, RLHF, Constitutional AI, DPO and LoRA, and When an SAP Team Should Fine-Tune at All — Analytics Legends section illustration for the SAP Analytics knowledge base (concepts, studies, Academy)

As of 2026-10-04

Explains how models become assistants after pre-training (supervised fine-tuning, RLHF, Constitutional AI and RLAIF, DPO), how LoRA and QLoRA make adaptation affordable, and what SAP documents on SAP AI Core: consumption and orchestration of provider models, grounding, prompt optimisation, the SAP-trained SAP-ABAP-1, in-context prediction with SAP-RPT-1 and the restriction on generating synthetic training data. It then gives a decision framework for RAG versus fine-tuning versus in-context learning, grounded in published evidence (RAG beats fine-tuning for facts; LIMA and LoRA for behaviour) and an honest cost and effort framing in which data, evaluation, serving and base-model retirement dominate. The lab ends in a fine-tune-or-not memo.

What you will learn

  • Place pre-training, supervised fine-tuning, RLHF, Constitutional AI/RLAIF, DPO and continued pre-training in the model lifecycle and say which operation a client request actually describes
  • Explain the three steps of RLHF (SFT, reward model, PPO with a KL penalty) and why reward over-optimisation makes it a model-provider activity
  • Explain Constitutional AI's two phases and RLAIF, and the DPO loss on chosen and rejected pairs against a frozen reference model
  • Explain LoRA and QLoRA (frozen base, low-rank trainable matrices, quantised base) and the operational consequences: small swappable adapters, forgetting and quality limits
  • Read SAP's documentation precisely: what the generative AI hub offers (consumption, orchestration, grounding, prompt optimisation), what SAP fine-tuned itself (SAP-ABAP-1), where in-context learning replaces training (SAP-RPT-1), and the restriction on synthetic training data
  • Apply a decision framework (RAG vs fine-tuning vs in-context learning) with a golden set and an honest cost and effort estimate, and write a fine-tune-or-not memo

Module overview

Who this is for. You can build a grounded assistant on the SAP generative AI hub and you have now been asked the question every SAP AI team hears: "should we fine-tune the model on our own data?" This module explains how models are made helpful and safe after pre-training (supervised fine-tuning, RLHF, Constitutional AI, DPO), how parameter-efficient fine-tuning with LoRA works, and what SAP actually offers on SAP AI Core, so that you can answer with a decision, not a reflex. It assumes M333 (AI and LLM fundamentals) and benefits from M391 (LLM internals), M325 (generative AI hub hands-on) and M365 (LLM cost engineering). It is deliberately sceptical: for most SAP knowledge problems fine-tuning is the wrong tool, and the module shows why with evidence. Primary sources are the original papers and vendor documentation; cost and effort figures that are not published by SAP are labelled as editorial planning estimates.

Prerequisites

  • Completion of M333 (AI & LLM Fundamentals for SAP Consultants) or equivalent working knowledge of tokens, embeddings, grounding and context
  • Recommended: M391 (LLM Internals for SAP Architects) for attention, KV cache and quantisation background
  • Optional: access to an SAP AI Core tenant with the extended plan, or the SAP documentation on generative AI hub, orchestration and grounding, to run the baseline measurements in the lab

Outcomes

  • Say precisely which kind of adaptation a client request describes (style, task, facts, safety) and which method, if any, addresses it.
  • Explain RLHF, Constitutional AI, DPO and LoRA to a technical sponsor in two minutes each, with the right trade-offs.
  • Describe what SAP AI Core documents today for model consumption, SAP-trained models and in-context prediction, without promising a managed fine-tuning service that the reviewed documentation does not describe.
  • Run a golden-set comparison of prompt, few-shot, grounding and (if justified) a LoRA pilot, and defend the result.
  • Write a fine-tune-or-not memo with an honest cost and effort estimate that includes data, evaluation, serving and model retirement.

Full module available to members. The full module adds: the decision framework · the end-to-end scenario walkthrough · the KPI scorecard · the anti-patterns · the code blocks · the knowledge check · the diagrams.

Open in the app →