AI & Analytics Legends The knowledge platform for SAP Analytics
Academy module

AI Model Evaluation

AI evaluation pipeline: golden set feeds a judged evaluation run, scored on four axes, producing a continuous production outcome — architecture diagram for AI Model Evaluation, Analytics Legends Academy module M057

As of 2026-10-04

Evaluate four things — correctness/faithfulness, relevance, robustness/safety, form — per stratum, never as one score. Run it on the tools SAP ships in the generative AI hub: Evaluations (scenario genai-evaluations, simplified executable genai-evaluations-simplified) with computed metrics (BERT Score, BLEU, ROUGE, Exact Match, JSON Schema Match, Language Match, content-filter booleans) and LLM-as-a-judge metrics (Correctness, Answer Relevance, Instruction Following, Conciseness, RAG Groundedness, RAG Context Relevance); custom judge metrics with explicit rubrics; Prompt Optimization; and Inference Observability to sample production traffic. Test sets are dated claims re-validated by data owners; judges are instruments to calibrate. No invented sizes or thresholds.

What you will learn

  • Design a stratified, dated test set (happy path, edge, access-control, adversarial/out-of-scope) with five fields per entry and a named data owner for re-validation
  • Run an SAP AI Core evaluation (scenario genai-evaluations, executable genai-evaluations-simplified) comparing prompts and models with the right system-defined metrics — knowing which need a reference answer
  • Define a custom LLM-as-a-judge metric with an explicit rubric and examples, and calibrate any judge against human review and planted wrong answers
  • Separate retrieval evaluation (Retrieval API, RAG Context Relevance) from generation evaluation (RAG Groundedness, Correctness) for grounded assistants
  • Instrument production with Inference Observability (metadata vs full persistence, labels, feedback) and hand over a rehearsed evaluation runbook

Module overview

Evaluation is what separates "we shipped AI" from "we shipped AI that works — and can prove it". A prompt, a grounded assistant or an agent without a documented, repeatable evaluation is a risk nobody can size, a quality nobody can defend, and a regression the client will discover from users. This module teaches the discipline and anchors it on the tooling SAP actually ships in the generative AI hub of SAP AI Core: Evaluations, system-defined and custom metrics, Prompt Optimization, and Inference Observability. The method is portable; the tools are what let you run it inside the client's SAP landscape with evidence the client keeps.

Prerequisites

  • Completed or read the hands-on module M325 (orchestration, prompt registry)
  • Review core concepts first: C204, C118, C027

Outcomes

  • Build and version a stratified test set that survives data changes because every entry is dated and owned.
  • Choose SAP system-defined metrics correctly and add custom judge metrics that encode business rules.
  • Diagnose a wrong grounded answer as a retrieval or a generation problem before changing anything.
  • Hand over a rehearsed runbook plus production sampling through Inference Observability.

Full module available to members. The full module adds: the decision framework · the end-to-end scenario walkthrough · the KPI scorecard · the anti-patterns · the code blocks · the knowledge check · the diagrams.

Open in the app →