AI Model Evaluation
As of 2026-10-04
Evaluate four things — correctness/faithfulness, relevance, robustness/safety, form — per stratum, never as one score. Run it on the tools SAP ships in the generative AI hub: Evaluations (scenario genai-evaluations, simplified executable genai-evaluations-simplified) with computed metrics (BERT Score, BLEU, ROUGE, Exact Match, JSON Schema Match, Language Match, content-filter booleans) and LLM-as-a-judge metrics (Correctness, Answer Relevance, Instruction Following, Conciseness, RAG Groundedness, RAG Context Relevance); custom judge metrics with explicit rubrics; Prompt Optimization; and Inference Observability to sample production traffic. Test sets are dated claims re-validated by data owners; judges are instruments to calibrate. No invented sizes or thresholds.
What you will learn
- Design a stratified, dated test set (happy path, edge, access-control, adversarial/out-of-scope) with five fields per entry and a named data owner for re-validation
- Run an SAP AI Core evaluation (scenario genai-evaluations, executable genai-evaluations-simplified) comparing prompts and models with the right system-defined metrics — knowing which need a reference answer
- Define a custom LLM-as-a-judge metric with an explicit rubric and examples, and calibrate any judge against human review and planted wrong answers
- Separate retrieval evaluation (Retrieval API, RAG Context Relevance) from generation evaluation (RAG Groundedness, Correctness) for grounded assistants
- Instrument production with Inference Observability (metadata vs full persistence, labels, feedback) and hand over a rehearsed evaluation runbook
Module overview
Evaluation is what separates "we shipped AI" from "we shipped AI that works — and can prove it". A prompt, a grounded assistant or an agent without a documented, repeatable evaluation is a risk nobody can size, a quality nobody can defend, and a regression the client will discover from users. This module teaches the discipline and anchors it on the tooling SAP actually ships in the generative AI hub of SAP AI Core: Evaluations, system-defined and custom metrics, Prompt Optimization, and Inference Observability. The method is portable; the tools are what let you run it inside the client's SAP landscape with evidence the client keeps.
Prerequisites
- Completed or read the hands-on module M325 (orchestration, prompt registry)
- Review core concepts first: C204, C118, C027
Outcomes
- Build and version a stratified test set that survives data changes because every entry is dated and owned.
- Choose SAP system-defined metrics correctly and add custom judge metrics that encode business rules.
- Diagnose a wrong grounded answer as a retrieval or a generation problem before changing anything.
- Hand over a rehearsed runbook plus production sampling through Inference Observability.
Full module available to members. The full module adds: the decision framework · the end-to-end scenario walkthrough · the KPI scorecard · the anti-patterns · the code blocks · the knowledge check · the diagrams.