Choosing the Model Behind Your SAP AI — Claude, GPT, Gemini, Llama, Mistral and SAP's Own Models in the Generative AI Hub
As of 2026-10-04
Teaches the model-selection decision for SAP AI as an evidence loop: who chooses the model (SAP in Joule, you in the generative AI hub), what the hub verifiably offers today (Claude, GPT, Gemini, Mistral, SAP-RPT; Llama unverified), why vendor-current is not hub-current, how reasoning, caching and fallbacks work in Orchestration V2, why benchmarks only shortlist, a reproducible evaluation lab on masked own data, cost per correct task in capacity units (kept apart from AI Units), residency limits and the model lifecycle. Every model name is dated and marked verified, announced-only or unverified.
What you will learn
- Explain who chooses the model in SAP-delivered features (Joule) versus custom-built AI in the generative AI hub, and state only what SAP has published
- Build a dated model roster from the hub's discovery endpoint and the What's New log, marking every model verified, announced-only or unverified
- Configure reasoning (reasoning_effort, reasoning_content replay), prompt caching and fallbacks in Orchestration V2 and name which models each applies to
- Run a reproducible evaluation on masked own data with the hub's Evaluations capability, pinned versions, repeats and a validated judge
- Compute cost per task in GenAI tokens and capacity units, keep AI Units and vendor list prices separate, and rank models by cost per correct task
- Design for residency and lifecycle: operator, region and sovereign limits, model restriction list, version pinning versus latest, and retirement-date alerts
Who this is for. You already know how to call the generative AI hub (M325), you understand tokens, context windows and grounding (M333), and a client has now asked the question that every SAP AI project reaches within a month: "Which model should we use — Claude, GPT, Gemini, Mistral, Llama, or one of SAP's own?" This module treats it as an architecture decision with evidence. Everything below was checked on 4 October 2026 against SAP's AI Core documentation (including the "What's New for SAP AI Core" log, latest revision 7 September 2026), SAP News, SAP's AI pricing page and the vendors' own documentation. Where a fact lives only in a login-protected SAP Note, the module says so.
1. Who actually chooses the model — three layers
Before comparing models, ask who holds the choice, because the answer changes what you are allowed to optimise.
Prerequisites
- Completion of M325 (SAP Generative AI Hub — Hands-on) or equivalent: you can create an orchestration deployment and call it with a config
- Completion of M333 (AI & LLM Fundamentals for SAP Consultants) or equivalent working knowledge of tokens, context windows and grounding
- Access to (or a recorded export of) a generative AI hub tenant discovery call — the lab can be done from the What's New log if not
Outcomes
- Answer "which model should we use?" with a dated roster, a shortlist justified by constraints and a measured pass rate, not a leaderboard screenshot
- Tell a client precisely what is and is not known about the models behind Joule, and which model names are verified in the hub versus only at the vendor
- Produce a reproducible evaluation package (masked dataset, pinned versions, repeats, validated judge) that a second consultant can rerun
- Quote cost per correct task in capacity units with explicit assumptions, without mixing AI Units or vendor prices
- Deliver a lifecycle design with fallback chain, restriction list, rate-limit handling and a retirement-date review cadence
Full module available to members. The full module adds: the decision framework · the end-to-end scenario walkthrough · the KPI scorecard · the anti-patterns · the code blocks · the knowledge check · the diagrams.