AI & Analytics Legends The knowledge platform for SAP Analytics
Concept card

SAP AI Quality Gates — Accuracy, Latency, Cost, Safety Thresholds

SAP AI Quality Gates — Accuracy, Latency, Cost, Safety Thresholds — Analytics Legends section illustration for the SAP Analytics knowledge base (concepts, studies, Academy)

As of 2026-09-27

What is SAP AI Quality Gates?

AI quality gates are a release discipline, not an SAP product: before a generative AI feature goes live it must clear agreed thresholds on quality, latency, cost and safety, measured on the customer's own dataset. The generative AI hub supplies the instruments — Evaluations with system-defined and custom LLM-as-a-judge metrics, prompt optimization, inference observability — but the thresholds are yours to set and defend.

What a quality gate is — and what SAP provides

A quality gate is the release criterion that turns a probabilistic feature into something a business owner can sign off: this configuration of prompt, model, grounding and filters reaches this quality on our cases, answers within this time, costs this much per call and does not produce these harms. SAP does not ship a product called "AI Quality Gates", and it publishes no universal thresholds; earlier versions of this card quoted accuracy, latency and cost targets that no SAP source supports and that have been removed. What SAP does ship, inside the generative AI hub, is the measurement toolkit.

The instruments: Evaluations, metrics, optimization

Evaluations "provides tools for benchmarking large language models and prompts via orchestration configurations": you register a test dataset in your object store as an artifact, choose metrics, and run an execution that scores one or several prompt-and-model combinations — for example gpt-4.1 against gpt-4o-mini with the same template from the prompt registry, or several stored orchestration configs. Results can be compared across runs to detect regressions. The configuration uses the executable genai-evaluations-simplified and parameters such as the dataset path, metric IDs, the orchestration deployment used for judge calls, optional variable mapping and a test row count.

Why it matters

  • Without a written gate, ‘good enough’ is decided by whoever demos last — and nobody can say later why a model version was allowed into production.
  • Model upgrades, prompt versions, grounding changes and fallbacks all alter behaviour; a gate that is not re-run on each change certifies a system that no longer exists.
  • SAP publishes metrics and tooling, not thresholds: numbers borrowed from slides collapse the first time a business owner asks where they come from.

Key points

  • Quality gates are a customer release discipline; SAP provides measurement tools in the generative AI hub, not thresholds.
  • Evaluations benchmark prompt + model combinations or stored orchestration configs on your registered dataset (executable genai-evaluations-simplified).
  • Computed metrics: BERT Score, BLEU, ROUGE, Exact Match, JSON Schema Match, Language Match, Content Filter on Input/Output.
  • LLM-as-a-judge metrics: instruction following, correctness, answer relevance, conciseness (1–5); RAG groundedness and context relevance (1–3).
  • SAP recommends judge metrics over computed similarity for grounding; GPT-4.1 judge correlation 83.74 % / 87.07 % with ground truth.
  • Custom LLM-as-a-judge metrics with structured prompts; prompt optimization writes the best template back to the prompt registry.
  • Latency and tokens come from inference observability metadata; cost from GenAI-token rates (SAP Note 3437766) converted to capacity units.
  • Re-run the gate on every model, prompt, grounding, filter or fallback change; monitor production samples afterwards.

Terms used on this page

Quality gate
Documented release criterion with thresholds on quality, latency, cost and safety that a generative AI configuration must meet.
Evaluations
Generative AI hub feature that benchmarks prompts and models, as orchestration configurations, on a registered dataset with chosen metrics.
Computed metric
Deterministic score such as Exact Match, BLEU, ROUGE or JSON Schema Match, often requiring a reference answer.
LLM-as-a-judge
Metric computed by a second LLM applying a rubric, e.g. correctness or RAG groundedness.
Groundedness
Whether every factual claim of a response is supported by the provided context documents.
Context relevance
How well the retrieved context matches the user's query; measures the retrieval step of RAG.
Custom metric
User-defined LLM-as-a-judge metric written as a structured prompt, with lifecycle management.
Prompt optimization
Automated improvement of a registry prompt template against a dataset and target metric.

Sources

  1. SAP AI Core — Optimizations (SAP-docs)
  2. SAP AI Core — Evaluations (SAP-docs)
  3. SAP AI Core — System-Defined Evaluation Metrics (SAP-docs)
  4. SAP AI Core — Validated Metrics (SAP-docs)
  5. SAP AI Core — Create an Evaluation (SAP-docs)
  6. SAP AI Core — Prompt Optimization (SAP-docs)
  7. SAP AI Core — Inference Observability (SAP-docs)
  8. SAP AI Core service guide — metering and typical consumption patterns (PDF, 4 Sep 2026)
  9. EU AI Act — Article 15 (accuracy, robustness, cybersecurity)
  10. European Commission — AI Omnibus enters into force (27 Jul 2026)
  11. SAP News Center — the Operational Backbone of the Autonomous Enterprise: AI Agent Hub governance at scale (2026-09)
  12. SAP News Center — Secure AI Agents: how SAP and NVIDIA co-define enterprise-grade agent execution (NVIDIA OpenShell, 2026-05-12)

Full card available to members. What the full card adds: the full decision framework · the SAP vs Snowflake / Databricks / Fabric comparison · the common pitfalls and their fix · the cheat sheet · the architecture schemas · the code blocks · the facts worth quoting.

Open in the app →