AI Observability in SAP AI Core — Token Usage, Hallucination, Drift
As of 2026-09-27
What is AI Observability in SAP AI Core?
In SAP AI Core, observability is built from documented parts, not a ‘Model Gateway’: token usage returned on every orchestration response, inference observability (GA 8 May 2026) to record flagged requests with labels and feedback, and the evaluation service's LLM-as-a-judge groundedness metrics to detect hallucination and drift offline.
What AI observability means here
AI observability is the continuous measurement of how a generative AI system behaves in production — not in a demo, but on real traffic. Three questions matter: how many tokens each call consumes (cost), whether the answers are supported by the data the model was given (hallucination), and whether quality shifts over time as prompts, grounding sources or model versions change (drift). The failure mode that makes this a governance requirement is silence: a broken report throws an error, a hallucinating assistant returns a confident, well-formatted wrong answer.
Earlier versions of this card attributed token metering to an ‘AI Core Model Gateway’. SAP ships no component of that name. In SAP AI Core the instruments are the generative AI hub's orchestration service and harmonized API (which return token usage on every response), inference observability (general availability 8 May 2026), the evaluation service with its system-defined metrics, and the BTP usage and budget tooling. Joule's own model calls run under SAP's management and are not observable in your tenant; what follows applies to what you build on the generative AI hub.
Why it matters
- A hallucinating assistant fails silently: without recorded inferences and groundedness scoring, a systemic error is discovered by users acting on it, not by the team.
- Nothing is recorded by default in SAP AI Core — neither the audit log nor inference observability captures prompts unless the request headers ask for it.
- EU AI Act logging duties for high-risk systems (Art. 12, Art. 26(6)) apply from 2 December 2027 for Annex III systems; the evidence trail has to be designed, not retrofitted.
Key points
- No ‘Model Gateway’: token counts come from the usage object (prompt_tokens, completion_tokens, total_tokens) of every orchestration and harmonized-API response.
- Inference observability (GA 8 May 2026, Cloud Foundry, not sovereign cloud) records only requests flagged with ai-inference-observability-persistence-mode = metadata or full.
- Metadata mode: model name and version, input/output tokens, latency, up to 16 ext.ai.sap.com labels — no object store needed. Full mode adds request, response and feedback in an S3 store.
- Payloads are stored unmasked and unsanitised: consent for personal data and injection-safe consumption are the customer's job; metadata deletion does not purge S3.
- Hallucination: evaluation service (genai-evaluation) with Pointwise RAG Groundedness (1–3), Context Relevance (1–3) and Correctness (1–5) as LLM-as-a-judge metrics.
- SAP found computed metrics (BERTScore, BLEU, cosine similarity) weak for grounding evaluation compared with LLM-as-a-judge — thresholds must be calibrated on your own data.
- Drift: pin model.version, track deprecations in SAP Note 3437766, monitor intermediate_failures of fallbacks, re-run a golden dataset on every change.
- The AI Core audit log holds management events, not prompts; regulatory evidence needs full-mode recording or your own logs.
Terms used on this page
- Inference observability
- Generative AI hub capability (GA 8 May 2026) that records flagged inferences to orchestration or foundation models as metadata or full payload, with labels and feedback.
- Persistence mode
- Value of the header ai-inference-observability-persistence-mode: metadata (model, tokens, latency) or full (plus request, response and feedback in S3).
- Inference record
- Stored result of a recorded inference, retrievable by inference ID or label selector; deleting its data removes metadata and labels, not S3 payloads.
- LLM-as-a-judge metric
- Evaluation metric scored by a language model, such as Pointwise RAG Groundedness, Context Relevance or Correctness in the evaluation service.
- Groundedness
- Degree to which an answer is supported by the context provided to the model; the main hallucination signal for RAG.
- Drift
- Change in answer quality over time caused by prompt, grounding-source or model changes, detected by re-running a fixed evaluation dataset.
- intermediate_failures
- Orchestration response field listing why higher-preference configurations were skipped — the trace of a fallback.
Sources
- SAP AI Core — Inference Observability (SAP-docs)
- SAP AI Core — Record an Inference in Inference Observability (SAP-docs)
- SAP AI Core — Retrieving an Inference Record (SAP-docs)
- SAP AI Core — Deleting Inference Record Data (SAP-docs)
- SAP AI Core — System-Defined Evaluation Metrics (SAP-docs)
- SAP AI Core — Computed Metrics vs LLM-as-a-Judge Metrics (SAP-docs)
- SAP AI Core — Orchestration Workflow V2 (usage object, SAP-docs)
- SAP AI Core — What's New (inference observability GA, SAP-docs)
- SAP AI Core service guide — metering incl. inference observability (PDF, 4 Sep 2026)
- EU AI Act — Regulation (EU) 2024/1689 (EUR-Lex)
- OWASP — Top 10 for LLM Applications 2025 (LLM10 Unbounded Consumption, LLM09 Misinformation)
- NIST — AI Risk Management Framework 1.0 (Govern, Map, Measure, Manage; 26 Jan 2023)
- OpenTelemetry — Semantic conventions for generative AI (spans, metrics, events)
Full card available to members. What the full card adds: the full decision framework · the SAP vs Snowflake / Databricks / Fabric comparison · the common pitfalls and their fix · the cheat sheet · the architecture schemas · the code blocks · the facts worth quoting.
Guides that answer with this page
These guides cite this page as one of the sources their answer rests on.