AI & Analytics Legends The knowledge platform for SAP Analytics
Concept card

LLM Observability — Tracing, Token Accounting, Hallucination Detection

LLM Observability — Tracing, Token Accounting, Hallucination Detection — Analytics Legends section illustration for the SAP Analytics knowledge base (concepts, studies, Academy)

As of 2026-10-06

What is LLM Observability?

Traditional APM tools can't answer whether an LLM's answer was actually right — observability platforms fill that gap by sampling 1-5% of traffic plus 100% of the lowest-confidence answers against a judge model.

LLM observability is the discipline of seeing, for every model call a production system makes, what happened, why it happened, what it cost, and whether the answer was actually right. Conventional application performance monitoring answers the first two questions poorly and the third and fourth not at all — it was built for deterministic code paths, not for a system whose output varies from one nearly-identical call to the next. A newer generation of purpose-built platforms has grown up specifically to fill that gap, extending general distributed-tracing standards with attributes designed for generative AI: which model served the request, how many input and output tokens it consumed, what it cost, and how long it took.

The Four Surfaces That Matter

Distributed tracing of every model call is the foundation. Each user request becomes a tree of linked steps — user input, any retrieval calls, a reranking step, the model call itself, any tool calls the model triggers, and the final response — captured as a connected trace rather than a set of disconnected log lines. Without that trace, debugging a wrong or strange answer in production is close to guesswork, because there is no way to see which upstream step actually produced the bad input.

Why it matters

  • Four surfaces are non-negotiable: distributed tracing of the full span tree (input → retriever → reranker → model → tools → response), per-call token accounting tagged by tenant/workflow/model, sampled quality evaluation, and hallucination/groundedness detection.
  • The 2026 OpenTelemetry GenAI semantic-convention spec standardises span attributes like gen_ai.usage.input_tokens, adopted by every serious observability vendor (Langfuse, Arize Phoenix, LangSmith, Helicone).
  • Groundedness scoring catches RAG-specific failures (do the answer's claims appear in retrieved context?) while toxicity/PII/jailbreak classifiers catch safety failures — two distinct failure classes needing two distinct checks.

Key points

  • OpenTelemetry GenAI semantic conventions are the 2026 industry standard: `gen_ai.system`, `gen_ai.request.model`, `gen_ai.usage.input_tokens/output_tokens`, etc.
  • Distributed tracing of the full request span tree (input → retriever → reranker → model → tools → response) is non-negotiable for debugging.
  • Per-call token accounting tagged with tenant + workflow + model is the FinOps substrate (companion concept C249).
  • Quality evaluation: sample 1-5% of production traffic, replay against LLM-as-judge or golden dataset, trend the score, alert on drift.
  • Hallucination detection: groundedness scoring for RAG; toxicity/PII/jailbreak classifiers for safety; sample 100% of low-confidence answers.
  • Vendor landscape: Langfuse, Arize Phoenix, LangSmith, Weights & Biases Weave, Helicone — pick one and integrate via OTLP.

Terms used on this page

OpenTelemetry GenAI
Semantic-convention extension of the OpenTelemetry standard defining standard span attributes for generative-AI calls (model, tokens, cost, tools).
LLM-as-judge
Evaluation pattern where a separate LLM scores the quality of another LLM's output against a rubric; cheaper than human review, noisier than human review.
Groundedness
RAG-specific quality metric: do the claims in the answer appear in the retrieved context? Low groundedness = hallucination.
OTLP
OpenTelemetry Protocol — the wire protocol for shipping traces/metrics/logs from instrumented apps to collectors and backends.
Span
A single traced operation within a distributed trace tree — one retrieval call, one reranking step, or one model call — carrying attributes such as duration, token counts, and status; the basic unit OpenTelemetry GenAI semantic conventions standardise.
Drift detection
Monitoring whether a model's or pipeline's real-world output quality moves away from its established baseline, typically by periodically re-running an evaluation against a golden dataset or trending an LLM-as-judge score over time — the practice that turns a silent quality regression into an alert.
Confidence-weighted sampling
An evaluation-sampling strategy that reviews all or nearly all of a system's lowest-confidence responses in full, while sampling only a small percentage of high-confidence traffic — concentrating review effort where hallucinations and errors actually concentrate, rather than spreading it evenly.
SAP AI Core Inference Observability
SAP's own named capability for capturing per-call request/response data, token usage, and latency for calls made through the generative AI hub, documented in SAP's AI Core service guide — the SAP-native answer to this card's tracing and token-accounting surfaces, with a documented caution that captured payloads are stored unmasked.

Sources

  1. OpenTelemetry — Generative AI Semantic Conventions
  2. Langfuse — LLM Observability Documentation
  3. Arize Phoenix — Tracing and Evaluation
  4. OpenTelemetry — Observability primer
  5. NIST — AI Risk Management Framework knowledge base
  6. SAP Help — Generative AI Hub overview
  7. GitHub — open-telemetry/semantic-conventions-genai
  8. OpenTelemetry — Gen AI attribute registry
  9. OpenTelemetry Blog — Inside the LLM Call: GenAI Observability with OpenTelemetry (2026)
  10. MLflow Docs — OpenTelemetry GenAI Semantic Conventions
  11. OpenTelemetry — semantic-conventions-genai PR 498: Add gen_ai.skill.* attributes to the execute tool span (merged 29 Sep 2026)
  12. OpenTelemetry — semantic-conventions-genai PR 270: Add gen_ai.main_agent entity (merged 30 Sep 2026)
  13. OpenTelemetry — semantic-conventions-genai PR 520: Recommend the inference event be of severity level debug (merged 5 Oct 2026)
  14. InfoWorld — Observability for AI-native systems: New SLIs beyond latency and error rate (30 Sep 2026, opinion)

Full card available to members. What the full card adds: the full decision framework · the SAP vs Snowflake / Databricks / Fabric comparison · the common pitfalls and their fix · the cheat sheet · the architecture schemas · the code blocks · the facts worth quoting.

Open in the app →