AI-Ready Data — Quality, Semantics and Lineage
As of 2026-09-25
Gives senior SAP data consultants a defensible definition of 'AI-ready data': not a feature, but four checkable properties — quality, semantics, lineage, authorisation — each tied to a real, cited SAP mechanism. Covers SAP's own May 2026 framing of AI-readiness and the Reltio-based master-data harmonisation entering Business Data Cloud; SAP Data Quality Management microservices for location data; the SAP Datasphere Catalog's business glossary and KPI terms as the literal metadata JustAsk's vector search indexes before generating a query for Joule; Datasphere's Impact and Lineage Analysis (Data Analysis vs Dependency Analysis) and its built-in space-level authorisation boundary; and the principal-propagation pattern that keeps an agent's data access no wider than the requesting user's. Three exercises score a real data product, read a lineage diagram, and write a machine-usable glossary term. Prerequisite for M347, M348 and M349.
What you will learn
- Explain why 'AI-ready' resolves into four checkable properties — quality, semantics, lineage, authorisation — and reject 'we ran it through AI' as a quality control
- Name real SAP data-quality mechanisms (Data Quality Management microservices for location data; Reltio-based master-data harmonisation in Business Data Cloud) and the failure mode each addresses
- Describe what JustAsk's vector search over analytic-model metadata actually indexes, and why an undocumented model is invisible to it
- Use SAP Datasphere's Impact and Lineage Analysis, distinguishing Data Analysis from Dependency Analysis, to trace a column back to its source systems
- State why an agent's data access must inherit a human user's space-level authorisation rather than a broader service identity
Module overview
Who this is for. You know how to model a Datasphere space or a BW/4HANA InfoProvider; now a steering committee wants to know whether that data is "AI-ready" before Joule or an agent is allowed near it. This module gives that word a definition you can defend: not a feature flag, but three checkable properties — quality, semantics and lineage — plus the authorisation boundary that decides who, human or agent, may read what. It builds on M333 (AI & LLM fundamentals) and feeds directly into M347 (the semantic layer as LLM context), M348 (natural-language query) and M349 (vector search on analytic data).
Prerequisites
- M333 (AI & LLM Fundamentals for SAP Consultants) or equivalent working knowledge of grounding and RAG
- Working knowledge of at least one SAP data product (Datasphere space, BW/4HANA InfoProvider or S/4HANA embedded analytics)
- Familiarity with the Datasphere Catalog and space-based authorisation, or willingness to explore a trial tenant
Outcomes
- Score a real data product against a four-point AI-readiness checklist (quality, semantics, lineage, authorisation) and justify each score.
- Trace a column's provenance through Datasphere's Impact and Lineage Analysis and explain what it would and would not show under Dependency Analysis.
- Write a business glossary term structured the way JustAsk's retrieval step actually consumes it.
- Explain to a client, without hand-waving, why an ungoverned data product produces confidently wrong agent answers.
Full module available to members. The full module adds: the decision framework · the end-to-end scenario walkthrough · the KPI scorecard · the anti-patterns · the code blocks · the knowledge check · the diagrams.