AI-ready data — quality, semantics and lineage for AI
As of 2026-09-25
What is AI-ready data?
SAP's own language for its data strategy — "trusted, AI-ready data foundation" built on "semantically rich, trusted data products" — names three separate ingredients an agent actually needs: quality it can trust, semantics it can reason with, and lineage it can be accountable to. Each has a concrete SAP mechanism behind it, and skipping one produces a fluent agent that is wrong, not one that fails visibly.
Three words hiding inside one marketing phrase
SAP's November 2025 framing of its Business Data Fabric strategy, announced alongside the Snowflake partnership, describes the goal as a "trusted, AI-ready data foundation to harmonize SAP and non-SAP data," built on "semantically rich, trusted data products" that ground AI applications "in organizational knowledge." Read carefully, that sentence bundles three separate engineering problems that a delivery team has to solve separately, because none of them is solved by solving another: data can be perfectly clean and still meaningless to an LLM without semantics; it can be well-labelled and still untrustworthy if no one can trace where a number came from; and it can be both clean and well-labelled today and silently wrong tomorrow if quality is never monitored.
Why it matters
- "AI-ready" bundles three separate engineering problems — quality, semantics, lineage — that have different owners and different tooling; treating it as one checkbox is how a client ends up with a fluent, wrong agent.
- Zero-copy sharing (C349) solves data movement, not quality, semantics or lineage — a data product shared faster still carries whatever gaps it had before sharing.
- Lineage gaps and semantic gaps produce no visible defect until an audit, a GDPR Article 30 request or an EU AI Act traceability question forces the issue — by which point the agent has already answered thousands of times.
Key points
- SAP's own phrase, "trusted, AI-ready data foundation" on "semantically rich, trusted data products," bundles three separate problems: quality, semantics, lineage.
- Quality for AI needs continuous monitoring and drift detection, not one-time cleansing — a wrong answer delivered fluently is harder to catch than a visible null.
- Semantics comes from Catalog metadata (business terms, KPIs, descriptions); AI-assisted catalog content generation is a premium feature and still needs human approval.
- Lineage is auto-captured inside Datasphere for native objects, but needs manual Catalog API entries for external sources.
- Zero-copy sharing (BDC Connect to Databricks, Google, Snowflake) solves data movement, not any of the three AI-readiness ingredients.
- Score quality, semantics and lineage as three separate lines in an AI-readiness assessment — each has a different owner and remediation path.
Terms used on this page
- AI-ready data
- SAP's term for data that is simultaneously trustable (quality), reasonable (semantics) and accountable (lineage) — not simply movable.
- Drift detection
- Statistical monitoring that catches a data pipeline silently becoming incorrect after having been correct.
- Catalog API
- The Datasphere interface used to manually register lineage for data reaching the platform from external sources.
- Business Data Fabric
- SAP's strategic framing for an open, AI-ready data ecosystem spanning SAP and third-party platforms.
Sources
Full card available to members. What the full card adds: the full decision framework · the common pitfalls and their fix · the cheat sheet · the facts worth quoting.