Analytics Legends The knowledge platform for SAP Analytics
Academy module

Data Catalog Strategy

Data catalog flow from raw object to certified self-service product — architecture diagram for Data Catalog Strategy, Analytics Legends Academy module M080

As of 2026-08-16

A data catalog stops the platform becoming a swamp. Three jobs: inventory (searchable registry), business glossary (agreed term definitions linked to physical objects so "revenue" means one thing), classification & stewardship (sensitivity tags + named owner per product). Inventory-only is a phone book; glossary + stewardship are where governance lives. SAP Datasphere's catalog is metadata-powered and links lineage (M077); the strategy decisions are taxonomy, glossary ownership, and the certification process — not "switch it on". Active metadata (continuously harvested) beats passive. The catalog makes "data products with contracts" real (BDC's pattern, M068). Findability is the precondition for self-service that doesn't fragment into silos.

What you will learn

  • Design a data catalog governance model covering the three catalog jobs — inventory, business glossary, and classification & stewardship — with a named owner per glossary term and per data product
  • Define a certification process that moves a data product from raw to catalogued to certified for self-service consumption, and decide catalog scope by governance value rather than cataloguing every legacy object
  • Configure SAP Datasphere's catalog to harvest active metadata (continuously enriched: usage signals, freshness, quality scores) and link it to lineage (companion module M077) so a catalog entry shows verifiable provenance
  • Diagnose the shadow-dataset failure mode — business users rebuilding their own datasets because they cannot find or trust a central one — and design the findability and stewardship fixes that let self-service analytics scale without fragmenting

A data catalog is what stops a data platform from becoming a data swamp, and the moment that risk becomes real is entirely predictable: as soon as an organisation has accumulated dozens of data products spread across multiple Datasphere spaces, several BDC-connected non-SAP sources, and years of accreted BW content still in use, the binding constraint on the platform's value stops being "can we technically store and model this data" and becomes "can anyone actually find the right, trusted version of this data — and know, without asking around, that they can trust it." A catalog is the answer to that second question, and catalog strategy is fundamentally a governance discipline that happens to be implemented in software, not a tooling installation project that happens to touch governance.

Three jobs a catalog does

A catalog that only performs the first of these three jobs is not actually solving the problem it was bought to solve, which is why catalog strategy conversations that stop at "which tool" rather than "which governance model" tend to disappoint.

Prerequisites

  • Intermediate hands-on experience on SAP analytics projects
  • Review core concepts first: C041, C038, C037

Outcomes

  • Explain the three jobs a data catalog performs and why inventory alone is insufficient governance
  • Design a glossary-governance and certification model with named ownership per term and per data product
  • Configure active-metadata harvesting linked to lineage so catalog entries stay trustworthy as the estate evolves
  • Diagnose and remediate the shadow-dataset / self-service fragmentation failure mode with a findability and stewardship fix

Full module available to members. The full module adds: the decision framework · the end-to-end scenario walkthrough · the KPI scorecard · the anti-patterns · the code blocks · the knowledge check · the diagrams.

Open in the app →