Data Quality Frameworks
As of 2026-09-27
What is Data Quality Frameworks?
Data quality frameworks range from simple null checks to statistical drift detection on production pipelines — the discipline of rules, metrics and remediation that keeps data trustworthy.
What it is
A data quality framework is the set of rules, measurements and remediation workflows that decide whether a number is allowed to reach a decision-maker. It spans the trivial (a not-null constraint on a key) and the statistical (drift detection on a production pipeline), and its real subject is not correctness but accountability: who is told when a value is wrong, and what happens next.
Why it matters
On an SAP analytics engagement, quality is the discipline that decides whether a semantic layer survives contact with the business. A dashboard that is 3 % wrong is not 97 % useful — it is unusable, because no one can tell which 3 %. The framework is what converts "the data looks off" from an argument into a ticket with an owner.
It is also increasingly a compliance surface rather than an engineering preference. For AI systems in scope, the EU AI Act requires appropriate levels of accuracy and robustness and that data governance practices be documented (EUR-Lex — Regulation (EU) 2024/1689). A model trained on a pipeline with no quality controls is not merely risky; it is difficult to declare.
How it works
The canonical decomposition — completeness, uniqueness, timeliness, validity, accuracy, consistency — comes from the data management body of knowledge and is worth using because clients recognise it (DAMA — DMBOK). Each dimension becomes a measurable rule, each rule gets a threshold, and each threshold gets an owner who is paged when it breaks.
Why it matters in practice
- A drift-detection layer catches quality erosion that a one-time null check never will — the two belong at opposite ends of the same maturity curve, not as alternatives.
- Remediation workflows, not just detection rules, are what turns a quality alert into a fixed dataset — detection without a remediation path is a dashboard nobody acts on.
- This is depth-of-field knowledge consultants use to anchor rate negotiations — a client that trusts the data trusts the consultant who guarantees it.
Key points
- The set of rules, metrics, and remediation workflows that keep data trustworthy — not correctness alone, but accountability for who is told when a value is wrong.
- Ranges from simple null and referential-integrity checks to statistical drift detection on production pipelines.
- The DAMA-DMBOK dimensions — completeness, uniqueness, timeliness, validity, accuracy, consistency — give clients a recognisable vocabulary for each rule.
- A rule without a defined remediation path (block / quarantine / warn) is a metric, not a control — it becomes a red square everyone learns to ignore.
- Sequence structural checks (null, referential integrity, type consistency) before statistical drift detection — drift alerts on an unstable structural layer produce false signals.
- Data contracts attach quality guarantees to a named producer/consumer relationship, converting quality from after-the-fact policing into a pre-build agreement.
- In an AI pipeline, poor upstream data quality does not make the LLM fail loudly — it makes the LLM answer fluently from bad grounding, which is harder to catch.
- Tabular foundation models using in-context learning (SAP-RPT-1.6, TabPFN-3.5-Plus) have no training step to average out bad data — a malformed input table degrades that call's predictions immediately.
Terms used on this page
- Data owner
- The business stakeholder accountable for the correctness of a data domain (not the IT team).
- Data steward
- The operational role that maintains master data quality day-to-day.
- Data contract
- A named agreement between a producer and a consumer stating guaranteed quality dimensions and the notification path when a guarantee fails.
- Drift detection
- Statistical monitoring for a value distribution shifting away from its historical baseline — a maturity-stage control, not a starting one.
- Referential integrity check
- A structural rule verifying that a foreign-key value in one table actually exists in the table it references.
- Quarantine (remediation pattern)
- Isolating the rows that fail a rule so the rest of the load proceeds, instead of blocking the whole load for a minority defect.
- In-context learning
- A tabular foundation model's prediction mode (SAP-RPT-1.6, TabPFN-3.5-Plus) that reads a labeled table at inference time with no separate training step — making per-call data quality a live dependency.
- Eval suite
- A structured test set scoring a model's or agent's output against expected behavior — functionally a data quality framework applied to AI outputs.
Sources
- DAMA-DMBOK — data management body of knowledge
- EUR-Lex — Regulation (EU) 2024/1689 (AI Act), Article 10 data governance (2024-07-12)
- SAP Help Portal — SAP Master Data Governance overview
- SAP Help Portal — SAP Datasphere Data Access Control and quality-relevant validation rules
- SAP Community — orchestration service data masking and content filtering (grounding pipeline)
- ISO — ISO 8000-61:2016, data quality management: process reference model
- Harvard Business Review — Redman, "Bad Data Costs the U.S. $3 Trillion Per Year" (22 Sept 2016; headline estimate of the economic cost of poor data quality, an estimate not a measurement)
- PMC (Wang et al., 2023) — Overview of Data Quality: Examining the Dimensions, Antecedents, and Impacts of Data Quality (reviews Wang & Strong's four-category dimension framework)
- Bitol / Linux Foundation — Open Data Contract Standard (YAML standard for schema, quality, SLA and ownership of a data contract)
- Future of Life Institute AI Act Explorer — Article 10: Data and data governance (high-risk AI datasets: relevant, representative, free of errors and complete; unofficial text, check EUR-Lex)
- European Commission — AI Omnibus enters into force (27 July 2026; Annex III high-risk rules apply from 2 Dec 2027, Annex I from 2 Aug 2028)
- Great Expectations docs — GX Core overview (open-source framework for expressing data tests and validating data against them; vendor docs)
- Prior Labs — TabPFN-3.5 Technical Report (in-context tabular prediction; vendor research)
Full card available to members. What the full card adds: the full decision framework · the SAP vs Snowflake / Databricks / Fabric comparison · the common pitfalls and their fix · the cheat sheet · the architecture schemas · the code blocks · the facts worth quoting.