Data quality is not a property of a dataset. It is a relationship between a dataset and one specific decision, and once you see that, the readiness program you have been funding for two years stops making sense.
On 2 April 2026, the Hong Kong Monetary Authority issued its circular on Granular Data Reporting 3.0, with provisional implementation parameters attached as an annex. It is a multi-year program: three phases running tentatively from 2026 to 2031, the first of which (2026–2028) is concerned with establishing “the foundation and initial scope for granular submissions” by preparing data governance and scalable infrastructure. Banks will progressively stop sending the regulator summary templates and start sending it the underlying rows.
Most of the commentary has treated this as a reporting-infrastructure story. Upgrade the pipes, hire the data engineers, budget for 2031.
Buried in the implementation parameters is something far more interesting, and almost nobody has read it as an AI document. The HKMA specifies that data validation controls “should be anchored on ‘validation points’ defined for different data areas, i.e. key data points that can be systemically reconstructed by the granular data,” with “a predefined tolerance band (e.g., ±0.1%)” where necessary.
Because a regulator has just quietly answered the question your AI program has been arguing about for eighteen months.
It did not ask whether the data is good. It asked whether a named figure can be reconstructed from the granular record, within a *stated* tolerance. Quality is not a virtue, It is a measurement, taken at a specific point, against a specific number.
Meanwhile, three floors up, the AI steering committee is being told the models cannot go live because the data is not ready.
The consensus has the sequence backwards
Ask any institution why its AI deployment is behind and you will get the same answer: our data isn’t ready. Then comes the remediation program. Catalogue everything. Profile everything. Establish lineage across the estate. Stand up a governance council. Two years, a large number, and a roadmap whose final milestone is a state called “data readiness.”
Every one of those activities is defensible. The program is still built on a category error.
The consensus frames this as a trade-off along a single axis — speed against quality, deploy now and accept risk, or clean first and accept delay. Both camps agree on the axis and argue about where to sit on it. But the axis is invented. It exists because we have been treating data quality as an *attribute of a dataset*, like temperature, when it is in fact a *relationship between a dataset and a decision*, like fitness.
The thesis
Data quality is not a property of your data. It is a relationship between your data and one named decision, which means “our data isn’t ready” is never a true statement, only an unfinished one.
Finish the sentence and the contradiction between quality and speed does not get resolved. It was a segmentation error.
This is TRIZ in its most useful mode. When a system is being fought over on an axis where both ends are correct, the axis is usually wrong. Separate the problem along a different dimension, here, by use case rather than by time, and the two sides stop competing for the same resource.
What follows:
- Why enterprise-wide readiness programs are structurally incapable of terminating what they have plan so carefully
- The decision register that must exist before the data catalogue does
- Converting “clean” into a threshold you can actually pass or fail is a better way to approach the situation
- The three data populations most institutions are treating as one, and what it costs them
- A diagnostic to run before Friday that tells you, in one number, whether your data program has a destination
The framework: segment by decision, threshold by tolerance, reconstruct by point
Three tiers. The first is a fortnight of conversation and produces no engineering. The second changes what your pipelines are for. The third is where you stop reporting on quality and start proving it.
A framework that cannot be held in a practitioner’s head between reading it and the next steering committee is a document, not a tool.
Tier One — Foundation: name the decision before you touch the data
1. The decision register precedes the data catalogue
The universal pattern. Any quality standard requires a purpose to be measured against, because “sufficient” is a two-place predicate, sufficient *for what*. Remove the second term and the standard becomes unfalsifiable, and an unfalsifiable standard cannot be met, only funded. This is why enterprise data-readiness programs so rarely end. They are not failing. They are structurally incapable of concluding, because no one has specified the condition under which they would be allowed to stop.
The regulated-industries manifestation. Regulators reached this conclusion first and stated it in the plainest possible terms. The HKMA does not ask banks for good data; it asks for named validation points reconstructible within ±0.1%. MAS, in the *MindForge AI Risk Management Operationalization Handbook* published in March 2026 at the conclusion of Project MindForge phase two, organizes its guidance around AI inventories, materiality assessments and key risk indicators, all of which are use-case-scoped objects. Not one major regulator in this region has asked an institution to certify that its data is, in general, good. They have asked what specific thing it is being used to decide, and how well it needs to perform to be trusted with that.
Your translation. Before any further remediation spend is approved, produce a register of decisions the institution intends to make with AI over the next twelve months, underwriting triage, claims routing, suitability checks, surveillance alerting, whatever they are. One row per decision. Each row names the decision, the owner accountable for its outcome, and the consequence of getting it wrong. This is a business exercise conducted with the data office in the room, not a data exercise with a business stakeholder.
What to build by Friday. The register itself, at roughly twenty rows, deliberately incomplete. The purpose is not coverage. The purpose is to discover, reliably, and how many funded data work streams cannot be mapped to any row on it.
2. Convert “clean” into a threshold with a number in it
The universal pattern. A quality target expressed as an adjective cannot be passed or failed, which means it cannot gate anything. A target expressed as a number against a named field can. The tolerance band is not a compromise on rigor; it is the mechanism that makes rigor operable, because it converts an argument into a test.
The regulated-industries manifestation. The ±0.1% in the HKMA’s parameters is worth dwelling on, because it is doing something subtle. It concedes, in a formal supervisory document, that perfect reconstruction is not the standard. What is required is that the deviation be *bounded, stated in advance, and defensible*. That is a far more sophisticated position than the one most internal data policies take, which is to demand completeness in principle and then quietly tolerate whatever arrives in practice, a posture that produces neither quality nor speed, only ambiguity about which one was sacrificed.
Your translation. For each decision on the register, specify the fields that actually feed it, usually a startlingly small number, frequently under twenty, and for each field, a tolerance: acceptable null rate, acceptable staleness in days, acceptable variance against the system of record. The threshold is set by the decision owner, who bears the consequence, and reviewed by risk. He is the one who know what the decision can survive.
What to build by Friday. A single-page fitness threshold sheet for your highest-materiality use case. Fields down the left, three tolerance columns across, decision owner’s signature at the bottom. If nobody will sign it, you have learned that the decision has no real owner.
Tier Two — Process Maturity: stop treating three populations as one
3. Training data, inference data, and evidence data are different objects
The universal pattern. Systems that conflate distinct populations under one label inherit the strictest requirement of all of them and the clarity of none. Data used to build a model, data the model consumes in production, and data retained to explain the model’s behavior later have genuinely different quality profiles, different retention obligations, and different failure consequences. Governed as one estate, every one of them gets over-controlled in the places that do not matter.
The regulated-industries manifestation. The third population is the one institutions consistently underweight, and it is the one supervision runs on. When an examiner asks why a particular customer received a particular outcome fourteen months ago, the question is not answerable from the training set or the current production feed. It is answerable only from a preserved record of the inputs, model version and decision at that moment. This is where personal-data obligations bite hardest, and where the regional direction of travel is unambiguous: Hong Kong’s Privacy Commissioner for Personal Data launched compliance checks on 60 organizations in January 2026 and published results in May 2026 showing 95% used AI in day-to-day operations, with more than half running three or more AI systems, recommending governance structures, privacy impact assessments, AI audits and incident-response plans. Note the shape of that finding. The problem is that adoption has outrun the ability to reconstruct what was adopted.
Your translation. Tag every data flow feeding an AI use case as training, inference, or evidence. Apply different controls deliberately: training data needs representativeness and provenance; inference data needs freshness and availability; evidence data needs immutability and retrievability under a retention clock. The same physical table often appears in more than one population wearing a different obligation each time, and that is fine — provided somebody has noticed.
What to build by Friday. A one-page flow map for a single use case with the three populations colored differently. The gap that shows up in almost every first attempt is the evidence population, which frequently turns out not to exist at all. Its function assumed to be covered by application logs that were configured for debugging and are rotated out after thirty days.
4. Provenance is a deployment gate, not a remediation task
The universal pattern. Controls applied at the end of a process are cleanup; the same controls applied at the entry point are architecture. The cost differential is not marginal. It compounds with every downstream system that has already consumed the unverified input and made it load-bearing.
The regulated-industries manifestation. Provenance in a regulated deployment answers three questions: where did this data originate, what right do we have to use it for this purpose, and can we demonstrate both without a two-week archaeology exercise. The second question is the one that quietly kills projects late. A dataset lawfully collected for policy administration is not automatically available for model training under the purpose-limitation principles running through PDPA and PDPO alike. Institutions discover this at the point of deployment with unhelpful regularity, having already built the model.
Your translation. Add provenance to the go-live checklist as a gate with the power to stop a release, alongside security review and model validation. Three fields per source: origin system, lawful basis for this specific purpose, and named data owner. If any field is blank, the deployment does not proceed. The discipline is unpopular for roughly one quarter and then becomes invisible.
What to build by Friday. Three lines added to your existing go-live checklist. That is genuinely the whole deliverable, and it is the highest-leverage half-hour.
Tier Three — Scale: make quality reconstructible rather than reportable
5. Replace the quality dashboard with the reconstruction test
The universal pattern. A metric that is reported is a claim. A metric that can be independently rebuilt from primitives is evidence. Mature control environments in every domain move from the first to the second, and the transition is nearly always resisted, because the first is comfortable and the second occasionally embarrasses somebody.
The regulated-industries manifestation. This is precisely the HKMA’s move, and it is why GDR 3.0 deserves attention well beyond the reporting function. A validation point that can be “systemically reconstructed by the granular data” is a quality assertion that survives contact with an examiner, because the examiner can rebuild it. Contrast this with the standard institutional artifact, a red-amber-green data quality dashboard, sixty metrics, most of them green, none of them reconstructible by anyone outside the team that produces them. The pattern generalizes well past financial services: pharmaceutical batch records, grid telemetry, and clinical registries have all run this same migration from asserted quality to reconstructible quality, and all of them found it painful for about a year.
Your translation. Pick three validation points per material use case — figures a senior person would recognize and care about. Define each so it can be rebuilt from raw records. State the tolerance. Then run the rebuild monthly and report the deviation, not the color.
What to build by Friday. One validation point, fully specified: the figure, the reconstruction method, the tolerance, the owner. One is sufficient to expose whether reconstruction is currently possible at all.
A note on confidence
Worth being explicit about what is known here and what is argued.
Confirmed from public documents: the HKMA’s GDR 3.0 circular of 2 April 2026, its phasing, and the validation-point and tolerance-band language in the annexed implementation parameters; the MAS MindForge AI Risk Management Operationalization Handbook of March 2026; the PCPD compliance-check findings published in May 2026.
Strong inference from observed patterns: that enterprise data-readiness programs systematically fail to terminate, and that the evidence population is the most commonly missing of the three. I have seen this often enough across enough institutions to state it with confidence, but it is a pattern claim, not a measured one. Design hypothesis, offered for testing: that the decision register reliably outperforms the data catalogue as a starting artifact. The reasoning is sound and the early evidence is encouraging.
The data ready for a decision is the only data that is ready.
The Prompt Kit
Reading a framework and applying it to your own estate are separated by a surprising amount of work. Mostly the tedious kind, where you must hold a definition steady across twenty use cases without drifting. That is exactly the work a well-built prompt sequence absorbs.
Prompt Kit 09 runs the diagnostic in four sequenced steps: building the decision register from a plain description of your AI portfolio, deriving fitness thresholds per decision, separating the three data populations for a chosen use case, and specifying a reconstruction test with a defensible tolerance. It is designed to be run by the person who owns the decision, not only by the data team, which is rather the point of the whole essay.
Before Friday
Take your highest-materiality AI use case. Write down, without consulting anyone, the single figure that would tell you the model is being fed adequately. Now attempt to reconstruct that figure from raw records, not from the dashboard, from the underlying rows.
Time the attempt. That duration is your actual data quality score, and it is the only one an examiner will ever experience.
Most people find the reconstruction either impossible or possible in a way that required a specific individual’s tacit knowledge, which is the same finding with a different packaging. Either way, the exercise takes an afternoon and tells you more than the last quarterly data council pack did. Whether it is a comfortable afternoon depends largely on how long the dashboard has been green.
If you are standing up AI governance against a regulatory timetable and want a structured outside read on decision registers, fitness thresholds, or reconstruction testing, particularly across multiple jurisdictions where the obligations do not align neatly, mention it in your reply. I take one such conversation per month.
Replies to this post reach me directly. I read all of them.
Regulated Intelligence — TRIZ × AI | Regulated Markets. Written by JL CREPPY.


