Merely claiming that your data is accurate is insufficient. Especially in a worldwide financial network, when a single transaction can go through systems in five different jurisdictions before being included in a regulatory filing, accuracy with no provenance is simply not credible. The standard that can be defended is one in which you can demonstrate the origin of your data, what was done with it, and why.
The problem isn’t dirty data, it’s invisible data
Many data quality discussions in banking are focused on cleaning – whether it’s finding duplicate records, fixing incorrectly entered values, or making sure data adheres to some predefined format. But all the cleaning in the world only treats the symptoms of what can be an exponentially more insidious problem. The reality is, cleaning perpetuates the lofty illusion that existing processes are necessary in the first place. Most institutions don’t realize they have a problem with data quality until their reporting division and regulators have already sounded the alarm, and by then, it’s often too late.
The real problem with data quality is that most banks can’t trace the lineage of their business data – and especially financially significant data – from source to final report. Take, for example, a single customer record. It’s originally entered in an onboarding system and retrieved by a KYC platform. Perhaps that evening, an ETL process pushes records from the onboarding system and KYC platform into a data warehouse. A risk analyst runs a query, pulls the customer record from the data warehouse and pushes it into a spreadsheet. The next morning, your controller runs a query and pulls the customer record from the data warehouse and pushes it into a regulatory report. In each of these steps, the record could be transformed, kept, aggregated, truncated, slightly altered, or severely bungled. If it winds up incorrect on the regulatory report, no one can tell you which of those steps was the culprit.
What regulatory compliance actually requires
Regulators won’t take your word for it that the Asia Pacific operations only used clean, structured data from the corporate global data warehouse. They’ll ask for proof, and “I swear” isn’t it. Efforts to demonstrate without adequate tools can get ridiculous – you have to show your work, not just your answers, and in systems that have hundreds if not thousands of inputs to critical reports, that can involve a literal forest of paper or digital reports. A diligent auditor will not appreciate being handed a crate of ring binders.
Of course, proving your data’s provenance is just the first step. Accurate numbers are part of the game but, as noted, impressing auditor and regulator alike also means proving your mastersheet didn’t contain a formula error, an outdated risk metric, or an improperly debated tax scenario. This is where automated Data Lineage for Banking Institutions becomes a practical necessity rather than a technical preference. When complex data flows cross disparate international systems, visualization of those flows allows compliance teams to show exactly which source system fed a CCAR submission or a MiFID II disclosure, and what transformations occurred at each stage.
Manual mapping is where institutional risk hides
Spreadsheets are a lot more prevalent in mapping financial data than you might think. And people keep them up to date religiously. Until somebody departs, or a program evolves, or you add a new region and they have slightly different systems and start sending you different types of data. And it takes about three months for the spreadsheet to get out of date. And maybe six months before you even notice that it’s out of date. Because nobody’s updating it that fast. This isn’t a process failing. It’s a structural failing. The network is simply too dynamic for static documentation. The lift-and-shift of your Hadoop platform into the cloud created data movements and transformed data that our last batch mapping exercise didn’t capture. The lift-and-shift of your mainframe into modern architecture left whole subsystems running on the old box then feeding reports over a network share. Nobody knew.
This dynamism and complexity make the cliche that ‘operational risk is only increasing’ very real. Poor data quality has real financial consequences. The average financial impact of poor data quality on organizations is $9.7 million per year (Gartner), and that’s from 2019. The estimated financial loss suffered by financial institutions due to poor data quality is five times that of the average business – or $12.9 million (also Gartner). In banking, where the cost of a compliance breach or a misstatement in a capital adequacy report can reach multiples of that figure, this isn’t background noise.
Bridging technical teams and business stakeholders
One of the costs we talk about less is the fact that this type of poor data traceability erodes trust, and that’s expensive, too. Data engineers, the people closest to the infrastructure, don’t trust that the pipes won’t break at exactly the wrong time. Meanwhile, the rest of the business is left to rely, ultimately, on the systems of trust they’ve created – typically a small number of people who’ve earned a former colleagues email or a reputation for being “data whisperers”.
Risk officers and CFOs trust the numbers they’re given – because they have to – without context or confidence in the processes that produced them. When something looks off and the CFO has to sign a statement assuring the data’s accuracy, there often isn’t time for the software engineer to debate “what counts” or for SMEs to arbitrate the dispute.
Metadata management is the solution to this trust gap. Rather than a technical artifact produced as a side-effect of a project, metadata management serves as a shared, living asset. Regardless of where you work in the company or where your expertise lies, metadata is always there for the taking.
Accuracy is something you prove, not something you claim
Outdated spreadsheets and manual records cannot adapt to a financial network that transforms with each new market, retired legacy system, or amended regulatory guideline. The organizations that are successful in this space no longer view data accuracy as a consolidation effort managed by IT groups. They view it as an ongoing capacity of the entire organization.
If you’re unable to track your business data from the start to the final report, the truth is, you’re uncertain as to its accuracy. You’re merely assuming it is. And that point makes all the difference in the world when you’re asked to prove it, particularly in the middle of a regulatory inquiry.





Leave a Reply