Data Quality Frameworks & Methodologies for Data Managers

Executive Summary

Choosing the right DQ framework matters a great deal for your organisation’s data quality goals. It should integrate various techniques and tools, support different data types, and stay flexible and straightforward, with clear goals set for each stage.

Essential points discussed in this webinar are:

  • Developing a data quality framework is important
  • PDCA cycle involves planning, executing, checking, and acting for continuous improvement
  • Pareto principle is used to prioritise data quality issues
  • Root cause analysis is conducted to identify underlying reasons for data quality issues
  • Juran, Deming, and ISO offer different approaches to achieving data quality
  • Evaluating the cost of data quality issues is crucial
  • Developing taxonomies for data value realisation and cost reduction is important
  • Planning for data quality involves resolving data quality issues
  • Various frameworks can be utilised based on organisational needs and objectives

Webinar Details

Title: Data Quality Frameworks & Methodologies for Data Managers
Date: 6 July 2023
Presenter: Howard Diesel
Meetup Group: African Data Management Community Forum
Write-up Author: Howard Diesel

Introduction to Data Quality Framework (PDCA)

The session opened with developing a data quality framework as the main topic, and questions came up around the familiar PDCA framework. Monica, studying for an exam at the time, was asked specifically about what PDCA stands for, and Penelope explained the framework, stressing the importance of planning, executing, checking, and acting for continuous improvement.

Figure 1. PDCA Cycle

Data Quality Planning and Frameworks

During the planning phase, data quality issues get prioritised using the “Pareto principle,” targeting the 20% of issues causing 80% of the problems, with root cause analysis identifying what’s actually behind them before a remediation plan gets built. In the “do” phase, identified problems get fixed and new data quality targets get set.

The “check” phase verifies whether those targets were met and identifies the causes behind whatever problems remain, and the “act” phase addresses emerging issues so the whole plan-do-check-act cycle doesn’t need repeating from scratch. DQ frameworks offer various ways to reach specific goals, and larger contexts, company strategy included, can shape the plan too. Focusing on the top 20 issues is generally what covers the necessary ground for a migration or a big project.

Figure 2. DQ Frameworks

Data Analysis and Data Quality

If your data’s in good shape, you can move on to analysing data products or focusing on strategy. If data governance has flagged quality issues, though, those need prioritising and addressing first. It’s worth revamping your product catalogue while rolling out new data products, and running use case quality checks before handing data sets over to data scientists or BI teams matters a great deal.

Categorising data sets as either “salty water” (uncurated) or “freshwater” (curated) in the catalogue helps here too, since only curated sets should really be feeding BI and advanced analytics. The Plan-Do-Study-Adjust (PDSA) cycle is genuinely useful for data analysis processes, planning, researching, reviewing results, and adjusting, which lines up with Magnus’s own take on PDSA. It’s a topic worth revisiting further.

Figure 3. Use Data Analysis to Reconstruct the state of Data

Different Frameworks for Data Quality Assessment and Improvement

The Juran framework prioritises data accuracy over speed and brings its own distinct way of thinking to the table. It gets confused with the Deming framework fairly often, to the point where a specialist exam once mistakenly attributed Juran’s name to PDCA. Juran is also known for the Juran Trilogy, which covers three aspects of data quality assessment and improvement. Other notable frameworks in this space come from Larry English, Tom Redman, Laura Sebastian-Coleman, and Danette McGilvray.

PDCA typically gets applied to the assessment and improvement phases, though the speaker’s own work has focused mainly on assessment. Evaluating the cost of data quality issues matters a great deal, and building out a taxonomy for that cost, alongside taxonomies for data value realisation and cost reduction, helps with assessing and understanding data quality issues generally. Danette McGilvray’s POSMAD framework is another one worth noting, focused on the life cycle of data quality.

Figure 4 Definition: Quality Management Framework

Evaluating Data Quality Frameworks

These frameworks focus mainly on the operational side of data quality rather than definitions, principles, and policies. The process runs through several iterations, data value realisation, maturity assessment, developing a data quality strategy, establishing an operating model, and establishing practices by defining, operationalising, and ensuring continuous improvement matters a great deal throughout. Reviewing and improving quality policies and procedures as things progress is essential too, and standardising and reviewing frameworks means following specific criteria: planning, obtaining, storing, maintaining, applying, and disposing of data. Frameworks like POSMAD and ISO each bring their own distinct approach to achieving data quality.

Figure 5. DQ Framework Analysis

Planning for Data Quality

A comprehensive plan for ensuring high-quality data sits within the Excel spreadsheet, and the Juran Trilogy, quality planning, control, and improvement, gets discussed in detail. The focus here is resolving data quality issues rather than just planning for them, as the DMBOK itself suggests. Danette McGilvray’s approach covers various planning areas, control, assurance, and improvement among them, and ISO 8000-61 brings in the Plan-Do-Check-Act cycle along with additional elements around data architecture, IT, and HR. Different planning, control, assurance, and improvement breakdowns get presented, and which framework works best really depends on organisational needs and objectives, different frameworks suiting different outcomes, whether that’s implementing quality projects or changing organisational structures.

Figure 6 DQ – Data Quality Game Plan

Quality Control in Data Quality

The DMBOK is a valuable resource for understanding data quality, worth noting, but it only covers some of the frameworks, the Juran Trilogy, POSMAD, and ISO 8000 among them. Truly grasping data quality means recognising it as an ongoing process requiring continuous improvement, which is exactly what the “plan-do-check-act” cycle captures.

Technical debt and recurring data errors are persistent problems that need regular attention, and the Juran Trilogy stresses understanding the cost of poor data quality and working at a Quality Control level before moving to the next iteration. Developing data quality rules and gradually strengthening critical data elements is what facilitates the move to higher levels of quality control, going from 50% to 60%, say.

Figure 7. Plan, CONTROL, Improve

Quality Improvement in Healthcare Data

Improving quality is an ongoing process that takes consistent effort, and keeping data clean and accurate matters a great deal for getting there. Implementing a framework is what drives those enhancements, and the Canadian Institute for Health Information (CIHI) has built its own framework to ensure quality across the healthcare industry.

Working in the critical zone of data accuracy can be genuinely challenging and can carry unintended consequences, and setting expectations too low can end up causing dissatisfaction with data accuracy down the line. Regular assessments are necessary to evaluate data quality and adjust expectations as needed, and if quality falls below expectations, discontinuing the report may be the only option. Starting with lower expectations is actually the right move, though, since it establishes a foundation to build quality up from.

Figure 8. Building a Specialised Framework

Importance of Frameworks in Quality Management

Poor data quality is a genuinely significant risk for companies and calls for prompt intervention, sometimes as direct as hiring university students to contact customers and work through data-related issues. Depending on how severe things are, urgent risk management or more measured improvement work may be needed to bring data quality up to acceptable levels.

Choosing the right framework for data governance, NET, ISO, DMAIC among the options, matters a great deal, and standardising criteria is what lets different frameworks actually be compared and pulled into a coherent, comprehensive quality management system. Frameworks, in the end, are what make a comprehensive quality management system possible, and that’s what drives successful outcomes.

Figure 9. Framework Classification

Quality Management Frameworks and Data Certifications in Saudi Arabia

Several frameworks for quality management come up here: TQM, Total Quality ISO, Six Sigma, the European Foundation for Quality, Balanced Scorecard, Lean, and Lean Six Sigma. The suggestion is implementing a certification and standardisation framework for data quality across Saudi Arabia’s ministries, with an independent organisation assessing shared data quality before it’s used.

A lack of standardisation, in one cited case, led to real problems with birth dates. Six Sigma focuses on controlling the quality of data products, while Lean is more about reducing waste and improving efficiency, and comparing different frameworks, examining improvement techniques, and evaluating the dimensions and costs of data quality issues all matter here. ISO and ISTAT handle certification and control across various government ministries.

Figure 10. Distributed Systems & Operating Models

Integrating Quality into the Source System to Avoid Bad Data

Integrating data quality into the source system matters a great deal for avoiding problems downstream. Addressing ambiguous gender selection in the system is one way to improve quality, enhancing the system to ensure correct selection and manage the influx of poor-quality data that comes from getting it wrong. Extending data quality to the organisational level matters too, particularly when sharing data between organisations or group companies, and improving quality means considering the processes used to create, read, and update data, data warehousing among them.

Data quality services matter and should get reviewed from different perspectives, and when choosing a data quality framework, it’s worth looking for the ability to distribute governance across self-sufficient teams. Centralised quality broker services like ISTAT can support multiple organisational structures and cross-boundary use too.

Figure 11. Source System Comparisons

Data Quality Assessment and Improvement Process

The CDQ offers a comprehensive approach to ensuring high-quality data, covering state reconstruction, assessment, and improvement. State reconstruction means rebuilding the data, processes, and organisational structure to address quality issues and missing metadata, while assessment and measurement involve profiling data elements, analysing statistical distribution, and understanding the shape of the data before applying dimensions or data quality rules. Visual tools like histograms help analyse completeness levels and work toward 100% completeness, while also gauging whether that goal is actually feasible.

Figure 12. After Data Analysis: Assess & Improve

Importance of Root Cause Analysis and Improvement in Data Quality

Collaborating with data stewards during the assessment phase matters, and sometimes tackling data quality means lowering expectations rather than raising the bar unrealistically. Root cause analysis plays a key role in identifying what’s actually causing data issues, and the Fishbone (Ishikawa) diagram gets used often for this, alongside the Five Whys technique. Tracking and tracing data lineage matters for understanding data flow, and process analysis is another way to identify data issues. Improvement, when it comes to data quality, can be driven by data or by process, and both quick wins and long-term solutions are worth weighing.

Figure 13. RCA Methods & Techniques

The Juran Trilogy and Data Quality Targets

The Juran Trilogy and data quality targets’ main goal is reducing the error rate and improving overall data quality. The plan-do-check-act framework is genuinely helpful, though it doesn’t cover state reconstruction, which is where the assessment phase comes in, analysing data and understanding data requirements.

Various stakeholders and their expectations come into the conversation too, and identifying critical areas of data corruption using the processing matrix matters. Selecting quality dimensions, and distinguishing between objective and subjective metrics, rounds out the discussion, with the difference between those two types of metrics worth clarifying.

Figure 14 Common Phases & Steps

Qualitative and Quantitative Metrics in Data Analysis

Surveys can provide qualitative metrics for measuring subjective aspects like customer experience, and software usability can be evaluated subjectively in much the same way. People’s behaviour trends, eating habits included, can be used to measure subjectivity too, gauging reactions by, say, analysing the sales impact of swapping animal-based cheese for plant-based.

In health information analysis specifically, differentiating between data quality and information quality matters, since information quality involves analytics and statistical interpretation. Improving data analysis means evaluating costs, assigning responsibilities, identifying root causes of errors, and proposing genuinely qualitative improvement solutions. Improvement strategies can be data-driven or process-driven, and data-driven modifications tend to have a more direct impact on the actual value of the data.

Figure 15 Common Steps for Assessment Phase

Exploring Data Quality Issues and Quality Frameworks

Ensuring reliable data means avoiding compounding previous data issues, and data-driven approaches offer several techniques for that: obtaining high-quality data, using quality brokers, standardisation, record linking, ensuring trustworthiness, pinpointing errors, and making corrections. Process-driven approaches get compared against data-driven ones largely in terms of long-term versus short-term cost efficiency.

Defining accuracy and completeness at the concept level leads to better understanding and significance, and improving data quality really needs justifying through the reduced cost of poor quality. Defining different types of poor-quality data, structured, unstructured, semi-structured, matters too. Choosing the right quality framework, ISTAT, Dynamic Quality, Larry English’s Information Quality, and Wang among the options, takes careful consideration, since not all frameworks cover every aspect, data quality requirement analysis or process modelling among the gaps some leave.

Figure 16 Improvement Strategies & Techniques

Clarification on Frameworks and Metrics for Data Analysis

Howard stressed how important it is to select the right framework for data analysis rather than defaulting to the conventional plan-do-check-act (PDCA) framework, and pointed to various methods for comparing elements, cost qualification and mapping different areas and systems among them. A large-scale ERP system came up, along with the potential benefits of distributed data warehouse cooperatives and web-based systems.

Choosing suitable metrics for the framework matters too, standardised metrics from TDQM alongside the option to add personalised ones, and the discussion covered tools and methodologies for data collection and assessment. The speaker also clarified that the framework doesn’t explicitly mention distributed systems, even though they’re implicitly considered.

Figure 17. Subjective & Objective Measurements

The Impact of Distributed Systems on Data Quality

Magnus is focused on developing distributed systems and organisations that facilitate data sharing across various domains. Using Sera, distributed systems can support interaction and let users request specific data types while checking their quality level, and providers can improve data quality ratings by making changes or enhancements, with users notified when improvements land. Howard then shared an anecdote from their time at a central bank, where switching from Bloomberg to Reuters met real resistance, thanks to a long-standing perception that Bloomberg had once been the superior option.

Figure 18 DaQuinCIS Framework

Data Quality Assessment and Accuracy

When evaluating quality, accuracy becomes a genuinely essential factor to weigh. A Python interface and website offer an extension for calculating data quality that treats accuracy as a significant component, and data quality dimensions like completeness and accuracy get evaluated when setting user expectations.

Data quality requirements determine minimum and maximum accuracy levels and what they actually mean, and accuracy’s granularity can be assessed at various levels, element, row, dataset, data schema. Accuracy itself is measured against two concepts: agreement with the real world and agreement with a source. Several areas, employee data and data privacy among them, need accuracy validation.

Figure 19 DQ Expectation Template

Great Expectations: A Framework for Ensuring Data Quality

Howard brought up a workspace built around a data quality framework where many people can contribute, mentioning a tool called Great Expectations, which uses Python routines to assess data quality, employing probes to validate data quality against various frameworks. He later shared a link to the tool too, pointing to features like data profiling and seamless integration with Snowflake, and mentioned the partnership between Great Expectations and Data IQ to close things out.

Figure 20 Great Expectations

Figure 21 Edge Collection: DQ Frameworks

Scroll to Top