Data Warehousing, BI, Big Data & Data Science for Data Managers

Executive Summary

The landscape of data management covers several key components, Big Data, Data Warehousing, and Business Intelligence among them, each with its own challenges and opportunities. Key themes include building enterprise data warehouses, strategies for effective self-service reporting, and implementing robust data governance frameworks.

Understanding data warehousing architecture matters a great deal, particularly in relation to Big Data and what it means for organisational decision-making. Concepts like the Abate Information Triangle highlight why data integration and analysis matter, and the roles and responsibilities of data managers are pivotal in bridging data science with business performance measurement.

The adoption of models like Data Vault can also significantly reshape enterprise Data Warehousing, reflecting how data management keeps evolving alongside privacy concerns and the fourth paradigm of Data Science, which focuses on the key elements and processes needed for effective data management and statistical analysis.

Webinar Details

Title: Data Warehousing, BI, Big Data and Data Science for Data Managers
Date: 21 November 2024
Presenter: Howard Diesel
Meetup Group: African Data Management Community
Write-up Author: Howard Diesel

Big Data in Data Warehousing and Business Intelligence

Howard Diesel opened the webinar by pointing to how central Big Data has become within Data Warehousing and Business Intelligence (BI), highlighting its role in decision support and generating valuable insight for businesses. The data life cycle runs from data use and enhancement through to feeding Big Data back into planning, all aimed at elevating data assets through the Data, Information, Knowledge, Wisdom (DIKW) triangle, turning data into information, knowledge, and ultimately wisdom. He referenced the “Five Whys and Two Hows” framework too, encouraging exploration of questions like Who, What, When, Where, How, and How Many, which matter a great deal for understanding and measuring business events and metrics effectively.

Figure 1 How-to Analyse Data Using the 5W & 2H

Challenges in Writing the Revised Edition of DMBoK V2

The upcoming revision of the Data Management Body of Knowledge (DMBOK) Version 2 will highlight the differences between the second version and the new edition, which matters for anyone already familiar with Version 2, especially with Version 3 now in the works. As a member of the editorial board for Version 3, I’d welcome collaboration in reviewing the writing for accuracy and clarity, catching any errors along the way. Your input would genuinely help refine this document, since previous editions have had their share of mistakes.

Figure 2 Content of “How-To Analyse”

Figure 3 DW & BI Essential Components

Understanding the Role of Data Warehousing and Business Intelligence

The mind map of Data Warehousing lays out its definition and goals, stressing how important a solid technical environment is for Business Intelligence activities to actually support decision-making. The field started out as Decision Support Systems (DSS) before evolving to include Business Intelligence, which now takes in Data Science to help businesses make better decisions. Key business drivers include compliance, operational support, and using insight to drive innovation, all centred on how data can improve business operations and outcomes.

Figure 4 Focus on “Governance”

Strategies for Developing an Enterprise Data Warehouses

Developing an effective Enterprise Data Warehouse (EDW) starts with the business goals and stays focused on the outcomes you actually want. That means understanding all potential data inputs globally first, then building the system incrementally, starting with a single subject and working up to data marts. It’s worth avoiding aggregated data imports from applications too, aiming instead to retain the lowest level of detail to support thorough archiving.

Streamlining applications pays off too, allowing for a lean design that offloads data into the EDW. Following data retention policies then lets data be archived and disposed of appropriately, often needing aggregation for efficient storage. A 2012 implementation of in-memory databases is a good example of this working well, significantly improving performance by extracting data from applications into memory and showing the EDW’s value as a strategic archiving solution.

Data Strategies and the Challenges in Self-Service Reporting

Aggregate data and self-service data strategies are advancing quickly, particularly around reporting, and a well-defined reporting strategy matters a great deal, especially as organisations lean more on self-service tools to build regulatory reports that can deliver real value.

Past experience as a BI developer showed just how challenging precise reporting can be, often requiring a return to SQL Server Reporting Services because other visualisation tools fell short. That frustration says a lot about how complex producing accurate, detailed reports actually is.

Data Governance and Implementation Strategies

Data governance makes the case for getting business acceptance before going live, keeping consistency and quality across logical data sets intact. User satisfaction is worth monitoring through Service Level Agreements (SLAs), particularly around the quality of data movement and ETL processes, since those matter a great deal for timely decision-making. A release roadmap works well when built around the MoSCoW Prioritisation method (Must-have, Should-have, Could-have, Won’t-have), giving structure to project phase planning. Additionally, key elements of implementation include configuration management, cultural change, and tools like metadata repositories and data integration techniques, alongside methodologies like prototyping and self-service Business Intelligence.

Figure 5 Focus on “Implementation”

Figure 6 Focus on “Technology”

The Key Concepts in Data Processing and Analysis

Howard spent some time on core concepts in data processing and warehousing, batch change data capture (CDC) and historical data integration among them. The distinction between Inmon’s Corporate Information Factory (CIF) and Kimball’s dimensional modelling approach highlights two different ways of building data warehouses, with Kimball favouring quicker implementations through fact tables and conformed dimensions. He also touched on analytical methods like OLAP (Online Analytical Processing), ROLAP (Relational OLAP), and MOLAP (Multidimensional OLAP), along with central components like staging, reference, Master Data, data marts, and operational data stores (ODS), all of which matter for effective BI and data analysis strategies.

Figure 7 Focus on “Load Processing”

Figure 8 Focus on “DW Architecture Concepts”

Figure 9 Focus on “Essential Concepts”

Understanding Data Warehousing Architecture and Big Data

The chapter on Big Data and Data Science stresses how much the underlying architecture keeps evolving. Data lakes are widespread, but they haven’t actually eliminated the need for centralised data warehouses, which is part of what’s driven the recent rise of the Data Lake House concept, aiming to combine the functionality of warehousing with data lakes. That approach standardises data transformation and keeps data moving efficiently from lake to warehouse, getting it ready for operational use quickly. The key takeaway is how much value Big Data insights bring when they actually feed back into data warehousing practices.

Figure 10 Big Data and Data Science Essential Components

The Impact of Data Privacy and Digital Decisioning in Organizations

A case surfaced roughly 3-4 months ago involving someone who claimed an organisation denied them credit based on a decision made by a machine learning model. Ten years on, that individual took legal action against the organisation, citing the negative impact the decision had on their life, and the organisation struggled to produce evidence around the rejection, since it hadn’t retained sufficient data on the decisions its model had made. In the end, the court awarded damages to the individual, a case that shows how much human oversight matters in high-stakes digital decision-making, and why better data retention and accountability matter for algorithms with this kind of influence over people’s lives.

The Abate Information Triangle and Data Integration Strategy

The Abate Information Triangle highlights the connections between data, information, knowledge, and wisdom, and why Master Data matters for validating relevant information. As organisations collect streaming and social networking data, photographs among the sources, it’s worth making sure that data actually pertains to the organisation itself rather than to competitors.

The process means understanding business needs, selecting the right data sources, and integrating Master Data to provide context, which supports data analysis and needs monitoring to keep error rates down before deployment, after which the resulting insights may prompt further data requests. A governance framework matters a great deal here too, for evaluating the relevance and provenance of the collected data and making sure it’s properly vetted to support informed decision-making.

Data Integration and Analysis

The updated DMBoK Version 2 Revised edition brings some significant changes, particularly its emphasis on an integrated data system that consolidates data from various sources, something the previous version didn’t mention. Where Version 1 focused on providing decision support, Version 2 makes data integration a central element in its own right.

The focus on Big Data has shifted too, from posing unknown questions at the outset of analysis toward handling large volumes of data and the statistical analysis that comes with the Data Science paradigm. That reflects a broader understanding of data complexity and analytical approach, moving beyond the original concepts of known and unknown data states to take in methodologies like confirmatory and exploratory analytics.

Figure 11 Definitions for Data Warehousing and Business Intelligence and Big Data and Data Science

Figure 12 Alteration of Definitions

Big Data and Data Science in Business Performance Measurement

Data Science, at its core, is about uncovering answers and insight from various data types. Business Intelligence (BI) focuses on analysing historical data to understand what happened and why, identifying changes in performance, an increase or decrease in sales, say, and working out what caused them.

Evaluating business performance really comes down to five key questions: what happened in each period, why it happened, what might happen if current trends continue, what actions should be taken, and what critical information might be missing. These conversations with executives matter a great deal for sharing insight effectively.

Figure 13 Business Drivers

Understanding and Managing Data Flow

The DMBoK Version 2 Revised edition lays out a structured approach to Big Data strategy, covering understanding business needs, establishing environments, selecting data sources, and data acquisition and ingestion. It also covers developing hypotheses, integrating and exploring data, and communicating insights through data storytelling and visualisation. One aspect worth real attention is model drift, which happens when real-world conditions diverge from the original model, potentially chipping away at its value and accuracy.

That’s why ongoing feature engineering and model performance monitoring matter to prevent rising error rates that can otherwise reduce the perceived value of data assets, sometimes called applying a “haircut” to their valuation. Data drift came up too, changes in the statistical properties of data over time that further affect how well a model performs.

Figure 14 Inputs, Activities & Deliverables (HOW)

Figure 15 Definition of “Data Drift”

Responsibilities of Data Manager in Big Data and Data Science and Data Warehousing and Business Intelligence

The Data Manager owns key areas across Data Warehousing and Big Data, operations planning, model performance monitoring, and enhancement strategies among them. In Big Data specifically, metrics like usage, loading speed, and data lake ingestion time matter for evaluating how data, large videos included, gets processed. Effective storytelling helps pull insight out of data, surfacing “aha moments” that drive genuinely valuable business reports.

Big Data operates in a much more dynamic environment than the fixed structure of Data Warehousing, drawing on various analytics to answer critical questions about business performance and decision-making. The progression from data-enabled to knowledge-driven capabilities is really a marker of an organisation’s AI maturity, with assessments determining what level of question the existing data can actually answer.

Figure 16 Data Managers Role and Responsibility

Figure 17 Metrics (How Many)

Figure 18 “Information to Insights to Decisions to Action”

Figure 19 DIKW Diagram

Data Vault and its Impact on Enterprise Data Warehousing

Howard pointed to Data Vault’s growing importance in enterprise Data Warehousing, especially given how little coverage it gets in the DMBoK Version 2 Revised edition. Despite that limited textbook coverage, it showed up prominently in exam questions anyway. Data Vault has effectively replaced the traditional relational model in this space, addressing issues like handling unstructured data and the constraints on speed and flexibility that the relational model imposed on data management. The shift to Data Vault, overall, marks a genuinely significant evolution in enterprise Data Warehousing.

Figure 20 “BUT, What about Data Vault?”

Enterprise Data Warehousing and Data Vault Models

The customer journey in data management starts with identifying potential clients as “suspects” the moment they visit a website. Once they express interest, they become “prospects.” Traditional relational models make data entry harder here, given strict business rules and cardinality requirements that often leave essential details, customer names or IDs, for instance, uncollected. The Data Vault model simplifies this considerably process by requiring only a customer code to access and link additional data as the customer progresses.

The Data Vault consists of raw and business data components, which improves both data quality and design by allowing comparisons across multiple data sources in the staging area. That supports effective Master Data modelling and gives a foundation for data quality assessment and trust rules without needing to go back to original sources, a genuinely significant advance for enterprise Data Warehousing.

Figure 21 Data Vault Architecture

Fourth Paradigm of Data Science

Scientific paradigms have evolved from the empirical approach, simply describing natural phenomena, through the theoretical paradigm using mathematical models, to the computational paradigm that simulated complex systems. The fourth paradigm, which Jim Gray and Alex Szalay termed “e-science,” integrates experimental data, archives, and simulations into a comprehensive Data Warehouse, letting data scientists explore and analyse massive datasets.

That combination of computer science, business domain knowledge, and mathematics is what gave rise to Data Science. Moving from multidisciplinary through interdisciplinary to transdisciplinary collaboration, Data Science builds a genuinely holistic understanding of complex systems by merging computational models, algorithms, and extensive metadata with diverse knowledge sources.

Figure 22 Kimball: Dimension Data Warehousing

Figure 23 Data Science Paradigm

Figure 24 eScience – A Transformed Scientific Method

Figure 25 Fourth Paradigm of Science

Figure 26 “X-Info & Comp-X for Discipline X”

Figure 27 Levels of Discipline Integration

The Key Elements and Process of Data Science

The Data Science paradigm covers several key elements, starting with data acquisition, then collection, cleaning, integration, and transformation. Once the data’s ready, exploratory analysis follows, using statistical methods, data visualisation, and hypothesis testing, which means defining a hypothesis and working out whether it’s confirmatory or exploratory, leading to predictions or repeated probes of that hypothesis. From there, the focus shifts to model building and training, feature engineering and model selection to identify core features, before the model gets trained, evaluated for quality, deployed, and monitored for performance drift. The process finishes with data interpretation, storytelling, and decision-making.

Figure 28 Key Elements to Data Science Paradigm

Data Management in Data Science

Data Science depends heavily on effective Data Management, particularly for simulations and experiments involving user behaviour and customer data platforms. Concepts like “bring your own lake” have emerged too, letting platforms like Snowflake create virtualised views of vast amounts of data.

Traditional statistical methods often fall short applied to single files, since the sheer volume of information in data lakes can overwhelm them. It helps to think of analysing this data as searching for needles in an ever-expanding haystack, where careful integration, alignment, and aggregation genuinely matter.

Successful data analysis also needs transdisciplinary integration, domain knowledge alongside mathematical expertise, since statisticians and mathematicians can’t really deliver value without a solid grasp of the context they’re working in.

Figure 29 (Data) Science Needs Data Management

Figure 30 Data Analysis

Data Management and Analysis

Scientists face real challenges in data delivery, particularly in fields like astronomy, where observatories worldwide contribute vast amounts of data, often reaching petabytes and terabytes. Traditional file-by-file analysis is proving inadequate here, searching a petabyte of data can take up to three years and cost around $1 million, which is why researchers are pushing for better data management through indexing, not unlike Google’s approach, to enable efficient search and analysis. The emphasis is shifting from analysing files to using databases, which is why transitioning from data lakes to data warehouses matters, a concept popularised by recent innovations like Inmon’s data lake house model.

Figure 31 Data Analysis Expanded

Figure 32 Data Delivery: Hitting a Wall

Intricacies of Data Management and Statistical Analysis

Statistical analysis means creating uniform samples and filtering relevant data subsets, relying on data management capabilities like data models and controlled vocabularies. Ensuring data quality through completeness, and addressing bad data, matters a great deal for handling data ethically, and Business Intelligence (BI) automates certain procedures to help understand data’s “What” and “Why.”

Data Science improves hypothesis testing and likelihood calculations through structured storage, and recent developments, Microsoft enabling R and Python routines to run directly in databases, for instance, boost processing efficiency by executing queries server-side. Effective data management skills matter for visualisations too, and for building scalable Big Data platforms, which covers capturing, curating, analysing, publishing, and peer reviewing data before granting access.

Importance and Challenges of Data Management

The costs of data platforms are genuinely significant, estimated at around $1,000,000, with schema, ontologies, and provenance being the most expensive areas in Data Science. As more data gets generated digitally through sensors, IoT devices, and simulators, effective data management matters more than ever.

A data catalogue, which data scientists often call a digital data library, plays a genuinely important role in organising these diverse datasets and their accompanying research, giving access to valuable insight from both small and large data sets alike. This comprehensive approach is aimed at improving the visibility and understanding of data assets, which ultimately supports better organisational decision-making.

Figure 33 Analysis and Databases

Figure 34 Analysis and Databases Expanded

Figure 35 Data Science Paradigm Elements

Figure 36 Data Science Paradigm Elements & Costs

Data Warehousing and Big Data Challenges

Howard pointed to Data Warehousing’s continued relevance even as Big Data has risen, stressing the challenges that come with inconsistent transformations within data lakes, where different interpretations of data attributes can lead to real chaos. Participants stressed how important discipline in data modelling and cataloguing is for keeping a reliable semblance of reality intact, something that’s diminished somewhat in modern data practices.

The lake house concept aims to standardise transformations from data lakes through ETL processes, giving a more consistent data view. Concrete examples from astronomy show how a common schema has enabled federated data querying across multiple observatories, while health informatics shows ontologies being used to integrate data from various hospitals through shared metadata standards. Additionally, Howard then addressed a question regarding the shift from traditional ETL (Extract, Transform, Load) processes toward automated, API-driven integration systems using ontologies for data normalisation across various genomic projects, a shift that allows diverse data sources to integrate seamlessly and supports real-time data curation and analysis.

Initiatives like the James Webb Telescope and the MeerKAT array are good examples of standardised ontologies compiling large datasets from global observational efforts, enabling timely queries and insight without the usual constraints of weather-related data availability. That represents real progress in the automated handling and interpretation of complex data across genomics and astrophysics.

Figure 37 Digital Data Library (Catalogue)

Figure 38 Use of ontologies in Astronomy

Figure 39 Use of ontologies in RNA Structure Genomics

Scroll to Top