Executive Summary
This webinar covers the roles data professionals play in Business Intelligence (BI) and Big Data analysis, and makes the case for optimising analytical techniques and building a genuinely data-driven organisational culture. Howard Diesel works through why Data Warehousing matters and how it integrates with AI for effective business decision support, while examining how the data management landscape keeps evolving and shaping innovation.
Webinar Details
Title: Data Warehousing, BI, Big Data and Data Science for Data Citizens
Date: 05 December 2024
Presenter: Howard Diesel
Meetup Group: African Data Management Community
Write-up Author: Howard Diesel
Data Professionals in Business Intelligence and Big Data Analysis
Howard Diesel opened the webinar by noting this instalment is part of a Data Warehousing, BI, and Big Data training course, with the focus on how Business Intelligence and Big Data provide insight for decision-making. He also discussed the evolution of Data Science, referencing it as the “4th paradigm of science,” which uses data for hypothesis development and validation.
The presentation will use the integration of astronomical data from global observatories as an example, along with introducing a data citizen approach, and Howard again notes the focus on the roles data professionals play, BI developers, data engineers, and machine learning engineers among them.
Figure 1 How-To Analyse
Role of Analytics in Business Intelligence
The focus here is analytical maturity, and Howard worked through the various analytical methods used in Business Intelligence (BI), distinguishing descriptive analytics, which provides insight into past events, from diagnostic analytics, which answers what happened and why. That foundation supports the move to predictive analytics, forecasting future outcomes from previous data trends, budget performance, for instance. Key techniques here include statistical analysis, predictive modelling, and multivariable statistics, all helping organisations anticipate future scenarios and make informed decisions.
Figure 2 BIA DevOps
Figure 3 Data Management Maturity Growth Diagram
Figure 4 How-To Analyse Data (SIPOC)
Figure 5 Data Warehousing & Business Intelligence
Optimising Data Analysis Techniques
Howard pointed to how important it is to understand outcomes and optimise them through operations research, a key part of prescriptive analytics, which answers the question, “What should I do?” Prescriptive analytics draws on techniques like knowledge graphs for inferences and neural networks, and cognitive analytics matters too, for identifying unknowns and raising awareness of overlooked factors.
Analysts should aim to provide a range of options rather than a single answer, using optimisation to weigh pros and cons and inform decision-making. Howard referenced the DIKWA (Data, Information, Knowledge, Wisdom, Action) model here, which underlines why enhancing data maturity matters for answering more complex questions effectively.
Addressing Data-Disability in Organisational Culture
Organisational culture can shape a company’s shift toward being data-driven or knowledge-driven a great deal. A common challenge here is employees exhibiting a “data disabled” mindset, not trusting the quality of data presented through graphs and information, which is exactly why change managers need to get involved to address that scepticism and guide employees from resistance toward becoming “data-enabled.” Improving Data Quality, ensuring reliable dashboards, and building trust in BI reports is what lets organisations foster a more positive relationship with data, which leads to informed decision-making and a better grasp of underlying trends.
The Transition to Data-Driven Thinking
Moving to a data-driven approach means relying on accurate forecasting and predictions to guide decisions, trusting the data over subjective analysis or outside consultation. Understanding where you stand on the analysis continuum matters a great deal here, since it helps identify data sophistication, stakeholder positions, and the relevant business questions.
Building a strong data culture takes time and consistent effort, it’s not an overnight shift. A practical example shows how much trust and accountability matter in project management: fail to deliver as promised, and that undermines your reputation and invites closer scrutiny from stakeholders down the line. Demonstrating reliability over time is really what rebuilds that trust.
Understanding Data Warehousing and Big Data
In Data Management, business decisions need to rest on accurate data analysis, since incorrect interpretations can cause real setbacks. Effective data management takes a comprehensive understanding of assets, maturity, and measurement continuums to streamline processes.
The DMBoK (Data Management Body of Knowledge) framework differentiates between Data Warehousing, Business Intelligence (BI), and Big Data analytics, pointing to the need for rigorous methods to derive insight from large datasets. That includes moving from Version 2, which focuses on uncovering unknown questions, toward a more structured approach integrating the various stages of data processing. Data Managers establishing a solid foundation, in the end, is what frees up teams to focus on the analytical work that actually drives informed decisions.
Figure 6 DW/BI Vs Big Data Process Comparison
Data Warehousing and AI in Business Decision Support
The focus then moved to integrating Business Intelligence (BI) and Artificial Intelligence (AI) to support decision-making and empower knowledge workers in data analysis. Howard noted that an effective platform should consolidate operational data from various business processes to give a comprehensive overview, particularly useful when investigating declining revenue. BI and Data Warehousing can be systematically structured and controlled, giving clearer roadmaps and user acceptance testing, but Big Data projects tend to be less predictable and need more time to explore and validate hypotheses. Challenges in Big Data include potential delays in model deployment and the need for mature, accurate models to avoid false positives and negatives, which can frustrate users when expectations don’t get met.
Figure 7 Data Warehousing & Business Intelligence Process
Figure 8 Understand Requirements
Common Aspects of Data Management and Business Innovation
Both Business Intelligence and data analysis share common goals, mainly around fulfilling business requirements by posing the relevant questions. That means gathering and processing data across various structures, and building pilot data products to uncover valuable insight. These approaches can also generate ideas for business innovation, which shows that insight comes not just from Big Data but from a solid understanding of business processes too. Both methodologies also need deployment and monitoring of data operations, often called DevOps or ML Ops, to keep data and machine learning models deploying efficiently.
Understanding the Differences of Data Warehousing and Big Data
Data Management’s evolution is marked by the emergence of the Data Lake House, integrating a data warehouse framework on top of a Data Lake to standardise transformations across data scientists, a departure from traditional Big Data approaches. This model aims for a common data source for BI while still respecting the unique aspects of Big Data environments, preparing training datasets, for instance, and developing hypotheses from customer responses to refine decision-making through probing and sensing.
Integrating and aligning disparate data sources, often external ones, also takes a thorough understanding of metadata and master data, which sets it apart from conventional internal systems like ERP. Data Warehousing processes are well-defined with predictable outcomes, but Data Science brings real uncertainty around data completeness and the challenge of establishing causation versus correlation, which is where a lot of the complexity in deriving insight lies.
Figure 9 Define & Maintain DW/BI Architecture
The BI Process and its Importance in Business Performance Measurement
The ABI (Analytics Business Intelligence) process means understanding business goals, strategy, and performance to measure metrics like budget versus actual results effectively. BI excels at performance measurement but often struggles to explain the reasons behind trends, declining sales or customer retention, for instance. Addressing that means identifying stakeholders for each KPI and prioritising data products against business requirements, building a comprehensive requirements framework, and establishing a proper architecture with elements like data lineage and data catalogues. Additionally, keeping all metadata readily accessible is what lets business users trace report data back to its origins and seek clarification when needed, which strengthens data asset management overall.
Figure 10 DW/BI Technical Architecture
Figure 11 Define DW/BI Management Processes
Evolution and Impact of Data Warehouse Models
The Data Vault approach marks a shift away from traditional dimensional modelling and data warehouses toward a more flexible framework for data integration. Unlike classical models, which enforce strict quality measures that can inadvertently exclude valuable data, Data Vaults use a normalised data structure built around core business concepts and their relationships, creating “hubs” to capture essential identifiers like customer codes and “satellites” to store descriptive attributes. That structure allows for varied data sources, marketing interactions with customers who may not have complete information, for example, supporting a more inclusive Enterprise Data Warehouse that keeps analytics effective without compromising Data Quality.
Figure 12 Develop the Data Warehouse & Marts
Understanding the Dynamics of Data Warehousing and Integration
Data Warehousing’s management process typically runs across three main tracks: Data Architecture, technology requirements, and BI tools for data analysis. Key components include data integration, ETL processes, Data Quality, and Metadata Management, with a growing emphasis on metadata-driven integration, particularly using ontologies to simplify transformation rules.
Understanding the characteristics of data sources, structure, format, accuracy, matters for integrating diverse data effectively, and user needs matter too when designing reporting tools, since different stakeholders, executives, for example, might prefer high-level PDF dashboards over interactive BI tools.
The Gartner analytics framework underlines how much people, processes, platforms, and metadata matter for driving various analytics types, and effective data product management follows a release process guided by prioritised use cases, keeping things aligned with business needs and Data Strategy.
Figure 13 Taxonomy of Data Sources
Figure 14 Data Source Taxonomy Use-Cases
Figure 15 Source-To-Target Data Element Taxonomy
Figure 16 Populate the Data Warehouse
Figure 17 Implementing BI Portfolio
Figure 18 Applying Gartner Analytics Framework
Figure 19 Maintain Data Products
Figure 20 Release Process
Figure 21 Big Data & Data Science
Data Science and AI
The Data Science process means managing large amounts of data through platforms like cloud services and orchestration pipelines, Snowflake among them. A key distinction is between shallow AI, developing models on limited data, and deeper analyses that need extensive datasets, astronomical information, for instance. The goal is extracting insight and answers from data, even where initial findings only reveal correlations rather than causation, which is why ongoing hypothesis development and data acquisition matter for deepening understanding.
Communicating complex data insight effectively matters a great deal, using methods like data storytelling and visualisation. The process runs through identifying needs, selecting appropriate data sources, exploring and analysing data, and refining models until they reach maturity and reliability.
Figure 22 Data Science Process
Figure 23 Define Big Data Strategy & Business Needs
The Implications of Synthetic Data in Machine Learning
Synthetic data has gained real traction as organisations run into bandwidth limitations training machine learning models. It’s designed to replicate the original dataset and generate the volume needed for training, but it risks reinforcing existing patterns and biases rather than introducing genuine diversity. Additionally, using this data well takes robust governance, assessing data sources, identifying and mitigating bias, and understanding what various features in the dataset actually imply.
Using metadata to understand relationships between data columns, and building data trust evaluations around foundational criteria like granularity and reliability, matter a great deal too. Ethical considerations around privacy and re-identification risk remain genuine concerns as well, since current anonymisation methods often fall short of guaranteeing complete security against re-identification attempts.
Figure 24 Choose Data Sources
Figure 25 Basic Metadata (Facts about Data)
Figure 26 Data Source Evaluation (Data Trust Rules)
Figure 27 Choosing Data Sources Associated Risk
Figure 28 Choosing Big Data
Data Sources and Data Governance in Business
Choosing data sources means weighing various factors: the type of data (internal web data, synthetic human-generated data, machine-generated biometric data), along with provenance, frequency, hardware requirements, and the governance surrounding data acquisition. It’s worth avoiding the common pitfall of hastily bringing in new data without properly assessing whether it’s appropriate. Industries are increasingly aligning specific use cases with suitable data types instead, sensor server logs, social and geographic data, clickstream data, and structured and unstructured user engagement metrics among them. Additionally, that structured approach supports effective Data Governance and lets data leaders identify and access valuable data more efficiently.
Figure 29 Big Data Type According to Business Needs
Statistical Models and Data Integration in Data Analysis
Developing hypotheses and statistical models means understanding various states of information, which can be categorised into four quadrants: known questions with known answers, known questions with unknown answers, unknown questions with known answers, and unknown questions with unknown answers.
Dealing with unknown questions and answers often means working with genuinely chaotic data, which takes more time to resolve. Data integration matters a great deal here, combining information from multiple datasets, article titles and authors, or paper titles, authors, and journals, say, using techniques like joining or merging. That kind of integration deepens understanding by leveraging ontologies and constraints that describe the relationships within the data, which is where a lot of the field’s exciting advances are coming from.
Figure 30 Develop Data Hypotheses & Methods
Figure 31 Data Merging in Data Integration Systems
The Intersection of Statistical Modelling and Machine Learning
In predictive modelling, statistical modelling and machine learning models are worth distinguishing. Statistical modelling tends to focus on prediction through approximation, while machine learning uses algorithms to analyse decision-making probabilities. Key to both is carefully selecting training and testing samples from the available dataset, which can run to millions of records, a 50/50 or 80/20 training-to-testing ratio, for instance, might be used to make sure the model trains adequately for reliable outcomes. Dimensionality and feature reduction techniques can streamline the process by minimising the data needed for training, and once trained, the model gets validated against test data, then optimised and refined as new data comes in.
Figure 32 Exploring Data Using Models
Figure 33 Predictive Process: Step 5.1 Random Sampling
Figure 34 Predictive Process: Step 6.2 Build/Develop/Train Models
Monitoring Machine Learning Models in Data Management
Data professionals need to understand the usage and processes involved in determining when a model is actually ready for production and confirming its accuracy. Deployment isn’t the finish line, continuous monitoring of outcomes and error rates matters just as much, and if error rates climb, the model may need to go back to a test environment for troubleshooting, potentially retraining with additional features.
Assessing model performance through true negatives, false positives, false negatives, and true positives matters a great deal for keeping it effective. Various visualisation techniques help illustrate time series data, expectations, and relationships, giving deeper insight into model performance and informing decisions about whether to keep using it or retire it.
Figure 35 ROC Curve
Figure 36 Predictive Process
Figure 37 Explore Data Using Models
Data Deployment and Monitoring in Data Warehousing and Machine Learning
In deploying and monitoring models, stages get identified as blue, red, and green models: the green model is typically the one in production, while blue and red are in testing. That framework allows comparisons between the different models, supporting performance evaluation by analysing each one’s outputs. The DMBoK framework aims to clarify the distinctions between areas like Data Warehousing, BI, and machine learning Data Science, and the processes each one involves.
Figure 38 Deploy & Monitor
Figure 39 Crucial Data Executive Questions
The Importance of Ontology in Data Mapping and Integration
The discussion pointed to how central ontologies are for consolidating diverse data from various medical laboratories participating in the genome project. With so many research datasets collected in different formats, integrating that information into a standardised layout matters for efficiency and effectiveness, and as data sources multiply, relying on metadata and a machine-readable ontology becomes vital, enabling seamless transformations and mappings between datasets.
This approach addresses the challenges of unstructured data while also giving a unified view of research findings, illustrated well by collaborative efforts in astronomy to synthesise observations from hundreds of observatories worldwide. Establishing a coherent mapping from source to target data, in the end, is what advances research and gives clear insight into complex data landscapes.
Data Management and Analysis in Healthcare and Smart Cities
Data Warehousing and integration take a structured management process centred on data architecture, technology requirements, and Business Intelligence tools, with key components like data integration, ETL processes, Data Quality, and Metadata Management. Understanding data source characteristics matters for effective integration, and the Gartner analytics framework highlights the role of people, processes, platforms, and metadata.
In Data Science, the focus is managing large datasets to extract insight, with ongoing hypothesis development and data storytelling front and centre. The rise of synthetic data for machine learning training calls for robust governance to mitigate bias and meet ethical considerations, and choosing appropriate data sources takes careful assessment aligned with specific use cases to strengthen Data Governance. The intersection of statistical modelling and machine learning is worth distinguishing too, approximation-focused statistical models versus algorithm-driven predictive analysis, which is driving real advances in understanding complex datasets.
Managing Machine Learning Models in Changing Data Scenarios
With transitional data that changes frequently, monitoring machine learning (ML) models for drift matters a great deal, since drift signals that the relationships between features may be shifting in the real world. That drift can affect a model’s accuracy, which means reviewing parameters and potentially updating the model. Assessing how new data affects the current model matters too, for determining whether it needs retraining or can keep functioning effectively, and researchers recommend exploring techniques to measure and detect ML drift to manage that process well.
Feature and Dimension Reduction in Model Building
Building predictive models takes key steps like dimension reduction and feature engineering, since including too many features complicates training and multiplies the challenge of handling permutations. Identifying and isolating the multivariate relationships among features while removing the extraneous ones matters for getting a clearer read on model performance, and along the way, false positives and false negatives can increase, which means reassessing the relationships among features so significant factors don’t get overlooked. Howard closed with a recommendation for continuous monitoring and re-evaluation to keep the model effective.
- Executive Summary
- Data Professionals in Business Intelligence and Big Data Analysis
- Role of Analytics in Business Intelligence
- Optimising Data Analysis Techniques
- Addressing Data-Disability in Organisational Culture
- The Transition to Data-Driven Thinking
- Understanding Data Warehousing and Big Data
- Data Warehousing and AI in Business Decision Support
- Common Aspects of Data Management and Business Innovation
- Understanding the Differences of Data Warehousing and Big Data
- The BI Process and its Importance in Business Performance Measurement
- Evolution and Impact of Data Warehouse Models
- Understanding the Dynamics of Data Warehousing and Integration
- Data Science and AI
- The Implications of Synthetic Data in Machine Learning
- Data Sources and Data Governance in Business
- Statistical Models and Data Integration in Data Analysis
- The Intersection of Statistical Modelling and Machine Learning
- Monitoring Machine Learning Models in Data Management
- Data Deployment and Monitoring in Data Warehousing and Machine Learning
- The Importance of Ontology in Data Mapping and Integration
- Data Management and Analysis in Healthcare and Smart Cities
- Managing Machine Learning Models in Changing Data Scenarios
- Feature and Dimension Reduction in Model Building