Fuel for the Machine: The Semantic Foundation of AI with Howard Diesel

Key Takeaways

  • Adapting to Non-Deterministic Systems: Bringing AI in means moving data quality checks from deterministic rules to probabilistic assessment.
  • A Symbiotic Relationship: AI automates the tasks, and data management supplies the accuracy and context that keep bias and errors out.
  • Expanding Governance: Enterprise data governance needs to cover AI models and the metadata they generate, or systemic errors slip through.
  • The Necessity of Semantic Bridges: AI models need a semantic bridge to reach enterprise data with the right contextual meaning intact.
  • Human Oversight and Feedback Loops: Without context, AI generates errors, so a human needs to validate the output and feed corrections back in.
  • Independent Context Pipelines: Leaning on major data platforms risks vendor lock-in, so build your own context pipelines instead.
  • Data Protection Risks with Consumer AI: Unchecked AI tool use exposes sensitive corporate data to external models, and that’s a security risk most organisations underestimate.
  • Records Management and Accountability: Solid records management is what makes AI accountable, since it keeps the archival records behind a decision within reach.

Webinar Details

Title: Fuel for the Machine: The Semantic Foundation of AI with Howard Diesel
Date: 2026-07-20
Presenter: Howard Diesel
Meetup Group: DAMA SA Big Data
Write-up Author: Howard Diesel

How does AI Change Data Management Practices?

Bringing artificial intelligence into the business means shifting from traditional deterministic systems to non-deterministic data environments.

Data quality used to get evaluated inside deterministic systems, where the outcome was whatever the code explicitly dictated. AI doesn’t work that way. It operates non-deterministically, producing probabilistic responses that need a different kind of quality control and evaluation. Data management professionals now have to get ahead of that, making sure enterprise data is structured and semantically ready before an AI model ever touches it.

Organisations that want this to work need to use AI to improve data management, and at the same time apply solid data management principles to support AI development. It’s a two-way relationship, and it replaces reactive data handling with proactive semantic structuring.

Key Takeaways

  • AI models are non-deterministic systems, so they need dynamic data quality assessments, not static ones.
  • Traditional, code-based statistical process controls need to adapt if they’re going to evaluate AI-generated output properly.
  • Proactive data management is what keeps datasets properly contextualised for machine reasoning.

FAQ

  • What is a non-deterministic system in AI? A non-deterministic system produces probabilistic outcomes instead of relying on strictly coded, predictable rules.
  • How does AI change traditional data quality? Data quality now must evaluate AI-generated outputs and probabilistic reasoning, not just static deterministic rules.

Figure 1 Fuel for the Machine: The Semantic Foundation of AI

How Can Organisations Govern AI Artifacts Effectively?

Deploying AI effectively means governing AI artefacts, things like models, registries, and metadata, as rigorously as any other core enterprise asset.

AI generates metadata and content fast, and that speed brings its own risk: errors propagate just as fast. Use a large language model to help draft a corporate policy, for instance, and it might confidently invent a reference that doesn’t exist. AI makes a tireless digital assistant, but humans still have to resist the trap of trusting fluent output just because it sounds right.

The way to manage that risk is to break tasks into smaller, verifiable steps. Guide the AI through them sequentially, validate what it extracts at each stage, and you can use it safely for rapid content assessment and generation.

Key Takeaways

  • AI artefacts, model cards and business requirements included, need the same strict asset governance as anything else in the estate.
  • AI models make unverified assumptions more often than you’d think, and users need to actively challenge them.
  • Break a complex AI prompt into validated steps and the content that comes out is simply better.

FAQ

  • How should organisations govern AI content? Register AI models as assets and validate the metadata they generate. Then keep a human in the loop reviewing the output, that’s what stops errors from propagating.
  • Can AI write corporate policies? Yes, AI can speed up policy generation, but humans still need to validate factual accuracy so fictitious legal or historical references don’t slip through.

Figure 2 AI and Data Management Symbiosis

How do AI Models Automate Master Data Integration?

Trusted ontologies and semantic models let AI automate master data integration without anyone writing manual code.

When an external system hands over its data alongside a documented ontology, AI tools can use canonical models to process and transform that information automatically. That replaces manual data integration scripts, which are error-prone and hard to audit at the best of times. Building an enterprise-wide semantic model has always been difficult too, mostly down to complexity and limited resources.

Large language models speed up taxonomy creation considerably these days. AI can run zero-shot assessments and lexical analysis, then map synonyms to help shape a structured corporate vocabulary. Human validation still matters though; someone needs to check the accuracy of whatever taxonomic structure the AI produces.

Key Takeaways

  • Ontologies work as standardised blueprints, letting AI integrate diverse data sets automatically.
  • Large language models streamline taxonomy building by mapping synonyms and running lexical analysis fast.
  • Automated data transformation cuts down on manual, black-box Python integration scripts.

FAQ

  • What is a canonical model in data integration? A canonical model is a standardised data structure that, paired with ontologies, lets AI translate and map incoming data automatically.

How can We Ensure AI Models Improve Reliably?

AI systems need continuous contextual feedback and a trustworthy data foundation, otherwise hallucinations and silent errors creep in.

When an AI model can’t resolve a request, it usually escalates to a human operator. The human’s job there isn’t just to override the AI, it’s to explicitly teach the model why its decision got rejected, so that context carries forward. Skip that feedback loop and the model keeps relying on flawed nearest-neighbour vector retrievals, which means the same hallucinations keep happening.

A trustworthy foundation also demands watching for data drift constantly. Once production data starts deviating significantly from the training data, the model’s reliability drops off fast.

Key Takeaways

  • Human operators need to feed contextual corrections back into the model, or the same errors keep recurring.
  • Data drift happens when production data drifts away from training data, and that corrupts the model’s output.
  • A trustworthy foundation means knowing the specific data used to train the AI, in real depth, not just in outline.

FAQ

  • What causes AI hallucinations in data searches? Hallucinations often come from vector similarity searches. If the exact context is missing, the AI retrieves unrelated data and manipulates it anyway.
  • How do organisations handle AI data drift? Organisations need to track what data the AI model was trained on and watch for significant deviations once it hits live production data.

Should Organisations Build Independent Context Pipelines for AI?

Organisations should build their own context pipelines and guardrails to stay in control of AI logic and avoid vendor lock-in.

Feeding an AI accurate contextual fuel matters enormously. Get the specifications wrong and it generates irrelevant or incorrect output, no exceptions. To manage that, enterprises need to capture organisational context, decision flows and definitions among it, through dedicated context pipelines that sit outside the standard data pipeline.

Major data platform providers do embed AI harnesses and guardrails directly into their tools, but leaning on those locks an organisation into a single vendor’s methodology. Building proprietary context models keeps that door open, letting an organisation integrate multiple different AI tools and observability platforms without friction.

Key Takeaways

  • Get the contextual inputs wrong and AI models generate outputs that don’t match the situation at all.
  • Context pipelines capture tacit organisational knowledge and process flows, kept separate from data storage.
  • Proprietary context guardrails are what protect enterprises from getting locked into a single vendor.

FAQ

  • Why is vendor lock-in a risk for AI context? Relying exclusively on a vendor’s built-in AI guardrails restricts an organisation’s ability to adopt a multi-tool ecosystem or apply its own quality metrics.
  • What is a context pipeline? A context pipeline captures and manages business decisions, process changes, and organisational definitions, and keeps all of that separate from the raw data pipeline.

How does AI Improve Data Management Processes?

Artificial intelligence and data management share a symbiotic ecosystem. Each one enables the other’s strategic value across enterprise operations.

AI speeds up data management processes dramatically. It classifies sensitive information and generates data lineage without being asked, and it can turn natural language straight into executable data quality rules. In return, solid data management gives AI the authoritative grounding and semantic context it needs to avoid biased output.

Those efficiencies come with a catch: complex autonomous tasks like entity resolution still call for caution. AI is genuinely good at spotting non-deterministic relationships across huge document sets, but trust automated resolution blindly, without human oversight, and inaccurate metadata spreads across the enterprise fast. Augmented AI, with a human still in the loop, remains the safer bet.

Key Takeaways

  • AI accelerates data stewardship, automating metadata tagging and lineage tracking, along with security classifications.
  • Data management protects AI’s integrity by supplying enterprise data that’s unbiased and semantically rich.
  • Automated entity resolution needs strict human oversight, or the risk is large-scale data corruption.

FAQ

  • How does AI improve metadata management? AI reads schemas and runs data classifications fast, and it can map complex data lineage that would take a human months to trace by hand.
  • What is the risk of AI-driven entity resolution? Without master data integration and human oversight, AI can misidentify individuals or policies and inject widespread errors into the system.

Figure 3 The Symbiotic Ecosystem

Figure 4 The Fallacy of AI Independence

Figure 5 The Symbiosis Engine

Figure 6 Translating the Core: Foundations of Trust

Figure 7 Translating the Core: Operations & Analytics

How Do Semantic Bridges Enhance AI Data Access?

Deploying open-weight models and building semantic bridges lets organisations govern AI data access and evaluations securely.

To cut token costs and stop depending on centralised vendors, organisations are turning to open-weight models like Mistral and DeepSeek, which can be tailored to a specific operational context. To keep an eye on these models, developers build evals, evaluation datasets, into observability platforms, which dynamically assess an autonomous agent’s execution traces against defined business metrics.

AI agents shouldn’t be allowed to query raw enterprise databases directly. Organisations need to build a semantic bridge instead, a structured layer of context and meaning that acts as a secure intermediary between the AI model and the data underneath it.

Key Takeaways

  • Open-weight models offer a cost-effective, localised alternative to proprietary AI systems.
  • Dynamic evals help observability platforms monitor autonomous agent behaviour as it happens.
  • A semantic bridge stops AI from running flawed database queries by putting a managed contextual interface in the way.

FAQ

  • Why use an open-weight AI model? Open-weight models cut vendor lock-in and token costs, and they allow for localised data protection too.
  • What is a semantic bridge in AI? A semantic bridge is a contextual metadata layer that controls and interprets AI’s access to enterprise data, so there’s no direct, unguided database interaction.

Figure 8 The AI-Era Governance Perimeter

Figure 9 Evolving Quality: Fitness for Machine

Figure 10 The Velocity of Governance

Figure 11 The Semantic Bridge

How can Organisations Ensure Successful AI Deployments?

Successful AI deployments depend on harmonised corporate taxonomies and rigorous ethical feedback loops, without them errors just compound.

In a dynamic environment like enterprise marketing, AI systems need highly structured taxonomies. Assuming AI can accurately infer business definitions from raw text is a dangerous bet, terminology varies wildly from one department to the next. Standardise the glossaries, convert them into ontologies, load those into a knowledge graph, and organisations can be confident their AI models are reasoning with accurate semantics.

Organisations also tend to put off feedback loops until something critical breaks. When an AI-automated process makes an incorrect or biased decision, no audit trail means human reviewers can’t understand, let alone correct, the machine’s flawed logic. Building ethical feedback loops in from the start is what catches these issues before they compound.

Key Takeaways

  • Organisations need to actively harmonise their internal business glossaries before deploying AI, not after.
  • Knowledge graphs use taxonomies to keep AI agents reasoning with accurate context.
  • Feedback loops are what catch and correct AI errors before they cause real organisational harm.

FAQ

  • How do knowledge graphs improve AI reasoning? Knowledge graphs link every data node to a semantic metadata node, which gives the AI precise context for every piece of information it touches.
  • Why are AI feedback loops critical? Feedback loops let organisations audit AI decision-making, which means mistakes and biases get corrected systemically instead of one at a time.

How should Enterprises Ensure Trustworthy AI Governance?

Trustworthy AI needs enterprise governance that actually connects across the business. It also needs rigorous data security and meticulous records management alongside it.

The perimeter of enterprise governance has widened considerably, it now must cover data catalogues, AI models, autonomous agents, and the human operating frameworks around them. As employees reach for commercial tools like Microsoft Copilot more and more, organisations need to check their data protection configurations are actually holding. Without strict guardrails, uploading sensitive data, an HR spreadsheet, say, into a consumer AI platform can expose proprietary information to public training models without anyone noticing.

Traditional records management matters just as much for AI accountability. If an AI agent makes a consequential business decision, the enterprise needs to be able to pull up the exact archival records and metadata behind it. Lose that provenance and organisations are exposed to serious legal and operational liability.

Key Takeaways

  • AI governance must expand to cover human operating models and autonomous agent ones alike.
  • Upload corporate data to an unprotected AI tool and you’ve breached enterprise security and data privacy, full stop.
  • Rigorous records management is what gives automated AI decisions legal accountability.

FAQ

  • What are the risks of using consumer AI at work? Inputting corporate data into public AI models risks exposing proprietary and sensitive information to systems that have no business seeing it.
  • Why is records management important for AI? Corporate records provide the audit trail needed to validate and defend the automated decisions AI agents make.

Figure 12 The Seven Pillars of Trustworthy Enterprise AI

Figure 13 The Operational Foundation

Figure 14 ‘Data Protection when using Microsoft 366 Copilot Chat for work or school’

If you would like to join the discussion, please visit our community platform, the Data Professional Expedition.

Additionally, if you would like to watch the edited video on our YouTube please click here.

If you would like to be a guest speaker on a future webinar, kindly contact Debbie (social@modelwaresystems.com)

Don’t forget to join our exciting LinkedIn and Meetup data communities not to miss out!

Scroll to Top