Winning the Lexical, Semantic, Ontic Trifecta with John O’Gorman

Key Takeaways

  • Guiding AI to prevent hallucinations: Curate semantically structured data before it reaches an LLM and the output gets more reliable and much easier to explain.
  • Prioritising “strings” before “things”: Real semantic interoperability starts with the words themselves. Manage the lexical layer carefully before building it into enterprise ontologies, or communication breaks down downstream.
  • Creating atomic, reusable information (Quants): Break knowledge into quants, small reusable units of information that can be simple or complex depending on what they describe.
  • Strict, inherent classification: Give each lexical term one category, not several. That’s what keeps polysemy out and definitions clear.
  • Improving PII management through upstream disambiguation: Sort out homonyms and track data lineage early, and PII management gets simpler and safer down the line.
  • Driving adoption through “ego death” and practical language: Content creators need to let go of their own phrasing and lean on standardised vocabulary and metaphors people actually relate to, or stakeholders switch off.

Webinar Details

Title: Winning the Lexical, Semantic, Ontic Trifecta with John O’Gorman
Date: 2026-07-06
Presenter: John O’Gorman
Meetup Group: INs and OUTs of Data Modelling
Write-up Author: Howard Diesel

How can LLMs be Guided for Reliability?

LLMs drift off-topic or hallucinate unless you give them strict lexical and semantic guidance. Without that structure, generative AI runs on probability instead of certainty, and it will confidently make up a wrong answer without any warning sign.

Give an AI agent precise data and a solid semantic foundation, and it stays on track. A properly guided LLM stops being an unpredictable text generator and starts being a genuinely useful, reliable part of the business.

Key Takeaways

  • LLMs fail silently without strict contextual guidelines to keep them in line.
  • Structuring data upfront is what actually prevents AI hallucinations.

FAQ

  • Why do LLMs drift or hallucinate? Take away precise lower-level data and guidance, and an LLM has nothing to anchor to. It starts guessing, and guessing is where hallucination comes from.
  • How can businesses improve LLM accuracy? Give AI tools a solid lexical and semantic foundation and their outputs stay deterministic and reliable instead of drifting.

How does Lexical Alignment Achieve Semantic Interoperability?

O’Gorman calls this the lexical, semantic, and economic trifecta, and treats semantic interoperability the way you’d treat a supply chain. A grocery store can’t function without properly labelled stock arriving at the door, and business systems can’t build reliable taxonomies and ontologies without properly classified lexical strings arriving first.

Ontologies have traditionally focused on things (the real-world entities) rather than strings (the words businesses actually use to describe them). Full semantic interoperability needs lexical alignment to happen upstream, well before anyone starts structuring those words into complex ontologies.

Connect the language of the business (lexical) to logical meaning (semantic) and to real-world entities (ontic), and you get a knowledge bridge that actually holds together.

Key Takeaways

  • Treat semantic data modelling like a supply chain: categorise the inputs strictly before anyone uses them.
  • Lexical alignment must happen upstream, before terms get built into enterprise ontologies.

FAQ

  • What is the difference between lexical and ontic data? Lexical data is the strings, the actual language a business uses. Ontic data is the real-world physical or digital entities those strings point to.

Figure 1 Winning the Trifecta: Lexical, Semantic and Ontic

Figure 2 “What Started All of This”

Figure 3 Strategic SOCs

Figure 4 Bridges Create Complete Interoperability

Figure 5 Semantium’s Proposal

What is a “Quant” in Information Organisation?

A quant is a self-contained, atomic unit of information built for reuse without alteration. It can be as small as a single UI descriptor or as large as an entire process document, as long as it stays standalone and genuinely reusable.

To keep quants organised, the Semantium protocol sorts them into 19 mutually exclusive facets. Those facets map onto the natural dimensions of human communication: who, what, when, where, why, and how.

Categorising atomic units upstream standardises metadata and tracks lineage, so terms keep their fundamental meaning before they ever reach a knowledge graph or a complex model.

Key Takeaways

  • A quant is an atomic, self-contained piece of information built to be reused.
  • Quants get sorted into mutually exclusive facets based on natural language dimensions.
  • Categorise upstream and you avoid the edge-explosion that hits knowledge graphs further down the line.

FAQ

  • What makes information “atomic”? Information counts as atomic when it’s fully self-contained and can be reused repeatedly without needing any structural change.

Figure 6 Where Semantium Protocols “Fit”

Figure 7 Why this Separation Works

Figure 8 Semantium’s “Protocols-on-a-page”

What Defines Digital Templates as Atomic Information Units?

Digital templates count as quants because reusability is their whole point. A template on its own is an abstract, functional concept, but a specific template used for meeting minutes becomes an actionable digital asset.

Managing information at this level means tracking lineage strictly. Organisations need to record a quant’s original source alongside its contributor, meaning whoever or whatever last modified it.

If an original term changes (say, a specific qualifier gets added to an existing medical diagnosis), a new quant must be registered so the lineage and integrity of the original terminology survives.

Key Takeaways

  • Templates count as quants because they’re self-contained and genuinely reusable.
  • Accurate lineage tracking means logging the original source of a digital asset and every contributor who touched it afterward.

FAQ

  • Can a large document be considered a quant? Yes. A quant can be as large as a full document or template, as long as it’s reusable without any structural changes.

Figure 9 Semantium’s “Protocols-on-a-page” pt.2

What is Robust Semantic Architecture’s Primary Rule?

Solid semantic architecture comes down to a strict rule of engagement: a unique lexical term belongs to exactly one foundational classification category, or facet, and nothing else.

Classify terms by what they inherently are, not by how someone might use them later, and you keep the data technology-agnostic and flexible. Take the Utrecht train station: it first gets classified as a physical thing. Downstream systems can then treat it as a location, as long as its root physical identity stays intact.

Every term needs to be fully registered with its attributes before it enters an operational environment.

Key Takeaways

  • Classify terms by their known, inherent nature, not by how you expect to use them later.
  • Each unique lexical value gets assigned to one foundational facet, never more than one.
  • Get the upstream classification right and enterprise search starts working with the precision of GPS coordinates.

FAQ

  • Why should a lexical term only have one facet? One root facet per term prevents polysemy (multiple meanings) and gives downstream systems a single, definitive source of truth to work from.

Figure 10 Why these Protocols Matter

Figure 11 Clean SOCs = Complementary Deliverables

Figure 12 Semantium Protocols RoE

How can we Enforce Standardised Terminology in Organisations?

Strict semantic standards usually run into friction, mostly from users who’d rather invent their own terminology than follow the rules. Getting past that friction takes accessible framing and a hard line on content reuse.

O’Gorman leans on what he calls the sock drawer analogy to teach semantic classification to non-technical staff, and it works. Just like you’d find a clean, white, knee-length sock by filtering on physical attributes, or facets, business users can find digital information by categorising data the same systematic way.

Organisations also need to enforce what O’Gorman calls ego death among content creators. Writers give up the urge to draft their own definitions and stick to standardised, pre-approved corporate vocabulary instead.

Key Takeaways

  • Relatable metaphors, like sorting a sock drawer by attributes, help non-technical users grasp abstract data classification.
  • Corporate consistency means content creators put vocabulary reuse ahead of writing something original.

FAQ

  • What is “ego death” in content management? It’s the process where content creators drop their personal stylistic preferences and reuse established, standardised corporate terminology instead, which keeps the data clean.

How can we Disambiguate Words with Multiple Meanings?

Words with multiple meanings, like “bank,” which can mean a financial institution, an aircraft manoeuvre, or a river edge, need to be disambiguated at the upstream lexical level. Assign each distinct word sense its own unique identifier, and systems can produce semantically equivalent translations across platforms.

Separating lexical terms (strings) from ontological entities (real-world things) simplifies Personally Identifiable Information management considerably. Store a person’s name variations on the lexical side, link them by a unique ID to a secure ontic profile (birthdate, nationality, and so on), and organisations decouple sensitive identities from unstructured data.

Key Takeaways

  • Homonyms, words with identical spellings but different meanings, need to be separated with unique IDs upstream.
  • Separating lexical string names from real-world entity profiles makes for a genuinely secure architecture for PII management.

FAQ

  • How does semantic separation improve PII management? Keep name variations (lexicals) separate from actual person records (ontics) and bridge them with unique IDs. Sensitive identities stay abstracted and secure that way.

Figure 13 Rule #1 is the Most Challenging

Figure 14 Rule #2 is also Challenging

Figure 15 Translation Equivalence: Lexical / Semantic

Figure 16 Glossary Search: Lexical / Semantic

Figure 17 Ontological Properties: Semantic / Ontic

Figure 18 Entity Resolution: Lexical / Ontic

Figure 19 How Semantium Protocols affect Models

Figure 20 Where Semantium Fits Analytically

Figure 21 Traditional Vs. Semantium Analytics

Figure 22 Anticipated Benefits

How do LLMs Enhance Semantic Data Processing?

Pair LLMs with strict semantic protocols and AI shifts from a probabilistic guessing engine into a deterministic, explainable system.

LLMs naturally treat words as nothing more than statistical tokens. Semantic curation gets ahead of that by categorising tokens upfront and stripping out ambiguity before the LLM ever processes them. Feed an LLM semantically clustered, faceted data and its outputs start relying on established semantic equivalence instead of raw statistical similarity.

That combination gives you explicit, analysable dimensions to work with, so businesses can trust AI outputs and slice organisational data without worrying about hallucinations creeping in.

Key Takeaways

  • Unchecked LLMs run on statistical probability. Semantic protocols replace that with deterministic, explainable facts.
  • Feed semantically equivalent data into an LLM and the guesswork disappears, along with a lot of the explainability problem.

FAQ

  • How do semantic protocols improve Large Language Models? They curate and categorise words before the LLM ever sees them, which turns probabilistic token generation into something deterministic and reliable.

Figure 23 How Do LLMs Fit Here?

Figure 24 Addenda: Referents in Semantium

How can we Align Business Language for Success?

A semantic or taxonomy initiative only succeeds if it delivers business value right away, without burying stakeholders in dense library science jargon.

Getting executive and user buy-in means speaking the business’s native language, not forcing academic terminology onto people who didn’t ask for it. Teams don’t need to define corporate terms from scratch either. Starting from an authoritative online dictionary sets a baseline and heads off a lot of pointless debate and duplicated work.

Strict upstream vocabulary controls end up cutting backend integration costs substantially, mostly because domain owners are forced to collaborate and standardise their language across the enterprise.

Key Takeaways

  • Skip the academic terminology. Speak the business’s practical language if you want stakeholders to adopt it.
  • Baseline definitions against established dictionaries so teams stop inventing their own arbitrary terminology.

FAQ

  • How should organisations begin defining business terminology? Start with authoritative sources, like online dictionaries, to set a baseline, and only modify a definition when the specific business context genuinely requires it.

If you would like to join the discussion, please visit our community platform, the Data Professional Expedition.

Additionally, if you would like to watch the edited video on our YouTube please click here.

If you would like to be a guest speaker on a future webinar, kindly contact Debbie (social@modelwaresystems.com)

Don’t forget to join our exciting LinkedIn and Meetup data communities not to miss out!

Scroll to Top