What is this course about?
Enterprise AI pilots stall because the organization cannot agree what a customer is. This course covers entity resolution, semantic layers, and metadata as agent prerequisites, then argues for incremental delivery โ unify one domain, deploy one agent, repeat โ rather than a two-year modeling project that never ships.
Who is this course for?
- Data architects and chief data officers
- Enterprise architects
- Analytics engineering leads
- AI program leaders
What do I need before starting?
- Data modeling and warehouse experience
- Familiarity with your organization's core data domains
What will I be able to do afterwards?
- Diagnose whether a stalled AI program has a model problem or a data problem
- Perform entity resolution across systems that disagree
- Build a semantic layer encoding shared business definitions
- Establish metadata and lineage as agent prerequisites
- Deliver ontology work incrementally rather than as a multi-year program
What does each module cover?
Is it a model problem or a data problem?
45 minThe diagnostic, because teams reliably attribute data failures to model quality.
Objectives
- Distinguish model failures from data failures
- Diagnose a stalled pilot's actual blocker
- Communicate the diagnosis to a sponsor expecting a model answer
Topics
Activity. Diagnose a stalled pilot in your organization and identify its real blocker.
How do you resolve entities across systems?
60 minCustomer, employee, product, and account โ the same thing represented four incompatible ways.
Objectives
- Perform entity resolution across heterogeneous systems
- Handle identifiers that do not align
- Manage confidence and ambiguity in matches
Topics
Activity. Resolve one entity across three systems and quantify the residual ambiguity.
What is an 'active customer'?
55 minSemantic layers and the shared definitions two departments will otherwise never agree on.
Objectives
- Build a semantic layer encoding business definitions
- Resolve definitional conflicts between departments
- Version definitions as the business changes
Topics
Activity. Reconcile one contested definition between two departments and encode the result.
Why do agents need metadata and lineage?
50 minProvenance as an agent prerequisite rather than a data governance nicety.
Objectives
- Specify metadata an agent needs to answer responsibly
- Implement lineage sufficient to trace an answer to its source
- Expose freshness and reliability to the agent
Topics
Activity. Trace one agent-generated answer back to its source systems through lineage.
How do you do ontology work without a two-year project?
50 minIncremental design that ships value while the model is still incomplete.
Objectives
- Scope ontology work to a single domain
- Design for extension rather than completeness
- Ship value before the model is complete
Topics
Activity. Scope one domain's ontology to something deliverable in six weeks.
Who owns a definition when two departments disagree?
45 minGovernance for semantic conflict, which is a political problem with a technical surface.
Objectives
- Establish definitional ownership and escalation
- Allow legitimate departmental variation without fragmentation
- Resolve conflicts with a decision rather than a compromise
Topics
Activity. Run a definitional dispute to resolution with two departments represented.
How do you deploy an agent against the ontology?
55 minConnecting an agent to the semantic layer so it inherits the shared definitions.
Objectives
- Connect an agent to the semantic layer
- Verify the agent uses shared rather than ad-hoc definitions
- Handle queries the ontology does not cover
Topics
Activity. Deploy an agent against the semantic layer and verify definition inheritance.
Building one domain end to end
65 minThe build module: entity resolution, semantic layer, and one agent for a single domain.
Objectives
- Complete one domain end to end
- Deploy an agent that reads it
- Plan the next domain's increment
Topics
Activity. Complete the domain and deploy the agent, then plan increment two.
What is the capstone project?
One business domain unified with an agent reading it
Take one business domain end to end: entity resolution across the source systems, a semantic layer with reconciled definitions, lineage sufficient to trace answers, and a deployed agent that inherits the shared definitions.
Deliverable: A working domain ontology, a deployed agent, and a plan for the next increment.
How are learners assessed?
- Entity resolution assessed on residual ambiguity, honestly quantified
- Definitional conflict must be resolved by a decision, not a compromise that preserves both
- Agent verified to inherit rather than reinvent definitions
What ships with the course?
Facilitator guide
Session-by-session running order, discussion prompts, and the questions that reliably derail a room.
Learner workbook
Exercises, checklists, and the templates each module's activity produces.
Hands-on lab environment
A sandboxed ibl.ai deployment so exercises run against real agents, not screenshots.
Assessment bank
Scenario questions and rubric criteria mapped to each stated learning outcome.
Source bibliography
Every primary regulation and standard cited on this page, linked and dated.
Which AI agents does this course use?
The hands-on modules run against agents already deployable on the ibl.ai platform for enterprise.
Where does the course material come from?
Every module is grounded in primary sources โ the regulation, standard, or research itself, not a summary of it. Each was resolved at authoring time.
- AI Risk Management Framework
NIST
Data quality and provenance requirements in the Map function.
- Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks
Lewis et al., arXiv
How retrieval quality depends on the underlying knowledge representation.
- Ideas Made to Matter
MIT Sloan
Research on why enterprise data programs stall.
- Model Context Protocol
Anthropic
Integration layer connecting agents to the semantic layer in Module 7.
Delivery notes
Binding guidance for anyone preparing and delivering this course.
- Module 6 is a political problem wearing technical clothes. Run it as a real facilitated dispute with two departments represented, and require a decision โ a compromise that preserves both definitions is the failure mode.
- Module 5's incremental framing is the course's argument against the two-year enterprise ontology project. Do not soften it; those projects are why the audience is skeptical.
- Use the organization's genuinely contested definition. Every company has one โ active customer, qualified lead, headcount โ and the real one produces a better session than an example.
- Do not teach a specific ontology formalism as the answer. The value is in reconciled definitions and lineage; the representation is an implementation detail that will change.
- Coordinate with ENT-2 โ retrieval quality depends on this work, and the two courses should reference a shared example domain.
Why run AI training on a platform you own?
You own the course, not a licence to it
Course content, learner data, and the platform run inside your perimeter โ you own all the code and the data.
Model-agnostic delivery
Run the course's AI components on any LLM โ Claude, GPT, Llama, Gemini, Command โ and switch anytime.
No per-seat training licences
Usage-based or self-hosted, so cost tracks actual use rather than headcount.
Deploy anywhere
Cloud, private VPC, on-premise, or fully air-gapped โ including for cohorts that cannot use public AI tools.
Frequently asked questions
What does the Build the Data Ontology Before the Agents course cover?
Enterprise AI pilots stall because the organization cannot agree what a customer is. This course covers entity resolution, semantic layers, and metadata as agent prerequisites, then argues for incremental delivery โ unify one domain, deploy one agent, repeat โ rather than a two-year modeling project that never ships. It runs 6.5 hours across 8 modules across 8 modules, at advanced level, and closes with a capstone: One business domain unified with an agent reading it.
Who should take Build the Data Ontology Before the Agents?
It is written for Data architects and chief data officers, Enterprise architects, Analytics engineering leads, AI program leaders. Prerequisites: Data modeling and warehouse experience; Familiarity with your organization's core data domains.
Can we run this course on our own infrastructure?
Yes. ibl.ai is model-agnostic and deploy-anywhere โ cloud, private VPC, on-premise, or fully air-gapped โ and you own all the code and the data. Cohort data, submissions, and any material learners upload stay inside your perimeter, which matters for enterprise teams that cannot send work to a public AI tool.
How do we get access to Build the Data Ontology Before the Agents?
Request access and we will set it up for your cohort โ hosted by ibl.ai, or running against your own deployment. Tell us the group size and timing you need, and whether it should run inside your own perimeter.
How much does AI training for enterprise cost on ibl.ai?
There is no per-seat pricing โ you pay for usage or self-host and pay only for the infrastructure, so a 5,000-person rollout does not cost 5,000 licences. 1.6M+ users across 400+ organizations run the platform this way, including NVIDIA, MIT, and Syracuse University.