Skip to content

8 minute read

What Boston's world cup beer shortage teaches us about context debt

by Datavid on

Context debt is the gap between your data and the business meaning AI agents need. See the warning signs and how to pay it down with a company brain.

Table of contents

Quick Answer:

Context debt is the gap between the business knowledge decisions that depend on and what systems can connect and explain. It grows when definitions and rules remain scattered across systems, documents, and people’s heads. AI agents expose this debt by acting on incomplete context. Organizations can reduce it by making the meaning of business, relationships, and rules explicit through a governed context layer.

In June 2026, Scotland's World Cup fans drank some of Boston's best-known bars dry. The signals were public. Scotland had qualified in November 2025, and Tennent's had stocked 80 Boston-area bars before kickoff. What nobody had was a connected view of fans, venues, and stock. That gap has a name: context debt.

The story was shared by Datavid CEO Balvinder Dang during his session at the Progress Data Platform Summit 2026, where he explored the trust layer needed for enterprise AI. The question was simple: if an AI agent had been responsible for preventing Boston's shortage, what would it have needed to know?

This article is for CDOs and CIOs moving AI agents from pilot to production. It shows how to spot context debt and how Datavid has helped organizations like Syngenta and Roche pay it down.

At a glance

  • Context debt builds up when meaning and business rules stay scattered across systems and people.
  • Enterprise AI agents that move inventory or approve decisions can turn missing context into financial and compliance losses.
  • Knowledge graphs and GraphRAG add the relationships that standard RAG often misses.
  • Many organizations can pay down the debt one decision at a time without replacing their data platforms.

What is context debt (and why you're already paying it)

Context debt is the accumulated gap between the business knowledge your decisions depend on and what your systems can connect to, explain, and hand off to an AI model.  A metric with three competing definitions adds to it. So does a rule that lives only in one analyst's head. For data leaders, the balance tends to stay hidden until someone asks a question that crosses systems.

Context debt vs. technical debt

The idea borrows from technical debt, where a shortcut to ship faster works like a loan. Google researchers applied it to machine learning in their paper on hidden technical debt in machine learning systems. They warned that this debt compounds silently and named changes in the external world as a risk factor. The Boston example shows how quickly an external change can expose gaps in an organization's existing assumptions.

Context debt sits a level above code, in the meaning between your data. A recent Forrester analysis of enterprise AI bills directly links cost to AI. When company knowledge isn't machine-readable, agents rebuild the missing meaning on each call, and you pay for it in tokens.

Aspect

Technical debt

Context debt

Where it lives

Code, architecture, and infrastructure

Definitions, relationships, and rules across systems

How it builds up

Shortcuts taken to ship faster

New sources added without shared meaning

How do you pay interest

Slower releases and more defects

Manual reconciliation and AI answers nobody can explain

How do you pay it down

Refactoring and modernization

Ontologies, knowledge graphs, and governed metadata

How context debt builds up across enterprise systems

Context debt piles up through ordinary decisions that made sense at the time. A new CRM arrives with its own customer ID. A regulation change, and the update sits in a PDF on a shared drive. A senior expert retires, taking years of institutional knowledge with them. The hidden cost of data silos existed long before AI, and agents make it harder to ignore.

Case study: Syngenta puts 70 years of R&D knowledge back to work

The underlying problem was not simply that information was difficult to search. Scientific terminology, synonyms, regulatory concepts, and decades of research knowledge had to be connected before that information could become useful at scale.

Datavid built Synapse, an ontology-driven semantic search platform that connects this domain knowledge and helps researchers find relevant information in minutes rather than the two to three weeks previously required.

Why AI agents turn context debt into real losses

For years, people covered the gaps. An analyst knew which spreadsheet to trust, and a planner remembered last year's festival weekend. AI agents don't carry that memory unless you give it to them. As Balvinder put it, chatbots can get by without deep context, but agents struggle without it. For a CDO, that moves context debt from a data quality topic to a business risk.

How a plain LLM, RAG, and context-aware AI answer the Boston question

The session made this concrete with one question: "How much beer should we send to Boston?"

Aspect

Plain LLM

Standard RAG

Context-aware AI

What it knows

Generic seasonal demand

World Cup event plans and last year's distribution report

Scotland qualified, where its fans will gather, and which beer they prefer

What it misses

That Scotland qualified, and what's in the warehouse

The link between fans, venues, and stock

Anything not yet modeled or refreshed in the graph

Likely output

A seasonal estimate

A summary of related documents

"Move 8,000 cases to these sites before Friday, and here is why"

In the session's illustrative example, the reasoning behind that answer is what lets a planner approve the shipment with confidence.

Standard RAG retrieves passages that look similar to the question, without joining fans to venues or venues to stock. Microsoft researchers showed that vector RAG falls short for questions that span an entire dataset.

Their GraphRAG method, a form of knowledge graph RAG, produced more complete answers in their tests. More on this in " Why RAG breaks in production and how GraphRAG fixes it."

What happens when enterprise AI agents act on incomplete context

Enterprise AI agents are starting to change prices, move inventory, and contact customers. Each action inherits whatever context the agent was given. If that context is stale, the mistake can land on a shelf or in an audit file.

This is where context debt is reflected on the income statement. Gartner predicts that through 2026, organizations will abandon 60% of AI projects unsupported by AI-ready data. It also expects over 40% of agentic AI projects to be canceled by the end of 2027, citing rising costs and unclear business value. Anthropic's engineering team reported that agents use about 4 times as many tokens as in chat interactions.

For data leaders, the practical response is AI agent governance built on a governed context. When agents reason over definitions and rules your teams have approved, you can trace their actions back to the source.  Models will increasingly commoditize. Enterprise context will not.

Five signs your organization is carrying context debt

If you lead data or AI at a large enterprise and several of these sound familiar, your AI programs may already be paying interest.

  • The Same Metric Has Several Definitions: Finance and marketing calculate revenue differently, so an agent has no single definition to reason with.
  • Complex Questions Need Several Rounds of Prompting: Teams rephrase until the AI lands on something usable, which often means the business meaning isn't in the data.
  • Your AI Finds Information but Can't Explain the Relationships: The system can retrieve documents about a customer, product, regulation, or process, but cannot reliably connect those entities or explain how they affect one another.
  • Nobody Can Show the Evidence Behind an Answer: Compliance asks where a recommendation came from, and the trail goes cold. Weak data lineage and audit trails can stall a regulated AI pilot on their own.
  • Experts Spend Their Days as Human Routers: Specialists answer the same questions repeatedly because the knowledge is scattered across their heads and scattered documents.

Paying down context debt: build a shared enterprise context layer

You're unlikely to clear context debt in one project. You pay it down like any loan, starting with the balances that cost you the most. The goal is a company brain for enterprise AI. It's a governed foundation that connects your data, knowledge, and business rules so agents can reason over them. For many data leaders, it builds on operational data hubs they already run.

The enterprise context layer, one layer at a time

The session mapped the company brain as a stack. Each layer adds something that the layer above would struggle to get on its own.

Five-layer enterprise company brain stack: Sources, Knowledge and semantic layer, Context layer, AI agents and analytics, and Decision and action.
Knowledge and context layers form the "company brain," with data flowing upward to decisions.

  • Sources: The CRM, ERP, and document systems you already run.
  • Knowledge and Semantic Layer: Ontologies and a knowledge graph define each entity and its relationships to others.
  • Context Layer: Live state and recent events, plus permissions and provenance that tell agents what they may use and where it came from.
  • AI Agents and Analytics: GraphRAG and reasoning workflows turn connected knowledge into answers and recommendations.
  • Decision and Action: Price changes and stock moves happen here, with a human in the loop for high-stakes calls.

For the foundations, see how the semantic layer accelerates AI readiness.

Case study: A global pharma company unifies biobank research with a metadata knowledge graph

Researchers at a global pharmaceutical company had to rebuild workflows for each biobank because each environment used a different metadata model.

Datavid built an ontology-driven metadata knowledge graph on the Datavid Rover semantic layer, with agentic RAG workflows and guardrails. In an eight-week proof-of-value, the team unified two biobanks and replaced ad hoc coding with reusable, governed workflows.

Context graph vs. knowledge graph: why AI agents need both

A context graph is a live graph that connects business entities with situational signals such as time, events, and intent. A knowledge graph works like a city map, showing the streets and how they connect.

Knowledge graph vs. context graph linking Scotland, Tennent's, Tartan Army, and Boston bars.
The context graph adds live signals like match day, low stock, fans landing, and 3 a.m. closing.

A context graph works more like live GPS, adding where you are and the traffic around you. In Boston's case, the map links Scotland to its fans' favorite lager. The live view adds extended bar hours and match-week weather.

Start with one decision and show the evidence

Modeling the whole enterprise at once is a common way to stall. The session's live examples asked four questions in order. What happened? Why? What should we do next? Where is the evidence?

In a direct store delivery example, the knowledge graph explained why some territories underperform and showed the graph behind its answer. That evidence trail also helps you decide what to automate and what to keep human.

Case study: Roche restarts a stalled compliance knowledge project

Roche's policy content was spread across PDFs, intranet pages, and spreadsheets. An internal effort to centralize this knowledge had stalled after two and a half years. Using Datavid Rover, Datavid delivered a working prototype of a semantic knowledge base in six weeks. It reduced advisor workload and laid the groundwork for AI-powered compliance support.

How Datavid helps you close the context gap

Datavid was founded by former MarkLogic consultants and works with CDOs and CIOs in life sciences, publishing, and financial services. Most engagements follow a simple path: an initial call to identify your costliest gaps, a focused pilot on a single decision, then scaling what works.

  • Our knowledge graph services model the entities and relationships your decisions depend on.
  • GraphRAG grounds AI agents in that connected knowledge, so answers come with traceable evidence.
  • Datavid Rover accelerates the development of a governed enterprise knowledge graph across MarkLogic, Semaphore, Databricks, or Snowflake.

For your organization, the benefit is practical. Your experts spend less time hunting for information, and your AI pilots gain a clearer route to production. Lean teams of senior engineers keep that delivery focused and cost-efficient.

Pay down context debt before your AI agents act

Boston didn't lack warning. It lacked connections between signals that were already public. If you're a CDO or CIO preparing to give AI agents real authority, your organization may be in a similar position.

Before an agent changes a price or approves a decision, check whether your systems can hand it what it needs, with evidence attached. If they can't, start paying down the debt on your most expensive decision first. Treat the context layer as a product you maintain, because trusted AI depends on context that keeps pace with your business.

Balvinder closed his session with this line: "Everyone will rent the same mind. Only you define what it can find."

Ready to find out where context debt is costing you the most?

TALK TO DATA EXPERTS

 

Frequently asked questions

Is context debt the same as poor data quality?

Not quite. Data quality measures whether individual records are accurate. Context debt measures whether the meaning between records has been captured. Clean data can still carry heavy context debt if nothing defines how customers, products, and rules relate.

How does context engineering relate to context debt?

Context engineering is the practice of giving an AI model the right information at the right moment. Paying down context debt gives that practice governed, reusable material to draw on.

Does GraphRAG reduce AI hallucinations?

It can help. GraphRAG grounds answers in explicit entities and relationships from a knowledge graph, which makes unsupported claims easier to spot and trace. Results still depend on how current and well governed the graph is.

Can you pay down context debt without replacing your data platforms?

In most cases, yes. A knowledge graph and context layer can sit across existing data lakes, warehouses, and document stores.

How long does it take to see results from a company brain?

It depends on your scope and data maturity. In the Roche and biobank projects above, Datavid delivered results in six and eight weeks.