Articles on all things data | Datavid blog

Semantic Layer Tools: Build, Buy, or Model It First

Written by Datavid | Oct 6, 2026

Quick answer

Build when your stack is stable, your scope is metrics over one warehouse and engineers can maintain it for years. Buy a semantic layer platform when many BI tools and AI agents need shared definitions this quarter. Model it first when your business context is spread across documents as well as tables, since the model is the part you keep, regardless of which tools run it.

As a Chief Data Officer, you are choosing semantic layer tools just as AI agents, BI tools, and regulators start depending on the same definitions. The goal is to choose an approach that will still hold when your stack changes, so your foundation supports AI without a rebuild in eighteen months.

You can build it, buy it, or model it first, then let the model decide on the tools. If you need the basics first, this semantic layer explainer covers them. For regulated enterprises, the right route tends to depend on where your business meaning lives.

At a glance

  • Semantic layer tools fall into six categories, and the category largely shapes where your business logic is defined.
  • Building fits a stable, single-warehouse stack, and its costs tend to arrive later as drift and rework.
  • Buying can quickly get many tools and agents onto shared metrics, while the platform's modeling language shapes your model.
  • Modeling it first defines your entities and rules independently of any tool, so the build-or-buy decision stays open to revision.
  • For regulated enterprises, any of the three routes can work, and modeling first is one option when the meaning is in both documents and tables.

Semantic layer tools by category

Semantic layer tools generally fall into five common categories, with ontology and knowledge graph approaches representing a related but conceptually different option.

Each approach shapes where business definitions are managed and how easily they can move across systems. For a CDO, the underlying approach tends to matter more than the product name. It influences the lock-in you inherit and how much of the work survives a platform change.

Six categories of semantic layer software

The differences become clearer when you compare what each approach models and where that work is tied to a specific platform or tool. The table below shows those tradeoffs side by side.

Category

Example tools

What it models

Lock-in pattern

Code-first and headless

dbt Semantic Layer (MetricFlow), Cube

Metrics defined in code alongside data pipelines

Lower, since definitions sit in your repository

Universal semantic layers

AtScale, Dremio

Metrics across integrated data sources for many BI tools

Tied to the vendor's modeling language

Warehouse and lakehouse native

Snowflake Semantic Views, Databricks Metric Views

Metrics inside one platform

Tied to that platform's data architecture

BI-native models

Looker (LookML), Power BI, Tableau

Metrics for one BI tool

Tends to stay inside that tool

Catalog and metadata layers

Atlan, OvalEdge

Glossary terms, metadata, and lineage for data governance

Tends to stay in the catalog

Ontology and knowledge graph

RDF graph stores, MarkLogic

Entities and rules across tables and documents

Lower when built on open standards

Metrics layers and meaning layers

The first five categories model metrics over tables, while the sixth models entities and rules across tables and documents, where much of a regulated enterprise's meaning sits. A Snowflake semantic layer answers "revenue by region" well, while "which trials used this compound under which protocol" is better served by a graph.

You don't have to pick one category for each job. Knowledge graph solutions can sit alongside your existing metrics tools, so BI investments keep their value while AI agents gain document-level meaning.

When building a semantic layer in-house, it makes sense

Building is a sound choice under the right conditions, and many teams run a homegrown layer well. Open-source semantic layer options, such as dbt Core with MetricFlow and Cube, reduce the license bill. The ownership cost stays with your team, so the question for a CDO is whether these conditions will still hold in three years.

Conditions where a build tends to work

  • Single warehouse: One platform holds the data your metrics use.
  • Few consumers: One or two BI tools, with no AI agents planned soon.
  • Stable definitions: Metric logic rarely changes and has clear owners.
  • Metrics-only scope: The meaning you govern sits in tables.
  • Committed platform team: Engineers can maintain the layer for years.

Costs that tend to surface later

Build costs rarely appear in the first budget. They tend to arrive as drift when more tools connect and as rework for each new AI agent. Key-person risk can grow when the builders move on, and audit effort tends to climb when a regulator asks where a number came from.

Standards Australia: a stalled build, rebuilt in six weeks

At Standards Australia, a third-party vendor spent 18 months on a custom build that hadn't reached production. The issues lie in the custom code rather than the underlying MarkLogic platform, which is the main risk of the build route. The root causes were diagnosed in one week, and the core application was rebuilt on the same platform in six weeks. Responses dropped from about 10 seconds to 300 milliseconds per request.

For a team weighing the build route, a stalled build can often be recovered on the platform you already own rather than written off, which protects both your budget and your board's timeline.

When buying a semantic layer platform, it makes sense

Buying fits when speed and reach matter more than control. A semantic layer platform, including a universal semantic layer that spans multiple warehouses, can quickly get many BI tools and AI agents onto shared metrics. The trade-off is that the platform's modeling paradigm tends to shape your meaning, which can be hard to change later.

Signals that point a CDO toward buying

  • Several warehouses and BI tools: Definitions need to match across platforms.
  • AI agents on the roadmap:  Agents are set to query the metrics your dashboards use.
  • Results this quarter: The board expects progress before a build could land.
  • Reporting obligations: Regulators expect one definition wherever a number appears.

Portability and the Open Semantic Interchange

The Open Semantic Interchange, backed by Snowflake and Databricks among others, has entered the Apache Incubator as Apache Ossie (incubating). It aims to give definitions a vendor-neutral format that moves between the semantic layer software. For now, portability still depends largely on how you model.

ABN AMRO: extending the platform you already own

ABN AMRO faced the reporting obligations signal directly, with daily MiFID II reporting to support. Its existing TradeStore hub on MarkLogic was extended with real-time compliance logic, control checks, and end-to-end traceability. It now handles that reporting with straight-through processing, a pattern that carries over to data compliance work elsewhere.

If you already run a platform, extending it before buying can deliver much of what a new purchase promises, without a second modeling language to maintain. On a MarkLogic estate, MarkLogic consulting and disciplined enterprise data management make that a realistic path while past investment keeps working for you.

When modeling, it first makes sense

Modeling first means defining your entities, relationships, and business rules in a semantic model before choosing where it runs. A semantic model is a shared description of what things are and how they relate to one another. It's one of three routes, and its trade-off is more domain modeling up front. Because it sits apart from any tool, the build or buy decision becomes one that a CDO can revisit without starting over.

The model outlives the tools

Definitions defined in open standards such as W3C's OWL 2 can move with you when a BI tool, warehouse, or AI platform changes. That is how a semantic data layer keeps meaning attached to your data through a tool swap.

The model covers documents as well as tables

Metrics layers describe numbers in tables, while much of the meaning an AI agent needs sits in protocols, policies, and standards. The model links those documents to the entities your metrics use, so a dashboard figure and a policy clause can share one definition.

The model points to the tools you need

With entities and rules defined, tool selection becomes a fit question: a metrics layer for dashboards, a knowledge graph for documents and agents, or both. The ontology vs semantic layer question tends to settle here, too: the ontology holds the meaning, and the semantic layer serves it.

Roche Helios: what modeling first returns

At Roche Helios, study-specific forms were mapped into a common semantic model, with ontology-based checks and business rule validation running on top. Trial data processing became 5x faster, and manual errors fell by 80%. Operational costs dropped by 40%, with on-demand audit readiness for FDA and EMA.

The same approach applies to clinical trial data programs, and the Beyond the Model webinar with Progress shows it in practice through three ideas:

  • Semantic foundations give AI trusted business meaning to reason with.
  • Knowledge graphs make AI outputs explainable and traceable to their sources.
  • You can start with a focused use case and expand the model over time.

A semantic data platform such as Rover can carry reusable components that shorten the path from model to working data product. Sustained ontology management keeps the model current, so the asset keeps paying back.

What each route costs: a TCO view for CDOs

Total cost of ownership is where the three routes most clearly separate. These cost drivers can be priced against your own numbers rather than industry estimates, which makes them easier to defend in a budget review. Most differences appear after year one, as more tools connect and the first platform change arrives.

Cost drivers across build, buy, and model it first

The cost difference is not just about licensing or initial setup. Ongoing maintenance, integration work, switching costs, and audit requirements can significantly alter the economics over time.

Cost driver

Build

Buy

Model it first

Upfront effort

Moderate engineering time

Lower, mostly configuration

Higher, mostly domain modeling

Licensing

Low with open source

Recurring platform fees

Depends on the tools used

Maintenance and drift

Grows as tools connect

Shared with the vendor

Lower, with one model

New BI tool or AI agent

Custom integration

Covered if supported

Reuses existing entities

Switching tools later

Rework of logic

Remodeling in a new language

Model carries over

Audit and compliance

Manual tracing through code

Depends on platform lineage

Provenance held in the model

Build-and-buy can pay again at each stack change, while modeling first front-loads more of the cost and then reuses it. Which curve is easier to defend depends on how often your stack is likely to change, and model-held provenance supports AI decision traceability when auditors ask.

How to choose the best semantic layer tools for your stack

The best semantic layer tools for your organization depend on fit, not rank. Six questions tend to settle the route, and most of them have little to do with features. They ask where your meaning lives, who consumes it, and what an auditor is likely to expect from you.

Six questions for your steering meeting

Question

Build

Buy

Model it first

Where does your business meaning live?

Governed tables

Tables across platforms

Tables and documents

How many tools and agents will share definitions in two years?

One or two

Many

Many, including agents

How often do definitions change?

Rarely

Occasionally

Often, with several owners

What will an auditor ask you to trace?

Numbers

Numbers and lineage

Numbers back to sources and rules

Who maintains the layer in three years?

A committed platform team

The vendor with your team

A shared model with partner support

When does the board expect the first result?

Within the year

This quarter

This quarter, with accelerators

Tables and one tool tend to point to building. Many consumers and speed point to buy. Documents, audit exposure, or a likely stack change point are the first to be modeled. Working through them before a vendor demo gives a CDO a narrower shortlist, and it can take some purchases off the table.

What AI agents need from semantic layer tools

Agents tend to need governed definitions that are reachable via an API or an MCP server, with provenance attached and access rules enforced at query time. Metrics layers handle numeric questions, while questions spanning documents usually call for the graph. A semantic layer for AI tends to combine both, and an AI semantic layer built this way underpins AI readiness.

Choose semantic layer tools that fit your stack

Whichever route you choose, a durable semantic layer investment is one that survives your next platform decision. Where modeling first fits, reusable accelerators on the stack you already own can bring a working data product to life in 6 to 8 weeks.

Find out whether you should build, buy, or model it before you commit to a platform.