What Is a Semantic Model? One Term, Three Meanings, One Core Idea
A semantic model is a set of definitions that gives tables business meaning, so a metric is calculated the same way wherever it's queried. Power BI, dbt and database theory use the term for versions of that idea; customer-facing models must also carry tenant access.
Ship native customer-facing dashboards and self-serve reporting fast with Embeddable
Start buildingKey takeaways
- Power BI, dbt and database theory all use "semantic model" for versions of one idea: definitions that give tables business meaning.
- Most semantic models contain entities, relationships, dimensions and measures or metrics, and the measures are where business meaning is fixed.
- The semantic model is the definitions; the semantic layer is the service that exposes those definitions to the tools that query them.
- In a multi-tenant software product, tenant scope and row-level security are part of what a metric means, not a filter added afterwards.
- Where several dashboards, APIs or AI agents query the model, we think it belongs in versioned code, enforced below the dashboard.
What Is a Semantic Model? One Idea Under Three Names
We've watched one phrase stall a planning meeting: the engineer means a dbt YAML file, the data lead means a Power BI dataset, and each assumes the other is using the word wrongly. Usually they're describing the same thing.
A semantic model is a set of definitions that maps raw tables to business concepts, so a metric is calculated the same way wherever it's queried. It typically names the entities a business cares about (such as customers or orders), the relationships between them, the dimensions used to slice data (including time), and the measures or metrics that fix how each number is calculated.
The term appears in three common contexts. In Power BI and Microsoft Fabric, every semantic model except a streaming one represents a data model (Microsoft Learn). In dbt, semantic models are defined in YAML. In database theory, a semantic data model describes entities and how they connect. In each case the purpose is the same: giving data a shared business meaning that every report and query can rely on.
The streaming exception is worth noting. Microsoft's guide to semantic models in the Power BI service says streaming semantic models don't represent a data model, so the modeling definitions described above may not apply to them in the same way.
Telling the meanings apart is mostly a matter of vocabulary. If a source talks about workspaces, reports, Fabric or "datasets" (Microsoft's older name for semantic models, per Microsoft Learn), you're in Microsoft's business intelligence (BI) world, where the model is usually built in Power BI Desktop by an analyst or BI developer.
If it mentions YAML, entities, measures and MetricFlow, it's dbt, where analytics engineers keep the model in a project repository next to their transformations. If it cites Hammer and McLeod's 1981 semantic data model (ACM TODS), or talks about conceptual schema design, it's database theory, and nothing runs at all: it's a way of describing what data means.
Our advice is simple. Work out which of the three contexts the source is in, then read "semantic model" as the same idea every time: definitions that give tables business meaning.
What Is Inside a Semantic Model: Entities, Relationships, Dimensions and Measures
Ask three teams how many active customers you had last month and you can get three different numbers. One counts logins, one counts paying accounts, and one counts anyone who fired an event. Nobody is lying; they just never agreed on a definition.
Most semantic models settle that argument with four parts: entities, the relationships between them, dimensions to slice by, and measures (or metrics) that fix how the number is calculated. Here's "active customers" broken into those parts:
# Illustrative sketch only: not any vendor's syntax
entity:
name: customer
source: accounts
key: customer_id
relationships:
- from: events.customer_id
to: customer.customer_id # many events, one customer
dimensions:
- name: event_date
type: time
- name: plan_tier
source: accounts.plan_tier
measure:
name: active_customers
calculation: count_distinct(customer.customer_id)
filter: events.type in (qualifying_events)
window: trailing_30_days
Each context names these parts differently, but they're the same parts.
In Power BI and Microsoft Fabric, the entity is a table, relationships are drawn between tables, dimensions are columns you group or filter by, and the measure is written in data analysis expressions (DAX). Streaming semantic models are the exception: Microsoft Learn doesn't treat them as data models, so most of this doesn't apply there.
In dbt's semantic models, you declare entities, dimensions and measures in YAML, with relationships inferred through shared entity keys; dbt's latest spec replaces measures with simple metrics defined directly (dbt docs). Database theory's semantic data model stops earlier: it describes entities and relationships, and leaves calculations to whatever queries it.
That last point tells you where the weight sits. The entity and relationships are plumbing (necessary, but rarely contested). The dimensions decide how people can cut the data. The measure is where someone has to commit: which events qualify, whether a trial counts, why 30 days and not 28.
In our experience, almost every "the numbers don't match" thread traces back to a measure that was defined twice, once in each place someone needed it.
So whatever your tool calls them, most semantic models are built from entities, relationships, dimensions and measures (or metrics, in dbt's current spec), and the measure is where business meaning actually gets fixed. Get that one definition right and shared, and the rest of the model has something worth carrying.
Semantic Model vs Semantic Layer
The semantic model is the set of definitions; the semantic layer is the service that exposes those definitions to every tool and consumer that queries them. One says what "active customers" means. The other is how a dashboard, an API or an AI agent actually gets that answer.
The two get blurred because vendors often ship them together. dbt is a fair example: its semantic models are YAML definitions, and its MetricFlow engine compiles metric requests against them into SQL for downstream tools. Same vendor, two separate jobs.
It also helps to say what a model isn't. It isn't the query engine that turns a metric request into SQL, and it isn't the dashboard that draws the chart. It isn't your warehouse schema either (tables describe how data is stored, not what it means), and it isn't a cache that speeds up repeat queries. All of those sit around the model. None of them decide what a metric means.
We keep the two apart because fixing one doesn't fix the other. A clean model with no layer in front of it leaves every consumer copying the logic, so the finance dashboard and the customer-facing API each re-derive churn and slowly drift apart. A polished layer over a sloppy model just serves inconsistency faster, to more places. In our view that second failure is the more dangerous one. The numbers arrive quickly and look authoritative, so nobody questions them until two customers compare screenshots.
So the model says what a metric means, and the layer is how everything else gets hold of that meaning.
That distinction sharpens as agents become consumers; our piece on the semantic layer for AI agents explains why. For the layer as architecture, see semantic layer vs headless BI. Here, we stay with the model itself.
What Makes a Good Semantic Model, and What Makes a Bad One
A good semantic model is one where a metric can't mean two things. Every other quality comes second to that.
A common failure mode looks like this. "Active customers" means a login in the last 30 days in the churn report, and any activity in the last 90 in the board deck. Both numbers are correct by their own logic, so nobody notices the gap until a customer or a chief financial officer puts them side by side. At that point the discussion is no longer about the business; it's about whose query is right. We'd call that model broken, however tidy its tables look.
Here's how we tell a good model from a bad one, in order of priority:
- Metrics are defined once and referenced, not re-typed. dbt builds its metrics on simple metrics defined inside semantic models, which act as the building blocks for other metrics (dbt docs). The principle is one definition and many consumers.
- Grain and primary entity are explicit. If you can't say whether a row is a customer, a session or an invoice, every count is a guess.
- Dimensions are named and documented. "Region" should say whether it means billing region or shipping region, so the choice doesn't depend on the last person who edited it.
- Relationships have known cardinality. A many-to-many join that nobody noticed will double-count revenue without any visible error.
- Definitions are owned and reviewed. Someone approves a change to "active" before it ships.
- Report-level overrides are discouraged, or at least visible. A local redefinition that nobody can see is how the split between 30 and 90 days starts.
If a model fails the first test, the other five won't save it.
When Customers Query the Semantic Model, Access Becomes Part of the Definition
In a multi-tenant software-as-a-service (SaaS) product, the correct value of a metric depends on who's asking. Same definition, same tables, a different right answer for every tenant.
Here's a hypothetical example of a common failure mode. Imagine a customer opens their usage dashboard, sees a spike and asks support why. The number isn't theirs: tenant scope lived in a dashboard filter, someone built a new chart without it, and in this illustrative scenario the query sums rows from every account on the platform. The measure was defined once and defined correctly, yet it would return the wrong answer and could expose another tenant's data along the way.
So the metric didn't mean two things. It meant the right thing on the wrong rows.
That's why we don't treat tenant scope as a filter added afterwards. It's part of what the metric means. For customer-facing analytics, "revenue" really means "revenue for this tenant, visible to this user". Row-level security (RLS) expresses that second half: rules that decide which rows a given identity may read. Databases have supported this for years (PostgreSQL ships row security policies natively), and a semantic model can declare the same rules beside the measures they protect. If the rule sits outside the model, every new chart, export and query path has to remember it. Some won't.
Our view is that the access rule belongs inside the model, enforced below the presentation layer (at the data, semantic, query or governed runtime layer), so no consumer can request a metric without its scope attached. In Embeddable, access policies are declared alongside the data models for exactly this reason.
The same logic applies to every other consumer. An AI agent answering a customer's question about those metrics needs identical tenant scope, which comes up repeatedly when moving AI agents from prototype to production. Once customers are asking, scope is part of the definition.
Why the Semantic Model Belongs in Code, Enforced Below the Dashboard
Our view: when more than one consumer queries your semantic model, it should live in versioned code beneath all of them, not inside any one dashboard.
The reason is mechanical. If "revenue for this tenant, visible to this user" is defined in a report layer, every consumer that can't reach that layer has to rebuild it. A second dashboard copies the measure, the public API gets its own filter, and an AI agent answering questions inside your product gets a third version written by whoever wired it up. Each copy is another chance for the metric or the tenant scope to drift. When the access rule is part of the definition, drift isn't a reporting nuisance; it's a customer seeing rows they shouldn't.
Semantic model in versioned code
Code solves the coordination problem rather than the modelling one. A change to a measure or an access rule arrives as a pull request that someone reviews. It's versioned, so you can see who changed what and roll it back. It can be tested in continuous integration (CI) against fixtures for two tenants before it merges. Then it's promoted from development to production through environments instead of being edited live. That workflow also suits agent-assisted development, because an agent proposing a model change produces a diff, not an untracked click.
GUI modelling has real advantages, and we won't pretend otherwise. An analyst can fix a measure in minutes without opening a repo, and one-off edits are simply faster. Versioning isn't exclusive to code-first tools either: Power BI can save models as Power BI Project (PBIP) files in Tabular Model Definition Language (TMDL) and keep them in Git, and other tools can query a published model.
We built Embeddable so models and access rules live in the repo as code (defining data models in code), are enforced per tenant at query time below the presentation layer, and move through environments with the rest of your release.
So here's where we'd draw the line. If dashboards, an API and an agent all query the model, put it in versioned code beneath them, because a definition that includes who may see what shouldn't be left to each consumer. If the BI tool is the model's only consumer, its model-level row-level security (RLS), ownership features and easier analyst editing can matter more.
Shipping analytics to your customers?
Define metrics and tenant access once, in code, beneath every dashboard
With Embeddable, you declare data models and row-level access policies in your repo. They're enforced per tenant at query time and move through environments with the rest of your release.
When a BI Tool's Semantic Model Is Enough
If only your own analysts query the model, keep it in your BI tool.
That's a sound choice, and it isn't a code-versus-GUI purity test. Power BI's documentation covers an embed-for-your-customers setup with row-level security (RLS) applied for each customer. Its models can also be saved as code: Power BI Project (PBIP) files in Tabular Model Definition Language (TMDL), committed to Git. (For the mechanics, see our guide to row-level security in Power BI.) Analysts keep point-and-click editing, and you still get a diff to review. For internal reporting, that balance is often the right one.
Our recommendation changes once customers query the model and several consumers sit on top of it. At that point, we'd trade some of that easy editing for review, versioning and one place to enforce access. A definition that decides who sees which rows shouldn't depend on which dashboard, API or agent asked. Giving up some point-and-click convenience is the cost, and we think it's worth paying once customers and several consumers are involved.
Whichever path you take, check that your semantic model contains:
- Primary entities and their grain.
- Relationships and join keys.
- Dimensions, including a time dimension.
- Each metric, defined once.
- Tenant scope.
- Role-aware row access.
- An owner and a review path.
- Versioning and promotion between environments.
- A list of the consumers that will query it.
Frequently asked questions
What is the difference between a semantic model and a semantic layer?
The model is the definitions: what each metric means. The layer is the service that exposes those definitions to the dashboards, APIs and agents that query them.
Is a Power BI semantic model the same as a dataset?
Yes. Microsoft renamed Power BI datasets to semantic models, so it's the same thing under a newer name (Microsoft Learn).
What is a semantic model in dbt?
It's a YAML definition over a dbt model that names its entities, dimensions and metrics, which MetricFlow uses to calculate them the same way every time.
What is a semantic data model in database theory?
It's a conceptual description of what data means (its entities and how they relate), dating to Hammer and McLeod's 1981 work (ACM TODS). Nothing runs.
Does a semantic model need row-level security for customer-facing analytics?
We think so. Once customers query it, a metric's correct value depends on which tenant and user is asking, so access rules belong inside the model.