SeenPixel

Engineering notes

Why 'chat with your data' projects fail, and what a semantic layer changes

SeenPixel engineering · October 4, 2026 · 6 min read

The demo always works. Someone connects a language model to the warehouse, types “what was revenue last quarter?”, and a number comes back in four seconds. The room is impressed. Six months later the project is quietly shelved, and nobody can say exactly when it died.

This is now common enough to have statistics. Gartner expects more than 40% of agentic AI projects to be canceled by the end of 2027, citing cost, unclear value and inadequate risk controls (Gartner, 2026, as reported by Forbes). Stacklok's State of MCP in Software 2026 found that only 11–14% of MCP pilots reach production. And in dbt Labs' State of Analytics Engineering 2026, 71% of respondents said they worry about hallucinated data reaching stakeholders.

Those three numbers describe one failure from three angles. This note is about what that failure is, and why a semantic layer changes the odds.

What actually goes wrong

A text-to-SQL agent pointed at raw tables has to do three hard things at once: understand the question, guess which of your four hundred tables answer it, and reconstruct your business logic from column names. The first is what language models are good at. The other two are where projects die.

1. The metric has no single definition

Ask five people what “revenue” means and you get gross bookings, net of refunds, recognized revenue, revenue excluding one legacy product line, and whatever the board deck uses. Each lives in a different dashboard or dbt model. The agent picks one, or invents a sixth. Its answer is fluent, plausible and different from the number finance reported. The first time an executive notices, trust is gone.

2. Joins are guessed

Warehouses are full of traps that every analyst on the team knows and no schema documents: the orders table that fans out on line items, the customer table with soft-deleted rows, the events table in UTC next to a calendar table in another time zone. A model writing SQL from column names walks into each of them, and the query still runs.

3. Access control lives in the prompt

The quick pilot runs as a service account that can read everything, with an instruction not to reveal salaries. That is not access control. Once the agent is opened beyond the data team, someone asks whether a regional manager can see another region's numbers through it, and the honest answer stops the rollout.

4. Nobody can prove it is right

A dashboard is checked once, when it is built. An agent produces a new query for every question, so it has to be checked continuously. Without a set of questions with known answers, run on every change, “it seems to work” is the only evidence available.

5. Cost is unbounded

Each question is one or more model calls plus one or more warehouse queries, some of them full scans. dbt Labs' 2026 report found compute costs up about 50% while only 36% of teams have rising budgets. An agent with no cost ceiling adds to a bill already under pressure.

Why the protocol does not fix it

The Model Context Protocol made it much easier to give an agent tools, and adoption has been fast: Stacklok reports 41% of software organizations running MCP in limited or broad production. But MCP is plumbing. It standardizes how the agent calls a tool; it says nothing about what the tool should let the agent do. An MCP server whose one tool is run_sql has reproduced the original problem with better packaging.

Gartner put a number on this in May 2026: by 2028, 60% of agentic-analytics projects that rely solely on MCP will fail for lack of a consistent semantic layer. The same research says semantics can raise agent accuracy by up to 80% and cut cost by up to 60% (Gartner, May 2026). The protocol is necessary. It is not the governance.

What a semantic layer changes

A semantic layer is the place where metrics, dimensions and the joins between them are defined once, in code, under version control. Cube, the dbt Semantic Layer and LookML are the common implementations. It was built so that two dashboards could not disagree, and an agent needs it for the same reason.

  • The agent chooses; it does not compose. Instead of writing SQL, the agent asks for net_revenue by region for last quarter. The layer compiles that to SQL it already knows is correct. The space of wrong answers shrinks from “any query that parses” to “the wrong metric from a short list”.
  • Definitions are reviewable. Finance can read the definition of net revenue in a pull request and approve it once.
  • Joins are declared, not guessed. The fan-outs and soft deletes are handled in the model, by the people who know about them.
  • Access control has somewhere to live. Row- and column-level rules attach to the semantic model and the warehouse roles beneath it. The agent runs as the user, so it cannot return what the user could not query directly.
  • Evaluation becomes tractable. With a finite list of metrics, you can write golden questions with known answers and score the agent on every change.

What it does not change

A semantic layer is not a shortcut around data quality. If the underlying tables are late, duplicated or wrong, the layer will serve wrong numbers consistently. Deloitte's State of AI in the Enterprise 2026 reports that 72% of organizations lack unified, accessible data, and Gartner found that organizations with successful AI initiatives invest up to four times more, as a share of revenue, in data and analytics foundations (Gartner, April 2026). The layer is part of that foundation, not a replacement for it.

It also does not remove the need for scope. An agent that answers questions about two well-modelled domains is useful in week six. One that is supposed to answer anything about everything is not.

A sequence that works

  1. Start from the questions. Collect the thirty questions people actually ask, with the answers finance already trusts. These become the evaluation set.
  2. Model one or two domains. Define the metrics and joins those questions need, in the semantic layer, with tests.
  3. Expose the layer, not the warehouse. An MCP server with a few typed tools: list metrics, describe a metric, run a metric query. No free-form SQL.
  4. Put policy below the agent. SSO, row- and column-level rules, an audit record for every call.
  5. Run the evaluations in CI. Every change to a model, a prompt or a definition is scored before it ships.
  6. Set a cost ceiling and watch it. Cost per query is a metric like any other.

None of this is exotic. It is the ordinary discipline of data engineering, applied before the agent rather than after the incident.

We have drawn this sequence as a reference design, and it is how we run an ask-your-data pilot.

Sources

  • Gartner, 2026: more than 40% of agentic AI projects canceled by end of 2027 (reported by Forbes, July 2026).
  • Gartner, May 2026: lack of semantics causes inaccurate AI agents and wasted spending.
  • Gartner, April 2026: organizations with successful AI initiatives invest up to four times more in data and analytics foundations.
  • Stacklok, State of MCP in Software 2026.
  • dbt Labs, State of Analytics Engineering 2026.
  • Deloitte, State of AI in the Enterprise 2026.

Planning an ask-your-data project?

Bring the questions your team asks most. We will tell you what the semantic layer underneath would need to look like.