cd /news/ai-agents/the-agent-says-what-the-engine-says-… · home › topics › ai-agents › article
[ARTICLE · art-146664] src=motley.ai ↗ pub= topic=ai-agents verified=true sentiment=· neutral

The Agent Says What, the Engine Says How: Why We Need Yet Another (OSS) Semantic Layer

MotleyAI released SLayer, an open source semantic layer under the MIT license that ships a query engine, agent tooling, and context primitives out of the box, according to the project's GitHub repository. The project targets the glue code teams typically build around proprietary semantic layers, and points to the Open Semantic Interchange (OSI, now Apache Ossie) standard as a partial fix for lock-in that still leaves custom integration work tied to each vendor's layer. SLayer joins existing tools such as Cube, the dbt semantic layer, MetricFlow, and Malloy, which the author notes include query engines but more constrained semantics.

by read16 min views1 publishedOct 7, 2026
The Agent Says What, the Engine Says How: Why We Need Yet Another (OSS) Semantic Layer
Image: Motley (auto-discovered)

Most semantic layers come with a great declarative modeling language to define metrics and measures well<sup>1</sup>, but very soon after, you start building wrappers around the semantic layer that enable agentic work and context engineering, that work well with harnesses and interfaces, and that allow flexible queries at runtime, without the need to rebuild or recalculate the measures in YAML.

The other problem is that most data warehouse or analytics solutions basically lock you in with their proprietary semantic layer. The Open Semantic Interchange (OSI, new Apache Ossie) standard, making it possible to convert definitions from Snowflake to the dbt semantic layer, will reduce this, but all the glue code you build on top is still tightly integrated into that very semantic layer, or you might get some of the capabilities mentioned above from that proprietary software. So you might be using an open data format, or an OSS business intelligence tool to integrate measures into your semantic layer, but if that semantic layer is not open, you might still be locked in.

That’s why SLayer tries another approach and ships its open source Semantic Layer with a query engine, agent tooling, and context primitives out of the box, having all the wrappers that you’d most certainly build anyway directly with it. And everything is under the MIT license.

Why Yet Another Semantic Layer? #

If you follow the space, there are many semantic layer tools, starting as early as 1991 with SAP Business Objects, and more “Modern Semantic Layer” tools such as Cube or the dbt semantic layer, etc., which are mostly YAML and outside of the BI tools.

But why might we need yet another Semantic Layer tool today? Especially when there are new terms like Context Layer and Knowledge Graph floating around (I wrote a primer about it). This is a fair question.

The biggest reason is mostly this: if you start with any of the modern layers, once you have defined your metrics in YAML and have a good workflow of versioning them in git, you can change fast and re-deploy, roll back, and even agents can easily contribute, with changes mergeable and transparently seen in the git log with PRs if you want. But what you always end up with is writing glue code around your Semantic Layer.

For example, in 2018, when we built our own Semantic Layer at a previous job, a Metadata store held all our metrics in a simple SQLite database for our SmartAnalytics application, which then had all the complex queries pre-defined for users to use, querying in milliseconds against Druid as our OLAP engine.

But we ended up building an API to fetch/create new metrics. We had to build a lot of glue code so that the integration with Druid worked well, that the application could fetch the metrics in a way that worked with its C# application logic, that we could use it in our Dagster data pipelines, that data scientists could use the same metrics in Jupyter Notebooks, or in the web app where we showed millions of points that had to be plotted on a map when zooming around the map. The declaration was already a major win, as we could centralize the business logic in one place for everyone to use, but we had to add the APIs and glue code ourselves.

Today, with agents, there’s one more layer that you need to build: the implementation for working with agents and for collaboration between humans and agents. So there you implement another wrapper layer. Usually, the paid solutions might solve some of these problems, but shouldn’t that be included out of the box with open source? SLayer, another semantic layer, and the company Motley think yes.

Also missing is an included query engine, though. MetricFlow and Malloy have an engine to generate and query, but the semantics are more constrained. SLayer, for example, comes with an engine out of the box that supports powerful expressions like sum(revenue), avg(revenue), sum(revenue) / count(*), time_shift(sum(revenue), -1, 'year'), or multi-stage queries such as:

Orders grouped by customer activity bucket. First calculate order count per customer, bucket that, then use the bucketed result as a dimension to group orders by.

This is a multi-stage query, a combo of two queries, and something that can be common in an enterprise, so it helps if we can define this with a simple expression language, as it’s very difficult to express in SQL.

Why not just dump the text into the agent?

Another question you might have: why not just dump the text directly into the chatbox of an agent? The reason is that it usually works on a simple demo project or a small one, but does not scale to the enterprise. One showcase of larger context needs is LiveSQLBench-Large. Another angle is the reiteration of agents, re-deriving cardinality, join paths, and metric definitions on every session instead of reading them once, as well as two people asking the same questions and getting different answers; an integrated Semantic Layer avoids drift in the metrics returned.

Three Out-of-the-Box Capability Layers That Come with an Open Agentic-First Semantic Layer #

If we look at what the capabilities and requirements for a new agentic-first and fully open source semantic layer are, they are managing context and enabling queries easily based on it, flexibly and with an expressiveness that makes it unambiguously clear to humans, but also to the agents.

This needs tooling that has a more powerful expression language and a query engine that supports it, integrated deeply into the definitions but also into the query engine. Yann Ranchere, co-founder and CEO of Motley, says it this way:

A strong agent-first semantic layer, with an expressive query engine, agent tooling, and context primitives, is a hard problem and a key component of future agentic and analytics solutions and that it is better served by an open-source project so that developers can spend more resources and effort on context, agent harness, visualisations and interfaces.

To tackle all of these, we can compress them into three pillars: context primitives, an expressive and flexible query engine, and being agent agnostic, with all of these capabilities integrated into OSS tooling that isn’t gated behind a paywall<sup>2</sup>.

1. Context Primitives Out of the Box: No More Wrappers and Glue Code

Pre-agentic Semantic Layers were built for dashboards that replay the same queries each time, so joins, aggregations, and time logic get defined once in YAML at design time. This is great for large and complex business logic, like revenue.

But the definition alone was never enough. Think back to the Metadata store example from above: the metrics lived in one place, but everything around them was glue code we wrote ourselves: the API to fetch and create metrics, the Druid integration, the access from C#, Dagster, and Jupyter. Every consumer needed its own way to find out what metrics exist, look up how they are defined, and query them.

With agents, the same three needs come back, only stronger. An agent doesn’t know your business. It has to find the right measure among hundreds, inspect how it is defined and joined, and ideally learn from what worked the last time someone asked a similar question. This is what I call context primitives: the small, composable units a semantic layer exposes so an agent can retrieve exactly the business knowledge it needs for one question, without the whole model into its context. In SLayer, these are:

  • Models : measures, dimensions, joins with their cardinality, and grains, defined in YAML and versioned in git, or imported from dbt, Cube, or OSI.
  • Search : onesearch tool over entities (datasources, models, columns, named measures) and memories, so an agent findsorders.amount from a natural-language question instead of reading every model file.
  • Inspect : the rendered definition of exactly one entity, with descriptions and sample values, for when the agent already knows what it wants.
  • Memories : notes an agent writes about a part of the schema, optionally with an example query, tagged to the entities they belong to (“Card TPV excludes refunds after day 45”). The next agent finds them via search before it queries.

Together they form the search → inspect → query flow we’ll see later, and they ship as MCP tools, a REST API, a Python client, and a CLI. These are exactly the wrappers we would otherwise build again, around every semantic layer and for every harness.

In the end, we want to build fewer wrappers ourselves and less custom code, and have all semantic capabilities provided end-to-end as open source so developers can spend more resources and effort on context, agent harness, visualisations and interfaces. The same holds for the second capability: a flexible query language and engine belong to the semantic layer.

A semantic layer should be a governed query engine, not a SQL wrapper

order_total:sum
sum(order_total)

2. Expressiveness + Flexible Query Engine That Works with AI Agents, Agnostically

Better expressiveness and more integration also mean more efficiency with agents, instead of going through many wrappers and non-integrated, non-battle-tested code with much longer round trips, meaning that agents otherwise use lots of tokens, with more ambiguity and less speed.

This is essentially the origin story of SLayer: building a reporting automation on Cube, whose wrapper grew into a complete rewrite. But in order for agents to express their query intent in a way that’s as concise as possible, and just formal enough to be unambiguous, we need an expressive query language integrated with its query engine. The query engine of an integrated Semantic Layer needs to deterministically figure out the necessary joins, subqueries, and safeguards, both for correctness and to have the right context understanding, to avoid needing to query the full world or all of its huge context, saving context<sup>3</sup>.

This is also where the static, design-time model hits its limits. If an agent asks a question that wasn’t planned, the pre-defined pattern breaks. Or the workaround, a model that pre-declares every measure × aggregation × time-shift combination, explodes. But there are more complex things for which we need additional logic written, for example, doing an aggregation across different grains, connecting the missing dimension, and then aggregating based on the result set, all in a single concise expression. SLayer calls this re-aggregation: a multi-stage aggregation inline in one query, no second stage needed, and guaranteed fan-out safe (the multi-stage query introduced above does the same across two stages): avg(sum(orders.total, partition_by=[city, region])) grouped by region, the unweighted average of city totals per region.

Another feature some OLAP engines like Druid or ClickHouse bring is moving decisions to query time, so you can define metrics while querying your data (not at processing or development time). But you need the engine for it and now write custom integrations and APIs for each query engine again, whether you use Snowflake, Cube, or any other engine.

Handling different grains, cardinality when joining, and many-to-many relationships are hard problems, and every bit of integration in that direction helps the domain experts querying and defining complex metrics in large organizations, once. A Semantic Layer is much more than a database schema.

How a flexible query engine with context primitives works, here in red is what SLayer has built. Source: Principles SQL

An agent needs to express its query intent concisely and unambiguously. The joins, subqueries, and safeguards are the engine’s job, and they have to be deterministic. SLayer’s SQL principles show that by using an AST (Abstract Syntax Tree) end to end and SQLGlot transpilation. Check out Try SLayer for an example.

For the agent, this means it states the what, such as which measures, at which grain, etc. The engine deterministically owns the how, like any query planner on a database, but on top of our defined metrics, with the intent to produce the same SQL every time. This leads to more trust in the agent’s answer.

A Quick Evolution of the Semantic Layer

A quick recap of the history and evolution of semantic layers. Originally, and still, the key purpose of a semantic layer is to have a virtual layer between the raw data and the dashboard or end-user analytics, without needing to manage direct access and permissions to your raw data.

This short recap shows the evolution and where that comes from (originally shared here, but extended with the latest):

1991: SAP BusinessObjects Universe and BI semantic layer 1997: SSAS and MDX, with their logical modeling layer with MDX, define business metrics and dimensions in a structured way 2013: Kimball discussed the concept of a semantic layer in #158 Making Sense of the Semantic Layer 2016: Maturing BI tools with an integrated semantic layer, such as Tableau, TARGIT, PowerBI, Apache Superset, etc., have their own metrics layer definition 2019: Looker and LookML were popularized as the first semantic layer 2022: Declarative Semantics and metrics with tools such as Cube, MetriQL, MetricFlow, Minerva, and dbt arose with the explosion of data tools around 2022 2023/24: The evolution of data querying and natural language queries with the integration of LLMs and catalogs. Text-to-SQL, embeddings & Retrieval-Augmented Generation (RAG) 2025: Resurrection of Knowledge Graphs & Ontologies 2026: Agent-written Context Layer & Karpathy self-curating LLM Evolution from Semantic Layer. See also related Enterprise & Personal Knowledge Management Evolution, and you can find another opinion in The Shape and Feel of the Post-AI Data Stack.

3. Use Your Harness of Choice: Agentic-Tooling Native, with Unambiguous Semantics

Other challenges are quickly changing queries while we explore the data with agents, as well as that we might need to use Copilot on Microsoft because it’s integrated, or we like the way Claude works. So the semantic layer must integrate into these tools, not the other way around, where each semantic layer builds its own AI harness.

Similarly important is to define metrics unambiguously, so that the time dimension, for example, one of the hardest dimensions to define<sup>4</sup>, is clear and flexible: time_shift(sum(order_total), -1, 'year'). Imagine pre-defining these in YAML at pre-build time, for all dimension/granularity combinations. Instead, we make it flexible at query time, at runtime, for agents to simply adapt to their needs.

The goal: Combine different grains - be very flexible, not depending on technical limitations such as grain or cardinality, but based on business demands. The query engine and its semantics should handle the rest, and agents can help us with it. Letting the agent ask what it needs, instead of what the semantic layer can express in its limited way. We know the limitations from DSLs (Domain Specific Languages): they are good for non-technical people to not break anything or to just stay within the prepared queries, but not so much if you need new queries or have fast-changing requirements. So we don’t need a DSL; we need flexibility and more expressiveness.

Being able to use any harness, being agnostic, is also what keeps the tool out of vendor lock-in, as it works within your existing environment.

Egor Kraev, co-founder of Motley, says:

An agent should be able to express its query intent in a way that’s as concise as possible, and just formal enough to be unambiguous, but let the deterministic machinery figure out the necessary joins, subqueries, and safeguards, for both safety/correctness and context-saving reasons.

What is a “(AI-) harness”?

The term popped up on social media, and it means scaffolding around model (tool loop, prompts, context management), working autonomously, not needing to constantly ask us, but that can run for 20 minutes and solve a hard problem. Harnesses are Pi Coding Agent, OpenCode, Aider, Goose, Cursor, Amp, Claude Code, and Codex, which have now built in this functionality.

What an Embeddable, Agentic-First Semantic Layer Looks Like #

So with all of this in mind, what would yet another Semantic Layer look like with the three foundations of avoiding building wrappers, including an expressive query language and query engine, and being able to use our harness of choice?

Luckily, we can look at SLayer, an open source implementation of these principles. SLayer is focused on the common agentic search → inspect → query flow with a search tool for efficient discovery and a memory store for linking the relevant business context. Try it on GitHub, and start with the quick start:

git clone git@github.com:MotleyAI/slayer.git
uv tool install 'motley-slayer[all]'

And use the demo data of Jaffle Shop, preloaded with zero config needed, with your agent of choice, here Claude, but you can use any MCP integration, included in the OSS version. Try it and get a feel for it:

claude mcp add slayer_demo -- slayer mcp --demo

Or even better, with your own data - a Postgres DB in this case:

slayer datasources create 'postgresql://user:${DB_PASSWORD}@hostname/db_name'

All of the integration is better served open-source, and it is with SLayer.

What’s Next? #

The usual flow of building a semantic layer is that you start with a cloud warehouse, you start up something quick and fast, the dashboard or insight gets more useful, more people use it, more want new metrics added, so you patch and update the ETL pipeline. The pipeline gets slow, so people ask to improve it, so you build caching mechanisms and materialized tables to speed it up, but all of a sudden, you end up with a mess. All your metrics and business transformations are scattered around in the ETL pipelines, in the materialized views, in the newly patched scripts you build, and in ad-hoc dashboards.

So instead, we could use an agentic semantic layer that is prepared for fast-changing metrics and business requirements with its extendable query engine, integrated with any popular AI harness, as well as absorbing and defining the right amount of business context through context primitives.

In this article, we learned why we might need yet another semantic layer that is open source and built for today’s needs. If you decide on a semantic layer, look at agent-friendliness, query-time expressiveness, and entity and context discoverability. If it’s not included, you usually end up building it yourself. You want a semantic layer so agents don’t end up writing raw SQL.

Again, agents should be able to express their query intent in a way that’s as concise as possible, and just formal enough to be unambiguous, but let the deterministic machinery figure out the necessary joins, subqueries, and safeguards, for both safety/correctness and context-saving reasons. This increases query expressiveness and simplicity, so dimensions can be derived from aggregations, as well as multi-stage aggregations, or guaranteed fan-out.

Curious to learn more? See related blog articles and architecture decisions:

Footnotes #

e.g., Semantic Layer Measure Definition Examples↩ 2. Some of the semantic layers do have an expressive query language, or APIs or enhanced features to integrate into the existing data stack, but they are behind a paid service. ↩ 3. This is also a good place to mention that a semantic layer is not the same as newer upcoming terms such as Context Layer, which is more about managing and ingesting context at large, keeping it up to date and managed. A semantic layer is more about the definition of metrics and business-relevant KPIs, not so much scraping or updating the actual “context” or information.↩ 4. E.g., time dimension: When we define one year, is it from today, or the end of last month, or of the week? Time or natural language is hard for this. ↩

── more in #ai-agents 4 stories · sorted by recency
── more on @slayer 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/the-agent-says-what-…] indexed:0 read:16min 2026-10-07 · —