{"slug": "the-semantic-compression-problem-engineering-ai-ready-views-for-complex-sql", "title": "The Semantic Compression Problem: Engineering AI-Ready Views for Complex SQL", "summary": "A developer argues that large analytical SQL queries fail not because SQL is hard but because meaning is distributed across CTEs, joins, and derived metrics, and proposes semantic views as a compression layer that exposes business concepts like Gross Revenue and Net Revenue above physical computation. The writeup cites Snowflake's schema-level semantic views and text-to-SQL research including the Spider benchmark and RAT-SQL to show that language models must reconstruct schema meaning when the semantic layer is weak. The recommended approach is to split a large transformation into physical implementation and business semantics before modeling.", "body_md": "Large SQL queries rarely become difficult because SQL itself is difficult.\n\nThey become difficult because meaning gets distributed across the query.\n\nA 20-line query can usually be understood by reading it top to bottom.\n\nA 500-line analytical query is different.\n\nIts meaning may be distributed across:\n\nnested CTEs\n\nmultiple joins\n\naggregation levels\n\nderived metrics\n\nbusiness filters\n\ndate logic\n\nslowly changing dimensions\n\nwindow functions\n\naliases\n\nimplicit assumptions\n\ntechnical column names\n\nduplicated business rules\n\nAt that point, adding another abstraction isn't necessarily the answer.\n\nThe real question becomes:\n\nHow do we compress the physical complexity of a data system into a semantic representation that humans, BI systems, and AI agents can reliably reason about?\n\nThat is where semantic modeling becomes interesting.\n\nConsider a simple calculation:\n\nSUM(unit_price * quantity)\n\nTechnically, this is just an aggregation.\n\nBut suppose the business calls it:\n\nGross Revenue\n\nNow consider:\n\nSUM(unit_price * quantity * (1 - discount_rate))\n\nThe business might call that:\n\nNet Revenue\n\nThe SQL tells us how the value is calculated.\n\nThe semantic layer tells us what the value means.\n\nThat distinction becomes increasingly important as analytical systems become consumed by AI.\n\nSnowflake describes semantic views as schema-level objects that model business entities, relationships, dimensions, facts, and metrics on top of physical data.\n\nThe abstraction therefore becomes:\n\nPhysical Data\n\n      ↓\n\nSQL Computation\n\n      ↓\n\nBusiness Semantics\n\n      ↓\n\nHuman / BI / AI Consumption\n\nThe semantic layer isn't supposed to hide SQL.\n\nIt is supposed to expose the right meaning above SQL.\n\nImagine an analytical query:\n\nWITH orders AS (\n\n    ...\n\n),\n\ncustomers AS (\n\n    ...\n\n),\n\nproducts AS (\n\n    ...\n\n),\n\ndaily_orders AS (\n\n    ...\n\n),\n\ncustomer_metrics AS (\n\n    ...\n\n),\n\nregional_metrics AS (\n\n    ...\n\n),\n\nranked_products AS (\n\n    ...\n\n)\n\nSELECT ...\n\nThe query may be completely correct.\n\nBut correctness isn't the same thing as usability.\n\nA new analyst now has to understand:\n\nWhich table represents the customer?\n\nWhat is the grain of orders?\n\nWhat does revenue mean?\n\nWhich date should be used?\n\nWhich joins are one-to-many?\n\nWhich filters are mandatory?\n\nWhich aggregation is authoritative?\n\nWhich calculation is business-defined?\n\nWhich CTE exists only as an implementation detail?\n\nThis is semantic complexity.\n\nAnd semantic complexity becomes particularly important when an AI system needs to generate SQL.\n\nNatural-language-to-SQL research has repeatedly shown that generating correct SQL requires more than understanding the user's sentence.\n\nThe model also has to understand the database schema and relationships.\n\nThe Spider benchmark demonstrated this problem explicitly by evaluating text-to-SQL across 200 databases, 138 domains, and thousands of complex SQL queries.\n\nRAT-SQL later showed how important schema encoding and schema linking are when translating natural language into SQL, particularly when the system encounters previously unseen schemas.\n\nThis leads to an important architectural observation:\n\nNatural Language\n\n       ↓\n\nIntent\n\n       ↓\n\nSemantic Concepts\n\n       ↓\n\nSchema Mapping\n\n       ↓\n\nRelationships\n\n       ↓\n\nSQL\n\nIf the semantic layer is weak, the model has to reconstruct too much meaning from raw database structure.\n\nThat is a difficult problem.\n\nThis is one of the easiest mistakes to make.\n\nSuppose you have a 700-line SQL transformation.\n\nIt is tempting to think:\n\n\"I'll put the entire query inside a semantic view.\"\n\nBut that doesn't necessarily create a good semantic model.\n\nInstead, first separate the query into two categories.\n\nPhysical implementation\n\nCTEs\n\nTemporary transformations\n\nTechnical joins\n\nDeduplication\n\nIntermediate calculations\n\nStaging logic\n\nOptimization logic\n\nBusiness semantics\n\nCustomer\n\nOrder\n\nProduct\n\nRevenue\n\nProfit\n\nConversion Rate\n\nActive Customer\n\nOrder Date\n\nRegion\n\nProduct Category\n\nThe semantic layer should primarily expose the second category.\n\nA useful mental model is:\n\n```\n            700-line SQL\n                 │\n      ┌──────────┴──────────┐\n      │                     │\n```\n\nImplementation            Meaning\n\n          │                     │\n\n          ▼                     ▼\n\n     SQL complexity       Business concepts\n\n                                │\n\n                                ▼\n\n                         Semantic View\n\nThis is semantic compression.\n\nWe're not necessarily reducing the amount of computation.\n\nWe're reducing the amount of meaning a consumer has to reconstruct.\n\nBefore creating dimensions or metrics, ask:\n\nWhat does one row represent?\n\nFor example:\n\norders\n\n→ one row per order\n\norder_items\n\n→ one row per order item\n\ncustomers\n\n→ one row per customer\n\ndaily_sales\n\n→ one row per customer/product/day\n\nThis sounds basic.\n\nIt isn't.\n\nGrain determines whether a metric is valid.\n\nConsider:\n\nOrders\n\n1 customer\n\n   │\n\n   ├── Order A\n\n   ├── Order B\n\n   └── Order C\n\nNow imagine joining:\n\nOrders\n\n    ×\n\nOrder Items\n\n    ×\n\nProduct Events\n\nIf one order contains 4 items and each item has 3 events, careless aggregation can create:\n\n1 × 4 × 3 = 12 rows\n\nA metric such as:\n\nSUM(order_amount)\n\ncan now be multiplied unintentionally.\n\nThe SQL may execute successfully.\n\nThe result can still be semantically wrong.\n\nThis is why grain is more fundamental than syntax.\n\nA useful semantic decomposition is:\n\nEntity\n\n ├── Dimensions\n\n ├── Facts\n\n └── Metrics\n\nFor an e-commerce domain:\n\nCustomer\n\n ├── customer_id\n\n ├── country\n\n ├── segment\n\n └── signup_date\n\nOrder\n\n ├── order_id\n\n ├── order_date\n\n ├── status\n\n └── order_amount\n\nProduct\n\n ├── product_id\n\n ├── category\n\n ├── brand\n\n └── price\n\nThen metrics:\n\nTotal Revenue\n\nAverage Order Value\n\nOrder Count\n\nCustomer Count\n\nConversion Rate\n\nThe important transformation is:\n\nSUM(amount)\n\nbecomes:\n\nTotal Revenue\n\nand:\n\nCOUNT(DISTINCT order_id)\n\nOrder Count\n\nNow the model doesn't need to rediscover the meaning every time.\n\nA semantic model is not merely a dictionary of column names.\n\nIt is also a representation of relationships.\n\nCustomer\n\n    │\n\n    │ 1:N\n\n    ▼\n\nOrder\n\n    │\n\n    │ 1:N\n\n    ▼\n\nOrder Item\n\n    │\n\n    │ N:1\n\n    ▼\n\nProduct\n\nThese relationships constrain how questions can be answered.\n\n\"Revenue by product category\"\n\nrequires a valid path:\n\nRevenue\n\n  ↓\n\nOrder\n\n  ↓\n\nOrder Item\n\n  ↓\n\nProduct\n\n  ↓\n\nCategory\n\nWithout explicit relationship information, an AI system may have to infer the join path from schema names.\n\nThat is precisely the kind of schema reasoning that text-to-SQL research has identified as difficult.\n\nSnowflake's semantic-view guidance similarly emphasizes explicitly defining relationships required by the questions the model must answer.\n\nConsider this calculation:\n\nSUM(revenue) / NULLIF(COUNT(DISTINCT order_id), 0)\n\nTechnically:\n\nSQL expression\n\nSemantically:\n\nAverage Order Value\n\nOnce defined as a metric, it becomes reusable.\n\nInstead of every analyst writing:\n\nwe have:\n\nwith a single authoritative definition.\n\nSnowflake's current semantic-view model supports reusable metrics and filters specifically for this purpose.\n\nThis matters enormously for AI-generated SQL.\n\nAn AI shouldn't have to invent:\n\n\"What exactly does this organization mean by active customer?\"\n\nThe semantic model should already know.\n\nThis is where semantic modeling becomes much more interesting.\n\nWithout semantic modeling:\n\nUser\n\n ↓\n\nLLM\n\n ↓\n\nRaw database schema\n\n ↓\n\nInfer relationships\n\n ↓\n\nInfer metric definitions\n\n ↓\n\nGenerate SQL\n\n ↓\n\nHope it's correct\n\nWith semantic modeling:\n\nUser\n\n ↓\n\nNatural-language intent\n\n ↓\n\nSemantic concepts\n\n ↓\n\nKnown entities\n\n ↓\n\nKnown relationships\n\n ↓\n\nDefined metrics\n\n ↓\n\nVerified examples\n\n ↓\n\nSQL\n\nThis is not merely metadata.\n\nIt is structured context for reasoning.\n\nSnowflake's current semantic-view architecture supports verified queries — natural-language questions paired with validated SQL — as examples that can help Cortex Analyst understand how similar questions should be answered.\n\nThat creates an interesting bridge:\n\nSemantic Layer\n\n      +\n\nLLM\n\n      +\n\nVerified Queries\n\n      ↓\n\nMore constrained SQL generation\n\nAn LLM has an enormous output space.\n\nSQL doesn't.\n\nA generated query has to satisfy a formal grammar and the target database's schema.\n\nPICARD demonstrated this problem in text-to-SQL: unconstrained language-model generation can produce invalid SQL, while incremental parsing can constrain decoding to valid continuations. The work was published at EMNLP 2021.\n\nThis gives us an important architectural principle:\n\nDon't ask an LLM to infer everything that your data architecture already knows.\n\nIf relationships are known, expose them.\n\nIf metrics are defined, expose them.\n\nIf certain filters are mandatory, encode them.\n\nIf certain queries are verified, preserve them.\n\nThe semantic layer reduces the space of possible interpretations.\n\nThere is another misconception:\n\n\"If the semantic layer is correct, performance is automatically solved.\"\n\nSemantic correctness and execution performance are different dimensions.\n\nSemantic Correctness\n\n        │\n\n        ├── correct grain\n\n        ├── correct joins\n\n        ├── correct metric\n\n        └── correct business definition\n\nExecution Performance\n\n        │\n\n        ├── scan volume\n\n        ├── join cost\n\n        ├── aggregation cost\n\n        ├── pruning\n\n        ├── materialization\n\n        └── warehouse resources\n\nA semantic model can be conceptually excellent and still generate expensive SQL.\n\nTherefore:\n\nModel\n\n ↓\n\nGenerate SQL\n\n ↓\n\nEXPLAIN / PROFILE\n\n ↓\n\nMeasure\n\n ↓\n\nOptimize\n\n ↓\n\nRe-test semantics\n\nSnowflake currently supports materialization of selected semantic-view dimensions and metrics as one mechanism for improving performance, while also providing native semantic-view query syntax and management capabilities.\n\nA common architectural temptation is:\n\n```\n             EVERYTHING\n                 │\n   ┌─────────────┼─────────────┐\n   ▼             ▼             ▼\nSales         Finance        Marketing\n   │             │             │\n   └─────────────┼─────────────┘\n                 ▼\n         1 giant model\n```\n\nThat sounds convenient.\n\nIt often isn't.\n\nSnowflake's current modeling guidance recommends organizing semantic views around business domains and use cases, rather than simply mirroring the database. It also recommends starting with a manageable scope and avoiding irrelevant columns.\n\nA better structure might be:\n\n```\n                Semantic Layer\n                     │\n      ┌──────────────┼──────────────┐\n      ▼              ▼              ▼\nSales Analytics   Customer     Product Analytics\n                  Analytics\n```\n\nThe objective is not:\n\n\"Expose everything.\"\n\nThe objective is:\n\nExpose enough meaning to answer the intended class of questions reliably.\n\nname: csat_score\n\ndescription: \"Score\"\n\nversus:\n\nname: csat_score\n\ndescription: >\n\n  Customer Satisfaction Score measured on a 1–5 scale,\n\n  where 5 indicates the highest satisfaction.\n\nThese are technically similar.\n\nSemantically, they are very different.\n\nThe second provides information that an AI system can actually reason about.\n\nSnowflake's current guidance explicitly calls descriptions one of the most important elements for semantic-view accuracy and recommends clear business-oriented descriptions for tables and columns.\n\nSo metadata becomes part of the model's reasoning context.\n\nA simplified conceptual definition could look like:\n\nname: sales_analytics\n\ndescription: >\n\n  Business analytics model for understanding customer orders,\n\n  revenue, products, and regional sales performance.\n\ntables:\n\nname: customers\n\ndescription: >\n\n  One row per customer.\n\nprimary_key:\n\n  columns:\n\n    - customer_id\n\ndimensions:\n\nname: orders\n\ndescription: >\n\n  One row per customer order.\n\nprimary_key:\n\n  columns:\n\n    - order_id\n\nfacts:\n\nmetrics:\n\nThe exact syntax should always be aligned with the current platform specification; the important architectural idea is the decomposition into business entities, dimensions, facts, metrics and relationships. Snowflake's current semantic-view specification supports these concepts natively.\n\nThis is just as important.\n\nDon't blindly move everything upward.\n\nKeep implementation-specific transformations where they belong.\n\nRaw ingestion\n\n      ↓\n\nCleaning\n\n      ↓\n\nDeduplication\n\n      ↓\n\nNormalization\n\n      ↓\n\nBusiness transformations\n\n      ↓\n\nCurated analytical data\n\n      ↓\n\nSemantic layer\n\n      ↓\n\nBI / AI / Applications\n\nA semantic layer should not become a dumping ground for:\n\nETL\n\nELT\n\ndebugging SQL\n\ntemporary transformations\n\none-off reports\n\napplication-specific formatting\n\nOtherwise we simply move the complexity from one location to another.\n\nThis is perhaps the most important architectural perspective.\n\nThink about a semantic view as a contract between:\n\nData Engineering\n\n        │\n\n        ▼\n\nSemantic Model\n\n        │\n\n        ▼\n\nAnalytics / BI\n\n        │\n\n        ▼\n\nAI Systems\n\n        │\n\n        ▼\n\nApplications\n\nThe contract defines:\n\nWhat does this entity represent?\n\nWhat does this metric mean?\n\nWhat is its grain?\n\nHow are entities related?\n\nWhich filters are valid?\n\nWhich calculations are authoritative?\n\nWhich questions have been verified?\n\nThis makes semantic modeling closer to interface design than simply creating another database view.\n\nA semantic model should be tested like software.\n\nTest 1 — Grain\n\nDoes every logical table have a clearly understood grain?\n\nTest 2 — Join correctness\n\nCan every supported relationship be validated?\n\nTest 3 — Metric correctness\n\nDoes Total Revenue match the authoritative calculation?\n\nTest 4 — Aggregation safety\n\nDoes Revenue by Region equal total Revenue?\n\nTest 5 — Natural-language coverage\n\n\"Revenue by country\"\n\n\"Average order value by month\"\n\n\"Top 10 products\"\n\nCan the system generate correct SQL for each?\n\nTest 6 — Performance\n\nGenerated SQL\n\n     ↓\n\nQuery profile\n\n     ↓\n\nBytes scanned\n\n     ↓\n\nJoin behavior\n\n     ↓\n\nExecution time\n\nTest 7 — Regression\n\nEvery semantic change should be tested against existing verified questions.\n\nThis is where semantic modeling starts looking like software engineering rather than documentation.\n\nAfter working through all of this, the architecture becomes:\n\n```\n              RAW DATA\n                 │\n                 ▼\n          DATA MODELING\n                 │\n                 ▼\n         COMPLEX SQL LOGIC\n                 │\n                 ▼\n          GRAIN ANALYSIS\n                 │\n                 ▼\n      ┌──────────────────────┐\n      │ SEMANTIC DECOMPOSITION│\n      └──────────┬───────────┘\n                 │\n      ┌──────────┼──────────┐\n      ▼          ▼          ▼\n   Entities   Dimensions   Metrics\n      │          │          │\n      └──────────┼──────────┘\n                 ▼\n           Relationships\n                 │\n                 ▼\n          Business Rules\n                 │\n                 ▼\n         Verified Queries\n                 │\n                 ▼\n          SEMANTIC VIEW\n                 │\n      ┌──────────┼──────────┐\n      ▼          ▼          ▼\n     BI         AI       Applications\n                 │\n                 ▼\n           Generated SQL\n                 │\n                 ▼\n            Validation\n                 │\n                 ▼\n           Performance\n                 │\n                 ▼\n             Feedback\n```\n\nThis is what I mean by semantic compression.\n\nWe're taking a complicated physical system and exposing a smaller, more meaningful representation to the consumers that need to reason about it.\n\nThere is a broader lesson here.\n\nThe future of natural-language analytics isn't simply:\n\nLLM + database\n\nIt increasingly looks like:\n\nLLM\n\n +\n\nSemantic representation\n\n +\n\nSchema relationships\n\n +\n\nMetric definitions\n\n +\n\nVerified examples\n\n +\n\nQuery constraints\n\n +\n\nExecution feedback\n\nThe database contains the data.\n\nThe semantic layer contains the meaning needed to reason over that data.\n\nThat distinction becomes increasingly important as AI systems move from answering questions to autonomously generating and executing analytical queries.\n\nLarge SQL queries are not necessarily a problem.\n\nUnstructured meaning is.\n\nA 1,000-line query can be correct.\n\nA 100-line query can be semantically wrong.\n\nAnd a beautifully designed semantic view can still produce expensive SQL if its underlying relationships, grain, or execution strategy are poorly understood.\n\nThe goal, therefore, isn't to eliminate SQL complexity.\n\nIt is to put complexity at the correct architectural boundary.\n\nPhysical layer\n\n→ How data is stored\n\nTransformation layer\n\n→ How data is prepared\n\nSemantic layer\n\n→ What data means\n\nAI / BI layer\n\n→ What users want to know\n\nExecution layer\n\n→ How the answer is computed\n\nThe most useful semantic layer is not the one containing the most metadata.\n\nIt is the one that allows a human or an AI system to move from:\n\n\"What does this data mean?\"\n\nto:\n\n\"Which concepts do I need?\"\n\n\"Which relationships are valid?\"\n\n\"Which metric definition should I use?\"\n\n\"Generate the correct SQL.\"\n\nAnd that leads to the principle I keep coming back to:\n\nDon't make the AI understand your entire database. Give it a semantic representation of the part of the database it actually needs to reason about.\n\nThat is the real purpose of a semantic layer.\n\nResearch & References\n\nYu et al. — “Spider: A Large-Scale Human-Labeled Dataset for Complex and Cross-Domain Semantic Parsing and Text-to-SQL Task,” EMNLP 2018.\n\nA foundational benchmark demonstrating the difficulty of generating SQL across complex, previously unseen database schemas.\n\nWang et al. — “RAT-SQL: Relation-Aware Schema Encoding and Linking for Text-to-SQL Parsers,” ACL 2020.\n\nImportant research on schema representation, relationships, and mapping natural-language concepts to database structures.\n\nScholak, Schucher & Bahdanau — “PICARD: Parsing Incrementally for Constrained Auto-Regressive Decoding from Language Models,” EMNLP 2021.\n\nDemonstrates how constraining language-model generation can improve validity for formal languages such as SQL.\n\nSnowflake — Semantic Views: Modeling and Best Practices.\n\nCurrent guidance covering business-domain modeling, descriptions, relationships, metrics, filters, verified queries, and accuracy iteration.\n\nSnowflake — Semantic View YAML Specification.\n\nCurrent specification for logical tables, dimensions, facts, metrics, relationships, verified queries, tags, and native semantic-view objects.", "url": "https://wpnews.pro/news/the-semantic-compression-problem-engineering-ai-ready-views-for-complex-sql", "canonical_source": "https://dev.to/nikhil_ramank_152ca48266/the-semantic-compression-problem-engineering-ai-ready-views-for-complex-sql-44j6", "published_at": "2026-09-28 17:35:49+00:00", "updated_at": "2026-09-28 17:50:21.991467+00:00", "lang": "en", "topics": ["natural-language-processing", "large-language-models", "ai-agents", "structured-data", "ai-tools"], "entities": ["Snowflake", "Spider benchmark", "RAT-SQL"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/the-semantic-compression-problem-engineering-ai-ready-views-for-complex-sql", "markdown": "https://wpnews.pro/news/the-semantic-compression-problem-engineering-ai-ready-views-for-complex-sql.md", "text": "https://wpnews.pro/news/the-semantic-compression-problem-engineering-ai-ready-views-for-complex-sql.txt", "jsonld": "https://wpnews.pro/news/the-semantic-compression-problem-engineering-ai-ready-views-for-complex-sql.jsonld"}}