Context, Not Models: What Actually Made AI BI Reliable A data engineering analysis argues that reliable enterprise text-to-SQL depends less on model choice than on the context built around the model, citing production write-ups from Uber and LinkedIn and a dbt Labs benchmark in which in-scope accuracy rose from 26.9% in 2023 to 62.5% and 51.2% for two 2026 models. The piece reports that a 2026 paper found added context moved accuracy more than switching between frontier models, and that a semantic layer compiling metric definitions into SQL shrinks the model's task to selecting the right metric and dimensions. You're in a planning meeting and someone says, "Can't we just let the LLM write the SQL?" Everyone looks at the data team. The honest answer is "yes, if we build the right things around it." Here's what those things are. That answer has gotten more specific over the last couple of years. In 2024, Uber and LinkedIn published detailed write-ups of text-to-SQL systems running against warehouses with hundreds of thousands of datasets Uber, in a limited release and millions of tables LinkedIn . Both teams said what worked and what didn't. Since then, newer models have raised the ceiling, and benchmark authors have started to measure what actually drives accuracy. This article walks through that evolution from a data engineering and architecture point of view. It ends with the less glamorous question of whether AI can build the pipelines themselves. The time window starts in 2024, but I pull in sources through 2026 because that is where the strongest comparative evidence sits. I'll say so each time a source falls outside 2024. One caution before we start. The numbers below come from different datasets, different question sets, and different definitions of "accuracy". I won't line them up in one table as if they were comparable, and you shouldn't either. Where a figure is vendor-run, user-reported, or comes from a small test, I say so. Q: What is the actual claim here? Two claims, one better supported than the other. The second claim is the interesting one for architects. A raw LLM generating SQL against a raw schema is a probabilistic system. Add a semantic layer that compiles metric definitions into SQL, and the model's job changes from "write correct SQL" to "pick the right metric and dimensions". That second job has a much smaller space of wrong answers. I couldn't find a controlled measurement of run-to-run variance across model generations. So this is a hypothesis the sources are consistent with, not a measured result. Key insight: Among current frontier models, one 2026 paper found that adding context moved accuracy more than choosing between models per its abstract . Newer models clearly helped too. In dbt Labs' benchmark, in-scope text-to-SQL accuracy rose from 26.9% in 2023 to 62.5% and 51.2% for two 2026 models. Both levers are real, and the second one is the one you control. Roughly where each topic gets its evidence: | Area | What changed | Where the evidence comes from | |---|---|---| | AI BI / natural-language interfaces | From demos to production and limited-release systems with agents and retrieval | Uber, LinkedIn, Snowflake | | Context and knowledge graphs | Metadata, query logs, and curated descriptions became first-class inputs | | | Semantic layers | A compile step between the question and the SQL | dbt Labs, Rumiantsau and Fokeev | | Reasoning models | Little effect on Semantic Layer queries in one benchmark; no direct evidence otherwise | dbt Labs | | Evaluation | Benchmarks like Spider and BIRD stopped being enough | Snowflake, Uber, LinkedIn | | Pipeline authoring by agents | Weak in an early-2025 benchmark | ELT-Bench | Q: What do I mean by AI BI and natural interaction? I mean an interface where a person asks a data question in plain language and gets back a query, a result, or both. The model sits between the person and the warehouse. The idea is old. What changed in 2024 is that large companies wrote up their attempts in detail, and those write-ups list their limitations explicitly. Three of them anchor this article: Q: Why did this get hard at enterprise scale when the demos looked easy? Because a demo has a schema of ten tables, and an enterprise has far more. Uber's team said that the number of datasets hundreds of thousands prevents complete evaluation coverage. LinkedIn's data warehouse holds millions of tables. At that scale the first problem is not generating SQL. It is finding which table to use. Uber's write-up gives some sense of the stakes. Uber's data platform handles about 1.2 million interactive queries a month, and the Operations organization contributes about 36% of them. Query authoring took about 10 minutes before QueryGPT and about 3 minutes with it. The limited release reached about 300 daily active users, and 78% of users said it reduced the time spent writing queries from scratch. Those are useful numbers with an important caveat. The 10-to-3-minute figure and the 78% figure are user-reported or estimated, not controlled measurements. Treat them as a direction, not a benchmark. Snowflake's engineering post makes the case clearly. It argues that public benchmarks such as Spider and BIRD show 80–90%+ accuracy but fall short on real business use, and it names four gaps: In Snowflake's own internal evaluation, GPT-4o with a single prompt scored 51%. The evaluation had 150 questions across sales, marketing, and finance, in three levels filtering, aggregation, trend analysis , with multiple gold queries per question. Snowflake reports 90%+ accuracy for Cortex Analyst, which uses a semantic model. Watch out: That is a vendor evaluating its own product on its own question set. The direction context beats a bare prompt matches what an independent paper reports in its abstract, but the specific numbers shouldn't be quoted as a general truth about the market. Placeholder image — replace images/image-1.jpg with your generated image keep the same filename . The full self-contained generation prompt is Image 1 in 05 image prompts.md . Q: If a single prompt isn't enough, what does a production system look like? Both Uber and LinkedIn converged on multi-agent designs, where each agent handles one narrow step. Uber's QueryGPT uses an Intent agent, a Table agent, and a Column Prune agent. The Column Prune agent exists for a very practical reason: to manage token usage on very wide tables. If a table has hundreds of columns, you cannot paste all of them into every prompt. LinkedIn's SQL Bot, built inside its DARWIN platform, is also multi-agent and is backed by a knowledge graph. The follow-up paper describes three components: a knowledge graph, a text-to-SQL agent that retrieves context, generates queries, and corrects errors, and an interactive chatbot. Roughly, the flow looks like this. It is my synthesis of the common shape, not a reproduction of either company's diagram, and the dotted feedback edge is a possible extension that neither source describes. php flowchart TD A User question -- B Intent classification B -- C Question enhancement: