{"slug": "context-not-models-what-actually-made-ai-bi-reliable", "title": "Context, Not Models: What Actually Made AI BI Reliable", "summary": "A data engineering analysis argues that reliable enterprise text-to-SQL depends less on model choice than on the context built around the model, citing production write-ups from Uber and LinkedIn and a dbt Labs benchmark in which in-scope accuracy rose from 26.9% in 2023 to 62.5% and 51.2% for two 2026 models. The piece reports that a 2026 paper found added context moved accuracy more than switching between frontier models, and that a semantic layer compiling metric definitions into SQL shrinks the model's task to selecting the right metric and dimensions.", "body_md": "You're in a planning meeting and someone says, \"Can't we just let the LLM write the SQL?\" Everyone looks at the data team. The honest answer is \"yes, if we build the right things around it.\" Here's what those things are.\n\nThat answer has gotten more specific over the last couple of years. In 2024, Uber and LinkedIn published detailed write-ups of **text-to-SQL** systems running against warehouses with hundreds of thousands of datasets (Uber, in a limited release) and millions of tables (LinkedIn). Both teams said what worked and what didn't. Since then, newer models have raised the ceiling, and benchmark authors have started to measure what actually drives accuracy.\n\nThis article walks through that evolution from a data engineering and architecture point of view. It ends with the less glamorous question of whether AI can build the pipelines themselves. The time window starts in 2024, but I pull in sources through 2026 because that is where the strongest comparative evidence sits. I'll say so each time a source falls outside 2024.\n\nOne caution before we start. The numbers below come from different datasets, different question sets, and different definitions of \"accuracy\". **I won't line them up in one table as if they were comparable, and you shouldn't either.** Where a figure is vendor-run, user-reported, or comes from a small test, I say so.\n\n**Q: What is the actual claim here?**\n\nTwo claims, one better supported than the other.\n\nThe second claim is the interesting one for architects. A raw LLM generating SQL against a raw schema is a probabilistic system. Add a semantic layer that compiles metric definitions into SQL, and the model's job changes from \"write correct SQL\" to \"pick the right metric and dimensions\". That second job has a much smaller space of wrong answers.\n\nI couldn't find a controlled measurement of run-to-run variance across model generations. So this is a hypothesis the sources are consistent with, not a measured result.\n\n**Key insight:** Among current frontier models, one 2026 paper found that adding context moved accuracy more than choosing between models (per its abstract). Newer models clearly helped too. In dbt Labs' benchmark, in-scope text-to-SQL accuracy rose from 26.9% in 2023 to 62.5% and 51.2% for two 2026 models. Both levers are real, and the second one is the one you control.\n\nRoughly where each topic gets its evidence:\n\n| Area | What changed | Where the evidence comes from | \n|---|---|---|\n| AI BI / natural-language interfaces | From demos to production and limited-release systems with agents and retrieval | Uber, LinkedIn, Snowflake | \n| Context and **knowledge graphs** | Metadata, query logs, and curated descriptions became first-class inputs |  | \n| Semantic layers | A compile step between the question and the SQL | dbt Labs, Rumiantsau and Fokeev | \n| Reasoning models | Little effect on Semantic Layer queries in one benchmark; no direct evidence otherwise | dbt Labs | \n| Evaluation | Benchmarks like Spider and BIRD stopped being enough | Snowflake, Uber, LinkedIn | \n| Pipeline authoring by agents | Weak in an early-2025 benchmark | ELT-Bench | \n\n**Q: What do I mean by AI BI and natural interaction?**\n\nI mean an interface where a person asks a data question in plain language and gets back a query, a result, or both. The model sits between the person and the warehouse.\n\nThe idea is old. What changed in 2024 is that large companies wrote up their attempts in detail, and those write-ups list their limitations explicitly. Three of them anchor this article:\n\n**Q: Why did this get hard at enterprise scale when the demos looked easy?**\n\nBecause a demo has a schema of ten tables, and an enterprise has far more. Uber's team said that the number of datasets (hundreds of thousands) prevents complete evaluation coverage. LinkedIn's data warehouse holds millions of tables. At that scale the first problem is not generating SQL. It is finding which table to use.\n\nUber's write-up gives some sense of the stakes. Uber's data platform handles about 1.2 million interactive queries a month, and the Operations organization contributes about 36% of them. Query authoring took about 10 minutes before QueryGPT and about 3 minutes with it. The limited release reached about 300 daily active users, and 78% of users said it reduced the time spent writing queries from scratch.\n\nThose are useful numbers with an important caveat. **The 10-to-3-minute figure and the 78% figure are user-reported or estimated, not controlled measurements.** Treat them as a direction, not a benchmark.\n\nSnowflake's engineering post makes the case clearly. It argues that public benchmarks such as Spider and BIRD show 80–90%+ accuracy but fall short on real business use, and it names four gaps:\n\nIn Snowflake's own internal evaluation, GPT-4o with a single prompt scored 51%. The evaluation had 150 questions across sales, marketing, and finance, in three levels (filtering, aggregation, trend analysis), with multiple gold queries per question. Snowflake reports 90%+ accuracy for Cortex Analyst, which uses a semantic model.\n\n**Watch out:** That is a vendor evaluating its own product on its own question set. The direction (context beats a bare prompt) matches what an independent paper reports in its abstract, but the specific numbers shouldn't be quoted as a general truth about the market.\n\n*Placeholder image — replace `images/image-1.jpg` with your generated image (keep the same filename). The full self-contained generation prompt is **Image 1** in `05_image_prompts.md`.*\n\n**Q: If a single prompt isn't enough, what does a production system look like?**\n\nBoth Uber and LinkedIn converged on multi-agent designs, where each agent handles one narrow step.\n\nUber's QueryGPT uses an Intent agent, a Table agent, and a Column Prune agent. The Column Prune agent exists for a very practical reason: to manage token usage on very wide tables. If a table has hundreds of columns, you cannot paste all of them into every prompt.\n\nLinkedIn's SQL Bot, built inside its DARWIN platform, is also multi-agent and is backed by a knowledge graph. The follow-up paper describes three components: a knowledge graph, a text-to-SQL agent that retrieves context, generates queries, and corrects errors, and an interactive chatbot.\n\nRoughly, the flow looks like this. It is my synthesis of the common shape, not a reproduction of either company's diagram, and the dotted feedback edge is a possible extension that neither source describes.\n\n``` php\nflowchart TD\n    A[User question] --> B[Intent classification]\n    B --> C[Question enhancement:<br/>add missing context]\n    C --> D[Table retrieval<br/>over metadata and docs]\n    D --> E[Column pruning<br/>for wide tables]\n    E --> F[SQL generation<br/>incremental steps]\n    F --> G{Executes?}\n    G -- no --> H[Error-driven<br/>self-correction]\n    H --> F\n    G -- yes --> I[Result and generated SQL<br/>shown to the user]\n    I --> J[User feedback and<br/>query logs]\n    J -.->|possible extension| D\n```\n\n**Q: What did these teams learn about each step?**\n\nA few lessons come straight from the write-ups, plus one that is my own reading:\n\nSplitting the pipeline into agents gives you narrow, testable steps. It also costs latency and more prompts to maintain, and errors compound across stages. Each stage has its own failure rate, so end-to-end accuracy is lower than the best stage.\n\nSplitting the pipeline is still worth it if you evaluate each stage separately, which I come back to in the evaluation section.\n\n**Q: What is the \"context\" everyone keeps talking about?**\n\nIt is everything the model needs to know that isn't in the question. For a data system, that means table and column descriptions, which tables are popular, how tables usually join, what the values in a categorical column look like, and what the business means by a given term.\n\nLinkedIn's SQL Bot makes this concrete, and it is the best-documented example I found. Its **knowledge graph** has:\n\nMost of that list is plain metadata management: catalogs, usage statistics, documentation. The AI use case made all of that metadata valuable in a way that \"please document your tables\" never managed on its own.\n\n**Q: Does improving the documentation actually move the numbers?**\n\nLinkedIn reports that dataset certification, meaning better table descriptions, significantly improved retrieval accuracy. That is a statement about the quality of human-maintained metadata. The model didn't get smarter. The input got better.\n\nFor a data architect, this reframes a familiar chore. Table descriptions, column definitions, and ownership tags used to be hygiene. Now they are an input to a production system, and bad ones have a measurable cost.\n\n**Key insight:** If you are deciding where to spend the first month of an AI BI initiative, the evidence here points to metadata quality and retrieval before model selection.\n\nAn earlier paper I came across reports a large accuracy jump from knowledge graphs, but I haven't read it, so I'm leaving the number out. LinkedIn's follow-up paper includes ablation studies on its knowledge graph, but I only read the abstract page, so I can't tell you the per-component results. If you are making an architecture decision on this, read that paper in full.\n\nA knowledge graph that goes stale will confidently feed the model wrong context. Someone has to own ingestion from the catalog, refresh usage statistics, and handle the crowdsourced knowledge. LinkedIn's graph depends on DataHub metadata and query logs. If you don't have a catalog or query logs worth mining, building the graph first is a much bigger project than the diagrams suggest.\n\n**Q: What is a semantic layer, and why does it matter for AI?**\n\nA semantic layer is a governed set of metric and dimension definitions that compiles a request into SQL. You define \"revenue\" once, with its filters and its grain. A consumer, human or model, asks for \"revenue by month by region\", and the layer produces the SQL.\n\nFor AI, the value is that the model no longer writes the SQL. It selects from a defined menu, and a deterministic compiler produces the query. That is the basis for my \"non-deterministic to more deterministic\" hypothesis.\n\nAn illustrative metric definition follows. It is not any specific product's exact syntax, and the names are invented.\n\n```\n# Illustrative only: a generic semantic model, not a specific product's syntax\nsemantic_model:\n  name: orders\n  table: analytics.fct_orders\n  entities:\n    - name: customer\n      key: customer_id\n  dimensions:\n    - name: order_date\n      type: time\n      grain: [day, week, month]\n    - name: region\n      type: categorical\n  measures:\n    - name: net_revenue\n      description: \"Gross order value minus refunds, excluding test accounts\"\n      expr: \"gross_amount - refund_amount\"\n      filter: \"is_test_account = false\"\n      agg: sum\n```\n\nThe lines that matter are the `description` and the `filter`. A model working from raw columns has to guess that test accounts should be excluded. A semantic layer states it once, and every query inherits it.\n\n**Q: Is there evidence that this actually helps?**\n\nYes, with caveats. dbt Labs (Jason Ganz and Benoit Perigaud) published a \"Semantic Layer vs. Text-to-SQL: 2026 Benchmark Update\" in April 2026. It reruns dbt's 2023 benchmark on current models, using the ACME Insurance dataset from data.world with 11 questions and 20 runs each. It compares four configurations:\n\nThe models tested were Claude Opus 4.6, Claude Sonnet 4.6, GPT-5.3 Codex, GPT-5.2, and GPT-4 from November 2023 as the baseline.\n\nThe findings I'm comfortable quoting come in two separate sets. With modeled data:\n\n| Modeled data (dbt Labs, 2026) | Text-to-SQL | Semantic Layer | \n|---|---|---|\n| claude-sonnet-4-6 | 90.0% | 98.2% | \n| gpt-5.3-codex | 84.1% | 100.0% | \n\nFor questions within the Semantic Layer's scope:\n\n| In-scope questions (dbt Labs) | Text-to-SQL | Semantic Layer | \n|---|---|---|\n| 2023, GPT-4 | 26.9% | 83.1% | \n| 2026, Sonnet 4.6 | 62.5% | 100% | \n| 2026, GPT-5.3 Codex | 51.2% | 100% | \n\nI couldn't pin down which of the four configurations the in-scope figures come from, so I've kept them in a separate table. Don't read across the two tables.\n\nTwo things stand out.\n\nFirst, **text-to-SQL improved a lot between 2023 and 2026** on the in-scope questions, going from 26.9% to 62.5% and 51.2% depending on the model. That is real progress from the models alone.\n\nSecond, **the Semantic Layer reached 100% on in-scope questions for both 2026 models.** The compile step plausibly removes most of the room for the model to be creative in the wrong place. Note that the Sonnet figure on modeled data is 98.2%, not 100%, so the model's choice of metrics is still not fully deterministic.\n\nThe benchmark's own caveats matter, and the authors state them:\n\nAdd to that the structural caveat: **this is a vendor benchmarking its own product on a small dataset of 11 questions.** It is a useful signal, not proof.\n\n**Watch out:** \"100% on in-scope questions\" depends entirely on how the scope is drawn. A semantic layer is deterministic about what it can answer. It does nothing for the questions it can't answer, and a user doesn't always know which kind they are asking.\n\n**Q: Does the same pattern show up outside a vendor's own benchmark?**\n\nAn arXiv paper by Rumiantsau and Fokeev, submitted in April 2026, tests Claude Opus 4.7, Claude Sonnet 4.6, and GPT-5.4 on 100 questions against the Contoso retail dataset in ClickHouse. Each model runs twice: with the schema only, and with the schema plus a 4 KB semantic document.\n\nPer the abstract (I could only read the abstract, not the full paper):\n\nThat last sentence is the strongest statement of my thesis in any source I found. Note the qualifier \"within tier\": these are three current frontier models from two vendors, and they landed within a point of each other once they had the same context. The paper's title also mentions hallucination results, which I couldn't extract, so I'm not making any hallucination claim from it.\n\nConceptually, a semantic document is a *document*, not a compiler. It gives the model better context, but it doesn't remove the model's freedom to guess. I'm not comparing its numbers with dbt's, because the datasets, question sets, and scoring differ.\n\nFree-form text-to-SQL can answer novel questions but fails by producing plausible, wrong SQL. A semantic layer only answers what it models but fails by saying \"I can't answer that\". A wrong-but-plausible query can quietly reach a dashboard, while a refusal gets noticed and gets someone to extend the model. For metrics that feed decisions, I'd take the refusal.\n\nI found no documented incident where an AI-generated query led to a bad business decision. Uber's note that hallucinated tables and columns remain unresolved is the closest. So treat the dashboard scenario as a risk, not a track record.\n\n**Q: Did reasoning models change data work?**\n\nI need to be careful here. I couldn't find a dedicated source on how reasoning models changed data work. What I have is a narrow signal: the dbt Labs benchmark found that reasoning effort mattered little for Semantic Layer queries.\n\nThat is at least consistent with the logic of the semantic layer. If the model's job is to choose among defined metrics, there is less to reason about. My guess is that reasoning matters more for open-ended, free-form SQL with multi-step logic, but that is a hypothesis, and no source I found tests it.\n\nThe 2023-to-2026 jump in text-to-SQL accuracy in the dbt benchmark suggests newer models are simply better at SQL. I can't attribute that to reasoning specifically, and I won't.\n\nOne thing worth testing on your own data, as an expectation I can't back with a source: as models get better at syntax, the remaining errors may shift from \"the query failed\" to \"the query answered a slightly different question\". If that holds, evaluation has to look at meaning, not just execution.\n\nIf reasoning effort matters little for constrained queries, then effort spent constraining the task (a semantic layer, curated context) may buy more than effort spent on heavier reasoning. That suggests, but doesn't prove, a priority order. The flip side is that a semantic layer has coverage limits, so open-ended questions still need free-form generation, and for those I have no evidence either way. Measure on your own questions before deciding.\n\n**Q: Why is evaluating text-to-SQL harder than it sounds?**\n\nBecause \"correct\" is slippery. Uber's team lists the problems in its own limitations:\n\nI'm paraphrasing Uber's post here, not quoting it.\n\nUber's evaluation framework measures intent accuracy, a table overlap score from 0 to 1, query execution success, and an LLM-judged similarity to a golden SQL query. Together those cover the classification, retrieval, execution, and output stages.\n\nLinkedIn's benchmark has 130+ questions across 10 product areas. About 60% of them have multiple valid answers. Its LLM-as-judge scores agreed with human evaluation 75% of the time.\n\nPause on that last number. A judge that agrees with humans three times out of four differs from the human call one time in four. That is fine for tracking trends, but it is not good enough to certify a single answer as correct.\n\nThis is my favorite illustration of the evaluation problem, and the LinkedIn sources supply both halves.\n\nThese are not contradictory. They measure different things: the first is user satisfaction, the second is a benchmark rating. I only read the paper's abstract page, so I can't tell you how the 53% benchmark was set up.\n\nThe lesson is clear anyway: **user satisfaction can be high while benchmark-measured correctness is moderate.** One plausible reading is that a \"passes\" rating means the query got the user close enough to edit and run. That is a real kind of value, and it is not the same as being right.\n\n| Metric type | What it tells you | What it hides | \n|---|---|---|\n| User rating | Whether the tool is useful in practice | Silent errors the user didn't catch | \n| Benchmark rating (correct or close to correct) | Whether output matches a gold answer | Valid alternatives, if the gold set is narrow | \n| Execution success | Whether the SQL runs | Whether it answers the right question | \n| LLM-judged similarity | Trend over time, cheaply | Disagreement with humans (75% agreement in LinkedIn's case) | \n\nBelow is a stripped-down sketch of a harness that respects the multiple-valid-answers problem. It compares result sets, not SQL text, and allows several gold queries per question. Snowflake describes using multiple gold queries per question; comparing result sets is my own addition. This is illustrative code, not any company's implementation.\n\n``` python\n# Illustrative sketch: compare results, not SQL strings.\nfrom dataclasses import dataclass\nimport hashlib\nimport json\n\n@dataclass\nclass EvalCase:\n    question: str\n    gold_queries: list[str]  # several valid answers per question\n    ordered: bool = False    # set True for ranking / top-N questions\n\ndef result_fingerprint(rows: list[tuple], ordered: bool = False) -> str:\n    # Order-insensitive by default so ORDER BY differences don't count as failures.\n    # Float formatting can still cause false mismatches; round values before hashing if needed.\n    canon = [json.dumps(r, default=str) for r in rows]\n    if not ordered:\n        canon = sorted(canon)\n    return hashlib.sha256(\"\\n\".join(canon).encode()).hexdigest()\n\ndef evaluate(case: EvalCase, generated_sql: str, run_query) -> dict:\n    try:\n        gold_prints = {\n            result_fingerprint(run_query(q), case.ordered) for q in case.gold_queries\n        }\n    except Exception as exc:\n        # A broken gold query is a test bug, not a model failure.\n        return {\"executed\": None, \"match\": None, \"error\": f\"gold query failed: {exc}\"}\n    try:\n        got = result_fingerprint(run_query(generated_sql), case.ordered)\n    except Exception as exc:\n        return {\"executed\": False, \"match\": False, \"error\": str(exc)}\n    return {\"executed\": True, \"match\": got in gold_prints, \"error\": None}\n```\n\nIt is deliberately small. In practice you'd add per-stage metrics (did retrieval return the right tables?), repeated runs to measure variance, and a human-reviewed sample to calibrate any LLM judge.\n\n**Q: What about measuring run-to-run variance?**\n\nThis is a gap. Uber says non-determinism requires longer observation periods, and the dbt benchmark runs each question 20 times. But I couldn't find a controlled measurement of run-to-run variance across model generations, so my \"less non-deterministic over time\" hypothesis rests on the semantic-layer results and on Uber's note about the problem, not on a direct measurement. If you adopt any of this, measure your own variance. Run each test question multiple times and look at the spread, not just the mean.\n\n*Placeholder image — replace `images/image-2.jpg` with your generated image (keep the same filename). The full self-contained generation prompt is **Image 2** in `05_image_prompts.md`.*\n\nA gold set of real questions takes time to collect, and multiple valid gold queries per question take more. Schemas change, so the set goes stale. An LLM judge needs a human-reviewed sample to calibrate against, and that sample has to be refreshed too. LinkedIn's 75% agreement figure is a reminder that skipping the calibration leaves you with a cheap number you can't fully trust. The cost is real, but without it you are flying on user satisfaction alone.\n\n**Q: Everything so far is about querying data. What about data engineering proper?**\n\nThis is where I have to be straight with you: the evidence is thinner and less flattering.\n\nAlmost all the strong sources I found are about AI BI and text-to-SQL for analytics. I couldn't find a firsthand team post-mortem on pipeline authoring, dbt model generation, orchestration, or self-healing pipelines. What I found was mostly listicles and vendor content, which I set aside.\n\nThe one solid source is academic: **ELT-Bench**, by Jin, Zhu, and Kang (arXiv, April 2025). It evaluates AI agents on building full ELT pipelines:\n\nThe best configuration, Spider-Agent with Claude-3.7-Sonnet and extended thinking, \"correctly generates only 3.9% of data models\". It averaged $4.30 and 89.3 steps per pipeline.\n\n**Watch out:** That is a snapshot of early-2025 tools and models, not the current state. Newer models may do better, and I didn't search for later results. Don't read 3.9% as \"agents can't do this.\" Read it as \"end-to-end pipeline generation was far behind SQL generation when this was measured.\"\n\n**Q: Why would building a pipeline be so much harder than answering a question?**\n\nMy explanation is an interpretation, not something the paper states, and it has three parts:\n\nThat's my reading of a single benchmark, so hold it loosely.\n\nSetting aside the benchmark, there are tasks where I'd expect an LLM assist to be a reasonable bet. None of the sources above measures these, so test them against your own stack:\n\nLinkedIn's data gives one hint about the last item. 80% of SQL Bot sessions use the \"Fix with AI\" debugging feature, which the post describes as needing minimal development effort. Usage is not the same as benefit, but a cheap-to-build feature that most sessions touch suggests assistance fits naturally right at the point of a concrete failure.\n\nFaster drafting shifts work from writing to reviewing. If your team writes twice as much SQL but reviews the same way, review becomes the bottleneck, and reviewers see more plausible-looking code that is subtly wrong. The tooling that helps most here is not generation. It is tests, data contracts, and diff-based validation of outputs.\n\n**Q: How did the data tools themselves change?**\n\nI'll mark the limits of my sourcing up front: I couldn't find a source that documents how specific tools evolved in 2024, such as Databricks Genie, MCP-based data tools, or catalog changes. I'm not going to write a feature history I can't back up.\n\nWhat the sources do show is a pattern in *where* AI got attached. Each point below says what it rests on:\n\nSo the shape of the evolution, from the evidence I have, is a move from \"a model that writes SQL\" toward \"a system that retrieves metadata, consults definitions, generates in steps, and gets graded on your data.\"\n\n```\nConceptual baseline                   Pattern in the sources (Uber, LinkedIn,\n(not from a source)                   Snowflake; my synthesis)\n-------------------                   ------------------------------------\nquestion ──> LLM ──> SQL              question ──> intent ──> enhance\n                                            ──> retrieve tables (catalog + graph)\n                                            ──> prune columns\n                                            ──> generate (stepwise)\n                                            ──> self-correct\n                                            ──> evaluate against your own gold set\n```\n\nIf you're an architect choosing tools, that diagram suggests a shopping list that has little to do with the model:\n\n**Q: Which approach should I pick: free-form text-to-SQL, a semantic layer, or a knowledge-graph-backed system?**\n\nThey are not mutually exclusive, because they solve different problems. dbt Labs' post has a diagram on when to use a Semantic Layer versus text-to-SQL, but I haven't seen its content. What follows is my own recommendation: use the semantic layer for questions you can predict, and text-to-SQL for exploration.\n\n| Dimension | Free-form text-to-SQL | Semantic layer | Knowledge-graph-backed retrieval (LinkedIn-style) | \n|---|---|---|---|\n| Main job | Generate SQL from a question | Compile defined metrics into SQL | Find the right tables, joins, and context | \n| Handles open-ended exploration | Yes | No, only modeled scope | Yes | \n| Consistency across runs (expected by design, not measured here) | Weakest | Strongest | Depends on the generator | \n| Upfront cost | Low | High (modeling) | High (metadata, graph) | \n| Scales to millions of tables | Not on its own | Not the point; scope is curated | Built for it at LinkedIn | \n| Key risk | Plausible wrong answers | Gaps in scope | Stale or poor metadata | \n| Best for | Analysts exploring | Governed KPIs and dashboards | Large, messy warehouses | \n\nA pragmatic architecture combines them:\n\n``` php\nflowchart LR\n    Q[Natural-language question] --> R{In semantic<br/>layer scope?}\n    R -- yes --> S[Semantic layer<br/>compiles governed SQL]\n    R -- no --> K[Knowledge-graph retrieval<br/>plus text-to-SQL]\n    S --> V[Result with<br/>metric definition shown]\n    K --> W[Result with generated SQL<br/>flagged as exploratory]\n```\n\nThe routing step is the hard part, and I'm not aware of a source that evaluates it. The design intent is that governed answers and exploratory answers should look different to the user, so nobody mistakes a draft query for a certified metric.\n\nIf your company has a small set of KPIs that drive decisions, a semantic layer pays for itself quickly. If your problem is a sprawling warehouse where people can't find tables, retrieval and metadata matter more than metric definitions. If you have both problems, you need both. Be wary of any pitch that tells you one component solves everything.\n\n**Q: What does a defensive implementation look like in practice?**\n\nBelow is an illustrative sketch of a wrapper that applies the lessons above. It checks that tables and columns exist (hallucinated tables and columns remain unresolved at Uber), filters out statements that don't look like queries, and attaches the provenance a human needs to review the answer. It is an example of the pattern, not production code.\n\n``` python\n# Illustrative sketch of guardrails around LLM-generated SQL.\nimport re\n\n# A cheap first filter only. It does NOT enforce read-only access:\n# \"SELECT 1; DROP TABLE x\" passes, and WITH can front data-modifying\n# statements in some engines. Real enforcement is a read-only database\n# role plus a proper SQL parser.\nLOOKS_LIKE_QUERY = re.compile(r\"^\\s*(with|select)\\b\", re.IGNORECASE)\n\ndef validate_generated_sql(sql, known_schema, referenced_tables, referenced_columns):\n    # known_schema: {\"table_name\": {\"col_a\", \"col_b\"}}\n    problems = []\n    if not LOOKS_LIKE_QUERY.match(sql):\n        problems.append(\"Statement must start with SELECT or WITH.\")\n    for t in referenced_tables(sql):  # e.g. via a SQL parser\n        if t not in known_schema:\n            problems.append(f\"Unknown table referenced: {t}\")\n    for t, c in referenced_columns(sql):  # (table, column) pairs, also via a parser\n        if t in known_schema and c not in known_schema[t]:\n            problems.append(f\"Unknown column referenced: {t}.{c}\")\n    return problems\n\ndef answer(question, retrieve_tables, generate_sql, run_query,\n           known_schema, referenced_tables, referenced_columns):\n    tables = retrieve_tables(question)          # retrieval over catalog metadata\n    sql = generate_sql(question, tables)        # LLM step\n    problems = validate_generated_sql(sql, known_schema, referenced_tables, referenced_columns)\n    if problems:\n        # Feed errors back for one correction attempt, then give up loudly.\n        sql = generate_sql(question, tables, feedback=problems)\n        problems = validate_generated_sql(sql, known_schema, referenced_tables, referenced_columns)\n        if problems:\n            return {\"status\": \"refused\", \"problems\": problems}\n    return {\n        \"status\": \"ok\",\n        \"sql\": sql,                     # always show the SQL\n        \"tables_used\": tables,          # and where the context came from\n        \"rows\": run_query(sql),         # run it under a read-only database role\n    }\n```\n\nThree design choices are worth copying even if you change everything else:\n\nValidation adds work to every request, and a failed check triggers a second generation call, so slow answers get slower. Refusals also cost something with users: someone who gets \"I can't answer that\" three times may stop using the tool. You can tune this by loosening checks for exploratory answers and keeping them strict for anything labeled governed. Either way, pick the trade-off on purpose, because a wrong answer that arrives quickly is the more expensive failure.\n\n**Q: What would I stand behind as a prediction?**\n\nBased on the sources, and tied to the thesis I opened with:\n\nWhat I'd *not* predict is a specific accuracy number or date. The sources show too much variation in definitions and datasets to extrapolate.\n\nThe through-line of 2024 to 2026, in the evidence I could find:\n\nIf you're a data engineer or architect, here is what I'd do in order. The specific numbers are my rule of thumb, not thresholds from any source.\n\nWhich part of your stack has changed the most since 2024: the interface people use to ask questions, the metadata underneath, or the way you evaluate results? I'd like to hear the numbers from your own setup in the comments.\n\n**Sources referenced**", "url": "https://wpnews.pro/news/context-not-models-what-actually-made-ai-bi-reliable", "canonical_source": "https://dev.to/datatechbridge/context-not-models-what-actually-made-ai-bi-reliable-572l", "published_at": "2026-10-02 03:40:29+00:00", "updated_at": "2026-10-02 03:44:25.932138+00:00", "lang": "en", "topics": ["large-language-models", "ai-agents", "mlops", "ai-tools", "ai-research"], "entities": ["Uber", "LinkedIn", "Snowflake", "dbt Labs", "Spider", "BIRD", "ELT-Bench"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/context-not-models-what-actually-made-ai-bi-reliable", "markdown": "https://wpnews.pro/news/context-not-models-what-actually-made-ai-bi-reliable.md", "text": "https://wpnews.pro/news/context-not-models-what-actually-made-ai-bi-reliable.txt", "jsonld": "https://wpnews.pro/news/context-not-models-what-actually-made-ai-bi-reliable.jsonld"}}