cd /news/large-language-models/context-not-models-what-actually-mad… · home › topics › large-language-models › article
[ARTICLE · art-143608] src=dev.to ↗ pub= topic=large-language-models verified=true sentiment=· neutral

Context, Not Models: What Actually Made AI BI Reliable

A data engineering analysis argues that reliable enterprise text-to-SQL depends less on model choice than on the context built around the model, citing production write-ups from Uber and LinkedIn and a dbt Labs benchmark in which in-scope accuracy rose from 26.9% in 2023 to 62.5% and 51.2% for two 2026 models. The piece reports that a 2026 paper found added context moved accuracy more than switching between frontier models, and that a semantic layer compiling metric definitions into SQL shrinks the model's task to selecting the right metric and dimensions.

by read24 min views2 publishedOct 2, 2026

You're in a planning meeting and someone says, "Can't we just let the LLM write the SQL?" Everyone looks at the data team. The honest answer is "yes, if we build the right things around it." Here's what those things are.

That answer has gotten more specific over the last couple of years. In 2024, Uber and LinkedIn published detailed write-ups of text-to-SQL systems running against warehouses with hundreds of thousands of datasets (Uber, in a limited release) and millions of tables (LinkedIn). Both teams said what worked and what didn't. Since then, newer models have raised the ceiling, and benchmark authors have started to measure what actually drives accuracy.

This article walks through that evolution from a data engineering and architecture point of view. It ends with the less glamorous question of whether AI can build the pipelines themselves. The time window starts in 2024, but I pull in sources through 2026 because that is where the strongest comparative evidence sits. I'll say so each time a source falls outside 2024.

One caution before we start. The numbers below come from different datasets, different question sets, and different definitions of "accuracy". I won't line them up in one table as if they were comparable, and you shouldn't either. Where a figure is vendor-run, user-reported, or comes from a small test, I say so.

Q: What is the actual claim here?

Two claims, one better supported than the other.

The second claim is the interesting one for architects. A raw LLM generating SQL against a raw schema is a probabilistic system. Add a semantic layer that compiles metric definitions into SQL, and the model's job changes from "write correct SQL" to "pick the right metric and dimensions". That second job has a much smaller space of wrong answers.

I couldn't find a controlled measurement of run-to-run variance across model generations. So this is a hypothesis the sources are consistent with, not a measured result.

Key insight: Among current frontier models, one 2026 paper found that adding context moved accuracy more than choosing between models (per its abstract). Newer models clearly helped too. In dbt Labs' benchmark, in-scope text-to-SQL accuracy rose from 26.9% in 2023 to 62.5% and 51.2% for two 2026 models. Both levers are real, and the second one is the one you control.

Roughly where each topic gets its evidence:

Area What changed Where the evidence comes from
AI BI / natural-language interfaces From demos to production and limited-release systems with agents and retrieval Uber, LinkedIn, Snowflake
Context and knowledge graphs Metadata, query logs, and curated descriptions became first-class inputs
Semantic layers A compile step between the question and the SQL dbt Labs, Rumiantsau and Fokeev
Reasoning models Little effect on Semantic Layer queries in one benchmark; no direct evidence otherwise dbt Labs
Evaluation Benchmarks like Spider and BIRD stopped being enough Snowflake, Uber, LinkedIn
Pipeline authoring by agents Weak in an early-2025 benchmark ELT-Bench

Q: What do I mean by AI BI and natural interaction?

I mean an interface where a person asks a data question in plain language and gets back a query, a result, or both. The model sits between the person and the warehouse.

The idea is old. What changed in 2024 is that large companies wrote up their attempts in detail, and those write-ups list their limitations explicitly. Three of them anchor this article:

Q: Why did this get hard at enterprise scale when the demos looked easy?

Because a demo has a schema of ten tables, and an enterprise has far more. Uber's team said that the number of datasets (hundreds of thousands) prevents complete evaluation coverage. LinkedIn's data warehouse holds millions of tables. At that scale the first problem is not generating SQL. It is finding which table to use.

Uber's write-up gives some sense of the stakes. Uber's data platform handles about 1.2 million interactive queries a month, and the Operations organization contributes about 36% of them. Query authoring took about 10 minutes before QueryGPT and about 3 minutes with it. The limited release reached about 300 daily active users, and 78% of users said it reduced the time spent writing queries from scratch.

Those are useful numbers with an important caveat. The 10-to-3-minute figure and the 78% figure are user-reported or estimated, not controlled measurements. Treat them as a direction, not a benchmark.

Snowflake's engineering post makes the case clearly. It argues that public benchmarks such as Spider and BIRD show 80–90%+ accuracy but fall short on real business use, and it names four gaps:

In Snowflake's own internal evaluation, GPT-4o with a single prompt scored 51%. The evaluation had 150 questions across sales, marketing, and finance, in three levels (filtering, aggregation, trend analysis), with multiple gold queries per question. Snowflake reports 90%+ accuracy for Cortex Analyst, which uses a semantic model.

Watch out: That is a vendor evaluating its own product on its own question set. The direction (context beats a bare prompt) matches what an independent paper reports in its abstract, but the specific numbers shouldn't be quoted as a general truth about the market.

Placeholder image — replace images/image-1.jpg with your generated image (keep the same filename). The full self-contained generation prompt is Image 1 in 05_image_prompts.md.

Q: If a single prompt isn't enough, what does a production system look like?

Both Uber and LinkedIn converged on multi-agent designs, where each agent handles one narrow step.

Uber's QueryGPT uses an Intent agent, a Table agent, and a Column Prune agent. The Column Prune agent exists for a very practical reason: to manage token usage on very wide tables. If a table has hundreds of columns, you cannot paste all of them into every prompt.

LinkedIn's SQL Bot, built inside its DARWIN platform, is also multi-agent and is backed by a knowledge graph. The follow-up paper describes three components: a knowledge graph, a text-to-SQL agent that retrieves context, generates queries, and corrects errors, and an interactive chatbot.

Roughly, the flow looks like this. It is my synthesis of the common shape, not a reproduction of either company's diagram, and the dotted feedback edge is a possible extension that neither source describes.

flowchart TD
    A[User question] --> B[Intent classification]
    B --> C[Question enhancement:<br/>add missing context]
    C --> D[Table retrieval<br/>over metadata and docs]
    D --> E[Column pruning<br/>for wide tables]
    E --> F[SQL generation<br/>incremental steps]
    F --> G{Executes?}
    G -- no --> H[Error-driven<br/>self-correction]
    H --> F
    G -- yes --> I[Result and generated SQL<br/>shown to the user]
    I --> J[User feedback and<br/>query logs]
    J -.->|possible extension| D

Q: What did these teams learn about each step?

A few lessons come straight from the write-ups, plus one that is my own reading:

Splitting the pipeline into agents gives you narrow, testable steps. It also costs latency and more prompts to maintain, and errors compound across stages. Each stage has its own failure rate, so end-to-end accuracy is lower than the best stage.

Splitting the pipeline is still worth it if you evaluate each stage separately, which I come back to in the evaluation section.

Q: What is the "context" everyone keeps talking about?

It is everything the model needs to know that isn't in the question. For a data system, that means table and column descriptions, which tables are popular, how tables usually join, what the values in a categorical column look like, and what the business means by a given term.

LinkedIn's SQL Bot makes this concrete, and it is the best-documented example I found. Its knowledge graph has:

Most of that list is plain metadata management: catalogs, usage statistics, documentation. The AI use case made all of that metadata valuable in a way that "please document your tables" never managed on its own.

Q: Does improving the documentation actually move the numbers?

LinkedIn reports that dataset certification, meaning better table descriptions, significantly improved retrieval accuracy. That is a statement about the quality of human-maintained metadata. The model didn't get smarter. The input got better.

For a data architect, this reframes a familiar chore. Table descriptions, column definitions, and ownership tags used to be hygiene. Now they are an input to a production system, and bad ones have a measurable cost.

Key insight: If you are deciding where to spend the first month of an AI BI initiative, the evidence here points to metadata quality and retrieval before model selection.

An earlier paper I came across reports a large accuracy jump from knowledge graphs, but I haven't read it, so I'm leaving the number out. LinkedIn's follow-up paper includes ablation studies on its knowledge graph, but I only read the abstract page, so I can't tell you the per-component results. If you are making an architecture decision on this, read that paper in full.

A knowledge graph that goes stale will confidently feed the model wrong context. Someone has to own ingestion from the catalog, refresh usage statistics, and handle the crowdsourced knowledge. LinkedIn's graph depends on DataHub metadata and query logs. If you don't have a catalog or query logs worth mining, building the graph first is a much bigger project than the diagrams suggest.

Q: What is a semantic layer, and why does it matter for AI?

A semantic layer is a governed set of metric and dimension definitions that compiles a request into SQL. You define "revenue" once, with its filters and its grain. A consumer, human or model, asks for "revenue by month by region", and the layer produces the SQL.

For AI, the value is that the model no longer writes the SQL. It selects from a defined menu, and a deterministic compiler produces the query. That is the basis for my "non-deterministic to more deterministic" hypothesis.

An illustrative metric definition follows. It is not any specific product's exact syntax, and the names are invented.

semantic_model:
  name: orders
  table: analytics.fct_orders
  entities:
    - name: customer
      key: customer_id
  dimensions:
    - name: order_date
      type: time
      grain: [day, week, month]
    - name: region
      type: categorical
  measures:
    - name: net_revenue
      description: "Gross order value minus refunds, excluding test accounts"
      expr: "gross_amount - refund_amount"
      filter: "is_test_account = false"
      agg: sum

The lines that matter are the description and the filter. A model working from raw columns has to guess that test accounts should be excluded. A semantic layer states it once, and every query inherits it.

Q: Is there evidence that this actually helps?

Yes, with caveats. dbt Labs (Jason Ganz and Benoit Perigaud) published a "Semantic Layer vs. Text-to-SQL: 2026 Benchmark Update" in April 2026. It reruns dbt's 2023 benchmark on current models, using the ACME Insurance dataset from data.world with 11 questions and 20 runs each. It compares four configurations:

The models tested were Claude Opus 4.6, Claude Sonnet 4.6, GPT-5.3 Codex, GPT-5.2, and GPT-4 from November 2023 as the baseline.

The findings I'm comfortable quoting come in two separate sets. With modeled data:

Modeled data (dbt Labs, 2026) Text-to-SQL Semantic Layer
claude-sonnet-4-6 90.0% 98.2%
gpt-5.3-codex 84.1% 100.0%

For questions within the Semantic Layer's scope:

In-scope questions (dbt Labs) Text-to-SQL Semantic Layer
2023, GPT-4 26.9% 83.1%
2026, Sonnet 4.6 62.5% 100%
2026, GPT-5.3 Codex 51.2% 100%

I couldn't pin down which of the four configurations the in-scope figures come from, so I've kept them in a separate table. Don't read across the two tables.

Two things stand out.

First, text-to-SQL improved a lot between 2023 and 2026 on the in-scope questions, going from 26.9% to 62.5% and 51.2% depending on the model. That is real progress from the models alone.

Second, the Semantic Layer reached 100% on in-scope questions for both 2026 models. The compile step plausibly removes most of the room for the model to be creative in the wrong place. Note that the Sonnet figure on modeled data is 98.2%, not 100%, so the model's choice of metrics is still not fully deterministic.

The benchmark's own caveats matter, and the authors state them:

Add to that the structural caveat: this is a vendor benchmarking its own product on a small dataset of 11 questions. It is a useful signal, not proof.

Watch out: "100% on in-scope questions" depends entirely on how the scope is drawn. A semantic layer is deterministic about what it can answer. It does nothing for the questions it can't answer, and a user doesn't always know which kind they are asking.

Q: Does the same pattern show up outside a vendor's own benchmark?

An arXiv paper by Rumiantsau and Fokeev, submitted in April 2026, tests Claude Opus 4.7, Claude Sonnet 4.6, and GPT-5.4 on 100 questions against the Contoso retail dataset in ClickHouse. Each model runs twice: with the schema only, and with the schema plus a 4 KB semantic document.

Per the abstract (I could only read the abstract, not the full paper):

That last sentence is the strongest statement of my thesis in any source I found. Note the qualifier "within tier": these are three current frontier models from two vendors, and they landed within a point of each other once they had the same context. The paper's title also mentions hallucination results, which I couldn't extract, so I'm not making any hallucination claim from it.

Conceptually, a semantic document is a document, not a compiler. It gives the model better context, but it doesn't remove the model's freedom to guess. I'm not comparing its numbers with dbt's, because the datasets, question sets, and scoring differ.

Free-form text-to-SQL can answer novel questions but fails by producing plausible, wrong SQL. A semantic layer only answers what it models but fails by saying "I can't answer that". A wrong-but-plausible query can quietly reach a dashboard, while a refusal gets noticed and gets someone to extend the model. For metrics that feed decisions, I'd take the refusal.

I found no documented incident where an AI-generated query led to a bad business decision. Uber's note that hallucinated tables and columns remain unresolved is the closest. So treat the dashboard scenario as a risk, not a track record.

Q: Did reasoning models change data work?

I need to be careful here. I couldn't find a dedicated source on how reasoning models changed data work. What I have is a narrow signal: the dbt Labs benchmark found that reasoning effort mattered little for Semantic Layer queries.

That is at least consistent with the logic of the semantic layer. If the model's job is to choose among defined metrics, there is less to reason about. My guess is that reasoning matters more for open-ended, free-form SQL with multi-step logic, but that is a hypothesis, and no source I found tests it.

The 2023-to-2026 jump in text-to-SQL accuracy in the dbt benchmark suggests newer models are simply better at SQL. I can't attribute that to reasoning specifically, and I won't.

One thing worth testing on your own data, as an expectation I can't back with a source: as models get better at syntax, the remaining errors may shift from "the query failed" to "the query answered a slightly different question". If that holds, evaluation has to look at meaning, not just execution.

If reasoning effort matters little for constrained queries, then effort spent constraining the task (a semantic layer, curated context) may buy more than effort spent on heavier reasoning. That suggests, but doesn't prove, a priority order. The flip side is that a semantic layer has coverage limits, so open-ended questions still need free-form generation, and for those I have no evidence either way. Measure on your own questions before deciding.

Q: Why is evaluating text-to-SQL harder than it sounds?

Because "correct" is slippery. Uber's team lists the problems in its own limitations:

I'm paraphrasing Uber's post here, not quoting it.

Uber's evaluation framework measures intent accuracy, a table overlap score from 0 to 1, query execution success, and an LLM-judged similarity to a golden SQL query. Together those cover the classification, retrieval, execution, and output stages.

LinkedIn's benchmark has 130+ questions across 10 product areas. About 60% of them have multiple valid answers. Its LLM-as-judge scores agreed with human evaluation 75% of the time.

on that last number. A judge that agrees with humans three times out of four differs from the human call one time in four. That is fine for tracking trends, but it is not good enough to certify a single answer as correct.

This is my favorite illustration of the evaluation problem, and the LinkedIn sources supply both halves.

These are not contradictory. They measure different things: the first is user satisfaction, the second is a benchmark rating. I only read the paper's abstract page, so I can't tell you how the 53% benchmark was set up.

The lesson is clear anyway: user satisfaction can be high while benchmark-measured correctness is moderate. One plausible reading is that a "passes" rating means the query got the user close enough to edit and run. That is a real kind of value, and it is not the same as being right.

Metric type What it tells you What it hides
User rating Whether the tool is useful in practice Silent errors the user didn't catch
Benchmark rating (correct or close to correct) Whether output matches a gold answer Valid alternatives, if the gold set is narrow
Execution success Whether the SQL runs Whether it answers the right question
LLM-judged similarity Trend over time, cheaply Disagreement with humans (75% agreement in LinkedIn's case)

Below is a stripped-down sketch of a harness that respects the multiple-valid-answers problem. It compares result sets, not SQL text, and allows several gold queries per question. Snowflake describes using multiple gold queries per question; comparing result sets is my own addition. This is illustrative code, not any company's implementation.

from dataclasses import dataclass
import hashlib
import json

@dataclass
class EvalCase:
    question: str
    gold_queries: list[str]  # several valid answers per question
    ordered: bool = False    # set True for ranking / top-N questions

def result_fingerprint(rows: list[tuple], ordered: bool = False) -> str:
    canon = [json.dumps(r, default=str) for r in rows]
    if not ordered:
        canon = sorted(canon)
    return hashlib.sha256("\n".join(canon).encode()).hexdigest()

def evaluate(case: EvalCase, generated_sql: str, run_query) -> dict:
    try:
        gold_prints = {
            result_fingerprint(run_query(q), case.ordered) for q in case.gold_queries
        }
    except Exception as exc:
        return {"executed": None, "match": None, "error": f"gold query failed: {exc}"}
    try:
        got = result_fingerprint(run_query(generated_sql), case.ordered)
    except Exception as exc:
        return {"executed": False, "match": False, "error": str(exc)}
    return {"executed": True, "match": got in gold_prints, "error": None}

It is deliberately small. In practice you'd add per-stage metrics (did retrieval return the right tables?), repeated runs to measure variance, and a human-reviewed sample to calibrate any LLM judge.

Q: What about measuring run-to-run variance?

This is a gap. Uber says non-determinism requires longer observation periods, and the dbt benchmark runs each question 20 times. But I couldn't find a controlled measurement of run-to-run variance across model generations, so my "less non-deterministic over time" hypothesis rests on the semantic-layer results and on Uber's note about the problem, not on a direct measurement. If you adopt any of this, measure your own variance. Run each test question multiple times and look at the spread, not just the mean.

Placeholder image — replace images/image-2.jpg with your generated image (keep the same filename). The full self-contained generation prompt is Image 2 in 05_image_prompts.md.

A gold set of real questions takes time to collect, and multiple valid gold queries per question take more. Schemas change, so the set goes stale. An LLM judge needs a human-reviewed sample to calibrate against, and that sample has to be refreshed too. LinkedIn's 75% agreement figure is a reminder that skipping the calibration leaves you with a cheap number you can't fully trust. The cost is real, but without it you are flying on user satisfaction alone.

Q: Everything so far is about querying data. What about data engineering proper?

This is where I have to be straight with you: the evidence is thinner and less flattering.

Almost all the strong sources I found are about AI BI and text-to-SQL for analytics. I couldn't find a firsthand team post-mortem on pipeline authoring, dbt model generation, orchestration, or self-healing pipelines. What I found was mostly listicles and vendor content, which I set aside.

The one solid source is academic: ELT-Bench, by Jin, Zhu, and Kang (arXiv, April 2025). It evaluates AI agents on building full ELT pipelines:

The best configuration, Spider-Agent with Claude-3.7-Sonnet and extended thinking, "correctly generates only 3.9% of data models". It averaged $4.30 and 89.3 steps per pipeline.

Watch out: That is a snapshot of early-2025 tools and models, not the current state. Newer models may do better, and I didn't search for later results. Don't read 3.9% as "agents can't do this." Read it as "end-to-end pipeline generation was far behind SQL generation when this was measured."

Q: Why would building a pipeline be so much harder than answering a question?

My explanation is an interpretation, not something the paper states, and it has three parts:

That's my reading of a single benchmark, so hold it loosely.

Setting aside the benchmark, there are tasks where I'd expect an LLM assist to be a reasonable bet. None of the sources above measures these, so test them against your own stack:

LinkedIn's data gives one hint about the last item. 80% of SQL Bot sessions use the "Fix with AI" debugging feature, which the post describes as needing minimal development effort. Usage is not the same as benefit, but a cheap-to-build feature that most sessions touch suggests assistance fits naturally right at the point of a concrete failure.

Faster drafting shifts work from writing to reviewing. If your team writes twice as much SQL but reviews the same way, review becomes the bottleneck, and reviewers see more plausible-looking code that is subtly wrong. The tooling that helps most here is not generation. It is tests, data contracts, and diff-based validation of outputs.

Q: How did the data tools themselves change?

I'll mark the limits of my sourcing up front: I couldn't find a source that documents how specific tools evolved in 2024, such as Databricks Genie, MCP-based data tools, or catalog changes. I'm not going to write a feature history I can't back up.

What the sources do show is a pattern in where AI got attached. Each point below says what it rests on:

So the shape of the evolution, from the evidence I have, is a move from "a model that writes SQL" toward "a system that retrieves metadata, consults definitions, generates in steps, and gets graded on your data."

Conceptual baseline                   Pattern in the sources (Uber, LinkedIn,
(not from a source)                   Snowflake; my synthesis)
-------------------                   ------------------------------------
question ──> LLM ──> SQL              question ──> intent ──> enhance
                                            ──> retrieve tables (catalog + graph)
                                            ──> prune columns
                                            ──> generate (stepwise)
                                            ──> self-correct
                                            ──> evaluate against your own gold set

If you're an architect choosing tools, that diagram suggests a shopping list that has little to do with the model:

Q: Which approach should I pick: free-form text-to-SQL, a semantic layer, or a knowledge-graph-backed system?

They are not mutually exclusive, because they solve different problems. dbt Labs' post has a diagram on when to use a Semantic Layer versus text-to-SQL, but I haven't seen its content. What follows is my own recommendation: use the semantic layer for questions you can predict, and text-to-SQL for exploration.

Dimension Free-form text-to-SQL Semantic layer Knowledge-graph-backed retrieval (LinkedIn-style)
Main job Generate SQL from a question Compile defined metrics into SQL Find the right tables, joins, and context
Handles open-ended exploration Yes No, only modeled scope Yes
Consistency across runs (expected by design, not measured here) Weakest Strongest Depends on the generator
Upfront cost Low High (modeling) High (metadata, graph)
Scales to millions of tables Not on its own Not the point; scope is curated Built for it at LinkedIn
Key risk Plausible wrong answers Gaps in scope Stale or poor metadata
Best for Analysts exploring Governed KPIs and dashboards Large, messy warehouses

A pragmatic architecture combines them:

flowchart LR
    Q[Natural-language question] --> R{In semantic<br/>layer scope?}
    R -- yes --> S[Semantic layer<br/>compiles governed SQL]
    R -- no --> K[Knowledge-graph retrieval<br/>plus text-to-SQL]
    S --> V[Result with<br/>metric definition shown]
    K --> W[Result with generated SQL<br/>flagged as exploratory]

The routing step is the hard part, and I'm not aware of a source that evaluates it. The design intent is that governed answers and exploratory answers should look different to the user, so nobody mistakes a draft query for a certified metric.

If your company has a small set of KPIs that drive decisions, a semantic layer pays for itself quickly. If your problem is a sprawling warehouse where people can't find tables, retrieval and metadata matter more than metric definitions. If you have both problems, you need both. Be wary of any pitch that tells you one component solves everything.

Q: What does a defensive implementation look like in practice?

Below is an illustrative sketch of a wrapper that applies the lessons above. It checks that tables and columns exist (hallucinated tables and columns remain unresolved at Uber), filters out statements that don't look like queries, and attaches the provenance a human needs to review the answer. It is an example of the pattern, not production code.

import re

LOOKS_LIKE_QUERY = re.compile(r"^\s*(with|select)\b", re.IGNORECASE)

def validate_generated_sql(sql, known_schema, referenced_tables, referenced_columns):
    problems = []
    if not LOOKS_LIKE_QUERY.match(sql):
        problems.append("Statement must start with SELECT or WITH.")
    for t in referenced_tables(sql):  # e.g. via a SQL parser
        if t not in known_schema:
            problems.append(f"Unknown table referenced: {t}")
    for t, c in referenced_columns(sql):  # (table, column) pairs, also via a parser
        if t in known_schema and c not in known_schema[t]:
            problems.append(f"Unknown column referenced: {t}.{c}")
    return problems

def answer(question, retrieve_tables, generate_sql, run_query,
           known_schema, referenced_tables, referenced_columns):
    tables = retrieve_tables(question)          # retrieval over catalog metadata
    sql = generate_sql(question, tables)        # LLM step
    problems = validate_generated_sql(sql, known_schema, referenced_tables, referenced_columns)
    if problems:
        sql = generate_sql(question, tables, feedback=problems)
        problems = validate_generated_sql(sql, known_schema, referenced_tables, referenced_columns)
        if problems:
            return {"status": "refused", "problems": problems}
    return {
        "status": "ok",
        "sql": sql,                     # always show the SQL
        "tables_used": tables,          # and where the context came from
        "rows": run_query(sql),         # run it under a read-only database role
    }

Three design choices are worth copying even if you change everything else:

Validation adds work to every request, and a failed check triggers a second generation call, so slow answers get slower. Refusals also cost something with users: someone who gets "I can't answer that" three times may stop using the tool. You can tune this by loosening checks for exploratory answers and keeping them strict for anything labeled governed. Either way, pick the trade-off on purpose, because a wrong answer that arrives quickly is the more expensive failure.

Q: What would I stand behind as a prediction?

Based on the sources, and tied to the thesis I opened with:

What I'd not predict is a specific accuracy number or date. The sources show too much variation in definitions and datasets to extrapolate.

The through-line of 2024 to 2026, in the evidence I could find:

If you're a data engineer or architect, here is what I'd do in order. The specific numbers are my rule of thumb, not thresholds from any source.

Which part of your stack has changed the most since 2024: the interface people use to ask questions, the metadata underneath, or the way you evaluate results? I'd like to hear the numbers from your own setup in the comments.

Sources referenced

── more in #large-language-models 4 stories · sorted by recency
── more on @uber 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/context-not-models-w…] indexed:0 read:24min 2026-10-02 · —