cd /news/large-language-models/llm-analytics-benchmark-the-best-mod… · home › topics › large-language-models › article
[ARTICLE · art-147593] src=blog.getcassis.com ↗ pub= topic=large-language-models verified=true sentiment=· neutral

LLM analytics benchmark: the best model for your analytics agent

Cassis built a repeatable benchmark combining 30 Formula 1 questions from the Spider 2.0 database with 60 business cases to compare LLM analytics agents, and found Anthropic's Sonnet 5.5 ranked first on the business questions at 75% correct for $0.12 per question, while Fable 5.1 scored 53% at $0.57. On the Formula 1 set, Sonnet 5, Sonnet 5.5 and Fable 5.1 each matched all 75 executed results across 25 answerable cases repeated three times, while GPT-6 Luna, GPT-6 Sol and Opus 5.5 each matched 72/75 (96%), GPT-6 Astra 69/75 (92%) and GPT-6.1 Sol 66/75 (88%). Cassis said provider benchmarks do not test models on a company's own definitions, so it built the harness-based benchmark to verify answer quality, consistency and cost before swapping the model behind its agent.

by read9 min views2 publishedOct 8, 2026
LLM analytics benchmark: the best model for your analytics agent
Image: source

When an analytics agent gets an answer wrong, there are three possible causes. The information it needed is missing or wrong in the context. The information is there, but the agent didn’t find it. Or the agent found it, and the model misread it or wrote the wrong query. At Cassis, most of our work is on the first two: keeping the context true as the business changes, and structuring it so an agent can find what it needs. We also run our own agent on that context, which answers questions in the Cassis web app and through our MCP server. The third cause depends on the model, and it changes with every release.

The three are not independent. How much a context needs to spell out depends on the model reading it, and a model that reasons more may read more of the context than the question needs. So each new model release brings the same question: should we update the model behind our agent? Benchmarks provided by LLM providers make it tempting, but they don’t test a model on a company’s own definitions. We need to verify our agent behavior before putting a new model into the hands of our users. Testing all of that by hand takes time we cannot spend on every release. So we built a repeatable benchmark to compare answer quality, consistency and cost across models. In this article, we explain the methodology behind our benchmark, and what we’ve learned so far.

Click here to jump directly to the results.

What we test #

We combined 30 cases on a public Formula 1 database with 60 business cases inspired by our internal real usage. Both sets include the cases our agent has to handle: a direct and easy question, a follow-up that modifies the initial request, an ambiguous question, and a very important case for us: questions that available data cannot cover.

The Formula 1 benchmark runs on the Formula 1 database from Spider 2.0, a public academic benchmark: races, drivers, constructors, points. The schema is small and documented, and how a Formula 1 championship works is already in the models’ training data. The benchmark itself may be too. We run each generated query and compare its results with the expected output.

The business benchmark is closer to what our users do. Its 60 questions are inspired by what we see in real usage, on a large business schema with its context: business definitions, metrics and rules. That context is far too large to fit in a model’s context window, so the agent has to find what it needs. We have the schema and the context but not the data, so we don’t run these queries: an LLM judge checks whether the generated SQL is equivalent to the reference.

Our harness is how our agent finds what it needs in that context, the second of the three causes above. It navigates the context as a tree of business domains, reads the tables, metrics and joins it needs, plans the query, then writes it. To measure what the harness adds, we also ran every model without it: the model searches the same context and schema with plain text search (grep) and reads the documents it finds.

The results #

On our business questions, Sonnet 5.5 came first: 75% correct at $0.12 per question. Fable 5.1, the most expensive model we tested, scored 53% at $0.57. The numbers in this section are with our harness, at medium effort: medium-effort results cover all eight selected models, keeping GPT-6 Sol and Sonnet 5 as reference points for their successors. The charts retain all available settings.

The Formula 1 benchmark is easily saturated by most models, making it less effective for fine-grained rankings, though it remains a useful low-cost regression check. With adaptive reasoning and medium effort, Sonnet 5, Sonnet 5.5 and Fable 5.1 each matched all 75 executed results across our 25 answerable Formula 1 cases, repeated three times. GPT-6 Luna, GPT-6 Sol and Opus 5.5 each matched 72/75 (96%); GPT-6 Astra matched 69/75 (92%) and GPT-6.1 Sol 66/75 (88%). Differences of a few points on 25 cases are not significant. When different model and setting combinations achieve identical scores, cost becomes the primary deciding factor.

The business benchmark separates the models: scores spread from 38% to 75%. Small gaps are still noise at this size: Sonnet 5.5 leads Sonnet 5 by 7 points, or 10 conversations out of 150. Its lead over Opus 5.5 (13 points), Fable 5.1 (22 points) and every GPT model (20 points or more) is not. So we are switching our default model to Sonnet 5.5: it scores highest, and it costs less than Sonnet 5, the only model within noise of it. GPT-6 Luna costs 25 times less per question, but answers 20 points fewer questions correctly.

Our harness makes the difference when the context is large. On Formula 1, results with and without it are similar for most models. On the business benchmark, it adds 10 points to Sonnet 5.5 (from 65% to 75%), 11 to GPT-6.1 Sol and 15 to GPT-6 Astra. Fable 5.1 is the exception: 58% without our harness, 53% with it. We can’t explain it yet. One hypothesis is that we tuned our harness on other models.

View results as a table #

Model Harness Effort Score Cost / task
GPT-6 Astra With low 37.3% $0.3845
GPT-6 Astra With medium 40.7% $0.4266
GPT-6 Astra With high 37.3% $0.5431
GPT-6 Astra Without medium 25.3% $0.1766
GPT-6 Luna With No thinking 43.3% $0.0037
GPT-6 Luna With low 46.0% $0.0039
GPT-6 Luna With medium 55.3% $0.0049
GPT-6 Luna With high 52.9% $0.0050
GPT-6 Luna Without medium 53.3% $0.0027
GPT-6 Sol With No thinking 40.7% $0.0931
GPT-6 Sol With low 44.0% $0.0875
GPT-6 Sol With medium 38.7% $0.1112
GPT-6 Sol With high 45.1% $0.1165
GPT-6 Sol Without medium 37.3% $0.0453
GPT-6.1 Sol With low 39.2% $0.0699
GPT-6.1 Sol With medium 38.0% $0.0726
GPT-6.1 Sol With high 35.3% $0.0996
GPT-6.1 Sol Without medium 27.3% $0.0313
Fable 5.1 With low 52.9% $0.5087
Fable 5.1 With medium 53.3% $0.5695
Fable 5.1 With high 63.1% $0.6053
Fable 5.1 Without medium 58.7% $0.8052
Opus 5.5 With low 76.5% $0.2322
Opus 5.5 With medium 62.0% $0.3004
Opus 5.5 With high 58.8% $0.3115
Opus 5.5 Without medium 56.0% $0.4801
Sonnet 5 With No thinking 58.7% $0.1268
Sonnet 5 With low 65.3% $0.1250
Sonnet 5 With medium 68.7% $0.1384
Sonnet 5 With high 62.7% $0.1400
Sonnet 5 Without medium 61.3% $0.2032
Sonnet 5.5 With low 82.4% $0.1154
Sonnet 5.5 With medium 75.3% $0.1237
Sonnet 5.5 With high 76.5% $0.1478
Sonnet 5.5 Without medium 65.3% $0.1223

How we evaluate the answers #

Developing the reference answer is the most time-consuming part. Existing dashboard queries and validated user queries are great starting points, but each reference needs to be carefully checked. We evaluate each model on the same set of questions, with the same context and tools, and repeat each test 3 times to measure consistency.

Two scenarios arise:

A question expects results. On Formula 1 we can compare query results with the expected output. We don’t compare column names or the SQL query, allowing harmless differences such as column aliases. #

A question expects clarification or expects an “I don’t know” answer. These cases and disagreements are assessed by an LLM judge and require human attention. A sensible request for further information shouldn’t be treated as an incorrect answer. The LLM judge is exactly the same (same harness, same model) for each model tested.

What the business failures tell us #

The failures we inspected are of the third kind: the agent had what it needed, and the model didn’t stick to the definition or to the question.

Some models do not follow the expected output conventions. GPT-6.1 Sol often expresses rates as percentages whereas the context defines them as ratios. The business calculation may be correct, but its representation differs from what our evaluation expects. 2. Increased effort can trigger more validation requests. With Opus 5.5 the number of interrupted conversations rises from 3 at the “low” level to 9 at “high” level. The agent offers the user more choices, which triggers a in our harness. These questions were already resolved by the context. 3. Increased effort can lead to going beyond the initial request. In our tests, the “high” mode consults more context (around 30% more) and sometimes adds columns that were not requested. While these additions can be useful, they can lead to potential errors: we observed a few cases where an extra column altered the grouping and changed the results.

To keep costs reasonable, we ran the low and high effort levels on a reduced set of questions. On that set, Opus 5.5 answered 39 of 51 conversations correctly at low effort and 30 at high effort, and the six extra s account for most of that gap. A more careful model should not score lower, so part of this is on us: our harness s whenever the plan offers a choice, even one the context already settled.

What’s next #

Keeping an analytics agent up to date requires a repeatable way to check both new models and changes to the agent. Updating a prompt, adding a tool, or a change in our context format can affect quality, cost and behavior. Our benchmarks help measure these effects and prevent regressions. Inspecting failures is a great way to improve our agent and its evaluation.

Our next step is to evaluate how the agent reaches the answer. How does it navigate through the context? Which domains were read and which relevant definitions were missed? We need to differentiate useful exploration from useless detours. This points to a broader idea: how can we ensure the reliability and consistency of the agent’s approach, beyond its final score?

── more in #large-language-models 4 stories · sorted by recency
── more on @cassis 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/llm-analytics-benchm…] indexed:0 read:9min 2026-10-08 · —