LLM analytics benchmark: the best model for your analytics agent Cassis built a repeatable benchmark combining 30 Formula 1 questions from the Spider 2.0 database with 60 business cases to compare LLM analytics agents, and found Anthropic's Sonnet 5.5 ranked first on the business questions at 75% correct for $0.12 per question, while Fable 5.1 scored 53% at $0.57. On the Formula 1 set, Sonnet 5, Sonnet 5.5 and Fable 5.1 each matched all 75 executed results across 25 answerable cases repeated three times, while GPT-6 Luna, GPT-6 Sol and Opus 5.5 each matched 72/75 (96%), GPT-6 Astra 69/75 (92%) and GPT-6.1 Sol 66/75 (88%). Cassis said provider benchmarks do not test models on a company's own definitions, so it built the harness-based benchmark to verify answer quality, consistency and cost before swapping the model behind its agent. When an analytics agent gets an answer wrong, there are three possible causes. The information it needed is missing or wrong in the context. The information is there, but the agent didn’t find it. Or the agent found it, and the model misread it or wrote the wrong query. At Cassis, most of our work is on the first two: keeping the context true as the business changes, and structuring it so an agent can find what it needs. We also run our own agent on that context, which answers questions in the Cassis web app and through our MCP server. The third cause depends on the model, and it changes with every release. The three are not independent. How much a context needs to spell out depends on the model reading it, and a model that reasons more may read more of the context than the question needs. So each new model release brings the same question: should we update the model behind our agent? Benchmarks provided by LLM providers make it tempting, but they don’t test a model on a company’s own definitions. We need to verify our agent behavior before putting a new model into the hands of our users. Testing all of that by hand takes time we cannot spend on every release. So we built a repeatable benchmark to compare answer quality, consistency and cost across models. In this article, we explain the methodology behind our benchmark, and what we’ve learned so far. Click here to jump directly to the results. the-results What we test We combined 30 cases on a public Formula 1 https://yale-lily.github.io/spider database with 60 business cases inspired by our internal real usage. Both sets include the cases our agent has to handle: a direct and easy question, a follow-up that modifies the initial request, an ambiguous question, and a very important case for us: questions that available data cannot cover. The Formula 1 benchmark runs on the Formula 1 database from Spider 2.0, a public academic benchmark: races, drivers, constructors, points. The schema is small and documented, and how a Formula 1 championship works is already in the models’ training data. The benchmark itself may be too. We run each generated query and compare its results with the expected output. The business benchmark is closer to what our users do. Its 60 questions are inspired by what we see in real usage, on a large business schema with its context: business definitions, metrics and rules. That context is far too large to fit in a model’s context window, so the agent has to find what it needs. We have the schema and the context but not the data, so we don’t run these queries: an LLM judge checks whether the generated SQL is equivalent to the reference. Our harness is how our agent finds what it needs in that context, the second of the three causes above. It navigates the context as a tree of business domains, reads the tables, metrics and joins it needs, plans the query, then writes it. To measure what the harness adds, we also ran every model without it: the model searches the same context and schema with plain text search grep and reads the documents it finds. The results On our business questions, Sonnet 5.5 came first: 75% correct at $0.12 per question. Fable 5.1, the most expensive model we tested, scored 53% at $0.57. The numbers in this section are with our harness, at medium effort: medium-effort results cover all eight selected models, keeping GPT-6 Sol and Sonnet 5 as reference points for their successors. The charts retain all available settings. The Formula 1 benchmark is easily saturated by most models, making it less effective for fine-grained rankings, though it remains a useful low-cost regression check. With adaptive reasoning and medium effort, Sonnet 5, Sonnet 5.5 and Fable 5.1 each matched all 75 executed results across our 25 answerable Formula 1 cases, repeated three times. GPT-6 Luna, GPT-6 Sol and Opus 5.5 each matched 72/75 96% ; GPT-6 Astra matched 69/75 92% and GPT-6.1 Sol 66/75 88% . Differences of a few points on 25 cases are not significant. When different model and setting combinations achieve identical scores, cost becomes the primary deciding factor. The business benchmark separates the models: scores spread from 38% to 75%. Small gaps are still noise at this size: Sonnet 5.5 leads Sonnet 5 by 7 points, or 10 conversations out of 150. Its lead over Opus 5.5 13 points , Fable 5.1 22 points and every GPT model 20 points or more is not. So we are switching our default model to Sonnet 5.5: it scores highest, and it costs less than Sonnet 5, the only model within noise of it. GPT-6 Luna costs 25 times less per question, but answers 20 points fewer questions correctly. Our harness makes the difference when the context is large. On Formula 1, results with and without it are similar for most models. On the business benchmark, it adds 10 points to Sonnet 5.5 from 65% to 75% , 11 to GPT-6.1 Sol and 15 to GPT-6 Astra. Fable 5.1 is the exception: 58% without our harness, 53% with it. We can’t explain it yet. One hypothesis is that we tuned our harness on other models. View results as a table | Model | Harness | Effort | Score | Cost / task | |---|---|---|---|---| | GPT-6 Astra | With | low | 37.3% | $0.3845 | | GPT-6 Astra | With | medium | 40.7% | $0.4266 | | GPT-6 Astra | With | high | 37.3% | $0.5431 | | GPT-6 Astra | Without | medium | 25.3% | $0.1766 | | GPT-6 Luna | With | No thinking | 43.3% | $0.0037 | | GPT-6 Luna | With | low | 46.0% | $0.0039 | | GPT-6 Luna | With | medium | 55.3% | $0.0049 | | GPT-6 Luna | With | high | 52.9% | $0.0050 | | GPT-6 Luna | Without | medium | 53.3% | $0.0027 | | GPT-6 Sol | With | No thinking | 40.7% | $0.0931 | | GPT-6 Sol | With | low | 44.0% | $0.0875 | | GPT-6 Sol | With | medium | 38.7% | $0.1112 | | GPT-6 Sol | With | high | 45.1% | $0.1165 | | GPT-6 Sol | Without | medium | 37.3% | $0.0453 | | GPT-6.1 Sol | With | low | 39.2% | $0.0699 | | GPT-6.1 Sol | With | medium | 38.0% | $0.0726 | | GPT-6.1 Sol | With | high | 35.3% | $0.0996 | | GPT-6.1 Sol | Without | medium | 27.3% | $0.0313 | | Fable 5.1 | With | low | 52.9% | $0.5087 | | Fable 5.1 | With | medium | 53.3% | $0.5695 | | Fable 5.1 | With | high | 63.1% | $0.6053 | | Fable 5.1 | Without | medium | 58.7% | $0.8052 | | Opus 5.5 | With | low | 76.5% | $0.2322 | | Opus 5.5 | With | medium | 62.0% | $0.3004 | | Opus 5.5 | With | high | 58.8% | $0.3115 | | Opus 5.5 | Without | medium | 56.0% | $0.4801 | | Sonnet 5 | With | No thinking | 58.7% | $0.1268 | | Sonnet 5 | With | low | 65.3% | $0.1250 | | Sonnet 5 | With | medium | 68.7% | $0.1384 | | Sonnet 5 | With | high | 62.7% | $0.1400 | | Sonnet 5 | Without | medium | 61.3% | $0.2032 | | Sonnet 5.5 | With | low | 82.4% | $0.1154 | | Sonnet 5.5 | With | medium | 75.3% | $0.1237 | | Sonnet 5.5 | With | high | 76.5% | $0.1478 | | Sonnet 5.5 | Without | medium | 65.3% | $0.1223 | How we evaluate the answers Developing the reference answer is the most time-consuming part. Existing dashboard queries and validated user queries https://blog.getcassis.com/assemble-context-for-analytics-agents/ are great starting points, but each reference needs to be carefully checked. We evaluate each model on the same set of questions, with the same context and tools, and repeat each test 3 times to measure consistency. Two scenarios arise: - A question expects results. On Formula 1 we can compare query results with the expected output. We don’t compare column names or the SQL query, allowing harmless differences such as column aliases. - A question expects clarification or expects an “I don’t know” answer. These cases and disagreements are assessed by an LLM judge and require human attention. A sensible request for further information shouldn’t be treated as an incorrect answer. The LLM judge is exactly the same same harness, same model for each model tested. What the business failures tell us The failures we inspected are of the third kind: the agent had what it needed, and the model didn’t stick to the definition or to the question. 1. Some models do not follow the expected output conventions. GPT-6.1 Sol often expresses rates as percentages whereas the context defines them as ratios. The business calculation may be correct, but its representation differs from what our evaluation expects. 2. Increased effort can trigger more validation requests. With Opus 5.5 the number of interrupted conversations rises from 3 at the “low” level to 9 at “high” level. The agent offers the user more choices, which triggers a pause in our harness. These questions were already resolved by the context. 3. Increased effort can lead to going beyond the initial request. In our tests, the “high” mode consults more context around 30% more and sometimes adds columns that were not requested. While these additions can be useful, they can lead to potential errors: we observed a few cases where an extra column altered the grouping and changed the results. To keep costs reasonable, we ran the low and high effort levels on a reduced set of questions. On that set, Opus 5.5 answered 39 of 51 conversations correctly at low effort and 30 at high effort, and the six extra pauses account for most of that gap. A more careful model should not score lower, so part of this is on us: our harness pauses whenever the plan offers a choice, even one the context already settled. What’s next Keeping an analytics agent up to date requires a repeatable way to check both new models and changes to the agent. Updating a prompt, adding a tool, or a change in our context format can affect quality, cost and behavior. Our benchmarks help measure these effects and prevent regressions. Inspecting failures https://blog.getcassis.com/derivation-distance/ when-more-context-gives-worse-results is a great way to improve our agent and its evaluation. Our next step is to evaluate how the agent reaches the answer. How does it navigate through the context https://blog.getcassis.com/your-data-context-should-be-debuggable/ ? Which domains were read and which relevant definitions were missed? We need to differentiate useful exploration from useless detours. This points to a broader idea: how can we ensure the reliability and consistency of the agent’s approach, beyond its final score?