Rippling Shares 2,100-Run AI Agent Benchmark Rippling President and CPO Matt MacInnis said on July 30 that the company ran roughly 2,100 scored attempts per model across 15 AI models on payroll, personnel, and finance tasks, with SaaStr's August 18 analysis finding leading models clustered within a few accuracy points while differing sharply in cost and latency. The benchmark, which included a 91.0% pass rate for Opus 4.6 and 89.5% for GPT-5.5 at its medium setting, underscores that model selection depends on the workload rather than one overall winner. Rippling Shares 2,100-Run AI Agent Benchmark Rippling President and CPO Matt MacInnis said on July 30 that the company ran roughly 2,100 scored attempts per model across 15 AI models on payroll, personnel, and finance tasks. SaaStr's August 18 analysis found the leading models clustered within a few accuracy points while differing sharply in cost and latency, underscoring that model selection depends on the workload rather than one overall winner. Rippling President and CPO Matt MacInnis said on July 30 that the company ran roughly 2,100 scored attempts per model across 15 open-weight and lab models on payroll, personnel, and finance tasks. The company executive described Grok as the value leader and said GLM performed close to the frontier models. SaaStr published a detailed analysis of the shared results on August 18. It says the test covered questions such as headcount and tenure distribution as well as actions including pay changes, onboarding, termination workflows, approvals, and entering payment amounts from a spreadsheet. The benchmark is a company-run production evaluation, not an independently reproduced public leaderboard. Accuracy was only part of the decision SaaStr reports a 91.0% pass rate for Opus 4.6 and 89.5% for GPT-5.5 at its medium setting, while noting that Rippling had spent five months tuning its setup for Opus. It also reports that seven models landed between 88.5% and 89.5%, a narrow range compared with their cost and latency differences. Those results do not establish a universal model ranking. They describe performance inside Rippling's own agent harness, task mix, scoring rules, prompts, and tuning history. The practical value is the evaluation design: score complete business outcomes, retain cost and tail-latency measurements, and test on the records and workflows the system will actually handle. Routing beats a single-model default The published comparison suggests different choices for different constraints. A team may accept a small accuracy tradeoff for faster interactive work, or choose a slower, cheaper model for overnight and batch jobs. That makes model routing and workload segmentation more useful than treating the highest pass rate as the only decision criterion. For teams evaluating agents, the next step is to reproduce this structure on their own tasks. Report the task set, pass/fail rubric, model settings, cost basis, latency percentile, and any model-specific tuning. Without those details, a headline accuracy number is difficult to transfer to another production environment. Key Points - 1Rippling said it ran roughly 2,100 scored attempts per model across 15 models on real payroll, personnel, and finance workflows. - 2SaaStr reported a narrow accuracy band among several leading models but much larger differences in cost and tail latency. - 3Because the benchmark used Rippling's own harness, tasks, scoring, and model-specific tuning, its results should guide workload-specific evaluation rather than a universal ranking. Scoring Rationale The benchmark offers a rare production-oriented comparison across business actions, cost, and tail latency. Its value is practical but bounded because it is company-run, model-specific, and not independently reproduced. Sources Primary source and supporting public references used for this report. Practice interview problems based on real data 1,625 SQL & Python problems across 15 industry datasets — the exact type of data you work with. Try 250 free problems /problems