{"slug": "rippling-ran-2100-scored-agent-runs-per-model-on-real-payroll-data-the-cheapest", "title": "Rippling Ran 2,100 Scored Agent Runs Per Model on Real Payroll Data. The Cheapest Model Tied the Most Expensive One.", "summary": "Rippling's President and CPO Matt MacInnis published a study testing 15 AI models on real payroll data, with about 2,100 graded attempts per model. Anthropic's Opus 4.6 led with a 91.0% pass rate at $1,453, while the cheapest model, GLM 5.2, tied the most expensive on accuracy at 88.7% for $621, showing that price and version do not predict quality.", "body_md": "### Rippling Tested 15 AI Models on Real Payroll Data. The Cheapest One Tied the Most Expensive One.\n\n** Rippling’s President and CPO Matt MacInnis published** something most B2B companies have and almost none share: a real test of AI models doing real work inside a real production system.\n\nNot a leaderboard. Not made-up tasks. About 2,100 graded attempts per model, across 15 models, on actual personnel, payroll, and financial records. Questions like headcount by department and tenure distribution. And actions like “increase base salaries of everyone who meets this condition by 10%,” onboard a new hire through the checklist, schedule a termination and route its approvals, enter payment amounts into a pay run from a spreadsheet.\n\nEvery attempt either passed Rippling’s production correctness checks or failed. An attempt that never finished counted as a failure. That is a much harder grader than most published benchmarks use.\n\nThree numbers run through the whole study, and they’re the only vocabulary you need:\n\n**Pass rate.** How often the model did the job correctly. No partial credit.**Cost.** What it cost to run the full set of tests once, in dollars.**The slowest 10%.** How long a task took on the bad days, not the average day. This is the number your customer actually feels.\n\n## The 15 models collapse to three real choices\n\nAlmost every model in the study is beaten outright by another one. What’s left is a three-way decision:\n\n**Opus 4.6 if being right matters most, and it’s nearly a tie.** 91.0% pass rate, $1,453, 154 seconds on the slowest 10%. The next model down, GPT-5.5 med, scored 89.5% for $1,435. That’s an $18 difference in cost, which is nothing, and a 1.5-point difference in accuracy that sits inside the margin of error on a test this size. Add the five months Rippling spent tuning its setup specifically for Opus 4.6, worth a point or two by MacInnis’s own estimate, and the two are a wash. What actually separates them is speed, and that goes to Opus 4.6.**GPT-5.5 low if speed matters most.** 88.8%, $1,308, 130 seconds on the slowest 10%. Cheaper*and*faster than Opus 4.6, giving up 2.2 points of accuracy. Worth noting that OpenAI’s own two settings of the same model land on either side of Opus 4.6, so the setting mattered more than the brand.**GLM 5.2 if speed doesn’t matter at all.** 88.7%, $621, 243 seconds on the slowest 10%. Same accuracy as GPT-5.5 low within a tenth of a point, half the price, and much slower. Every overnight job, backfill, and data enrichment run belongs on something like this.\n\nThe expensive models don’t make that list. Fable 5 matched GPT-5.5 low’s accuracy at 3.3x the price and twice the wait. Opus 5 was beaten on both price and speed by Grok 4.5, which scored 0.2 points behind it for 68% less.\n\n## Seven takeaways for B2B founders:\n\n## #1. Quality flattened at the top. Price did not.\n\nSeven models landed between 88.5% and 89.5%. That is a one-point spread. Add the leader at 91.0% and the whole competitive group is 2.5 points wide.\n\nHere’s what those models cost:\n\nGLM 5.2 and Fable 5 are one tenth of a point apart in accuracy. One costs seven times the other.\n\nIf your process is “use whatever the big lab just shipped,” you’re paying a 3x to 7x premium for a difference your customers cannot detect.\n\n## #2. The newest model was not the best model\n\nOpus 4.6 beat both of the newer Anthropic models on accuracy, cost 42% less than Opus 5, and cost a third of Fable 5. Fable 5 was the most expensive model tested and finished fifth.\n\nVersion numbers are a lab’s internal accounting, not a promise about your business. Nobody at Rippling would have known this without running the test.\n\nThen MacInnis proved it a second time. He re-ran the same 2,100 attempts on Grok 4.6 and posted the result: accuracy dropped from 87.3% to 85.9%, and typical response time nearly doubled, from 71 seconds to 131. The newer model was worse and slower at the same work.\n\nOne caution on reading those two numbers. The 87.3% he quotes for Grok 4.5 doesn’t match the 89.1% in his published table, because it came from a different run on a different day. Only compare results measured in the same run. A model that “improved two points” against a number you measured last quarter has told you nothing.\n\nSo when a new model ships, the default move is not to upgrade. It’s to re-run your own tests. Sometimes the answer is stay.\n\n## #3. Your own tuning is worth about as much as a whole model generation\n\nRippling spent five months tuning its instructions and tools around Opus 4.6. MacInnis estimates that work is worth a point or two of accuracy. The entire spread across the leading group is 2.5 points.\n\nSo the tuning is worth roughly the gap between the best model in the study and the cheapest one. Every other model ran with no tuning at all, which means their scores are floors, not ceilings.\n\nTwo things follow from that. First, your test suite and your instructions are a real asset, and they compound. Second, they’re also a switching cost you’re building against yourself, which is why you want the work to sit outside any one vendor’s model rather than baked into it.\n\nRippling deliberately kept its setup plain: one model, no routing logic, and two general-purpose tools that do all the work. If a bunch of custom plumbing were making the decisions, they’d be measuring the plumbing instead of the model.\n\n## #4. Cheap models are not cheap because they do less work\n\nModels are billed by the volume of text they read and write, measured in tokens. Grok 4.5 used 601,000 per task. Opus 5 used 599,000. Effectively the same amount of work. Grok cost $791. Opus 5 cost $2,509.\n\nThe whole difference is the price list. Grok charges $2 per million read and $6 per million written. Opus 5 charges $5 and $25. Fable 5 charges $10 and $50.\n\nThe cheaper models actually did *more* back-and-forth, not less. They aren’t cutting corners. They’re just priced differently.\n\nWhich means the metric to watch is dollars, not efficiency. With prices and usage both moving constantly, the invoice is the only number that stays honest.\n\n## #5. Model choice is a margin decision, not an engineering one\n\nRun it against your own P&L. If AI usage is your biggest variable cost, a 3x swing in that bill moves gross margin by double digits. Same product, same accuracy, same customer experience.\n\nThat makes this the cheapest margin improvement available to most AI-native B2B companies. No new features, no new headcount, no roadmap change. It’s a setting.\n\nIt also means someone in finance should own the number and review it on the same cadence as any other cost line.\n\n## #6. Speed and price are separate purchases\n\nTwo models with the same accuracy:\n\n- GPT-5.5 low: 88.8%, $1,308, and fast\n- GLM 5.2: 88.7%, $621, and slow\n\nOne is roughly 2.7x quicker. The other is half the price. Nothing is both.\n\nAnd for anything a customer is waiting on, the number to watch is the slowest 10%, not the average. The average is what founders quote in board decks. The slowest 10% is what a customer hits on the day they’re already annoyed.\n\nGLM 5.2 saves $687 against GPT-5.5 low at identical accuracy, and turns a 130-second worst case into a 243-second one. In a live support chat, four minutes of silence is an abandoned session and a ticket you now handle twice. That costs more than the $687.\n\nSo the cheap-model argument flips depending on the job. For work nobody is waiting on, price wins and GLM takes it outright. For live customer work, speed wins first and price second, and the winners are models nobody was calling a bargain. Opus 4.6 is the example: top accuracy, respectable speed, middle of the pack on price. It looks expensive on a cost chart and correct on a speed chart.\n\nMost B2B companies pick one model and run everything through it. The split worth making:\n\n**Customer waiting, live**(support chat, in-product copilot, sales assist on a call). Speed first, then accuracy, then price. Consider answering fast with a cheaper model and escalating to the better one only when the first answer fails a check.**Customer waiting, but not right now**(a report they’ll come back to). Accuracy and price. Speed barely matters.** Nobody waiting**(nightly runs, backfills, data enrichment). Price alone. Take the cheapest model that clears your accuracy bar and let it be slow.\n\nPylon made a related point on the SaaStr AI stage: the tickets your AI deflects are the easy ones. The tickets a fast cheap model handles well are the ones with the least revenue attached. That argues for tiering models by job, not picking one and hoping.\n\n## #7. The best model in the study still fails one job in ten\n\n91.0% is the ceiling here. Best model, fully tuned setup, on the exact work it was tuned for. One job in ten still fails.\n\nFor read-only questions, fine. Someone asks for headcount by department, gets a wrong number, asks again. For “increase the base salaries of all people who meet this condition by 10%” or “schedule a termination and route its approvals,” a 9% failure rate is a different kind of problem entirely.\n\nThe Grok 4.6 re-run produced the failure that should worry you most. It filled in 54% of the required fields and reported that it had filled in 100%. Not a wrong answer. A wrong answer labeled correct.\n\nNo accuracy score protects you from that one, because the model reports the same thing whether it worked or not. It also refused ordinary questions like who in Sales has access to Google Workspace, and crashed more often. All three are invisible in a summary number and very visible in production.\n\nI’ve written before about my own AI agent deleting 2,400+ production database records and then reporting no changes detected, and about a test suite reporting an 88% pass rate when the real number was 48%. Neither was a model quality problem. Both were verification problems.\n\nWhich is the part of this that lasts. The benchmark answers which model to buy, and the answer is “several of them work, take the cheap one.” The harder answer is that checking the work is the product. No model on this list saves you, including the $4,359 one.\n\n## What to do about it\n\n**Build your own test set on your own data.** Pass or fail, no partial credit, and anything that times out counts as a failure. Rippling’s entire advantage here is that they have one.**Re-run it every time a model ships,** including models you never planned to use. The answer may be “stay on the older one.”**Track the bill in dollars,** not usage.**Route by job.** Fast model where a customer is waiting, cheap model where nobody is.**Don’t hard-wire one vendor.** Being able to switch is architecture you build in early or pay for later.**Spend the engineering time on checking and undoing the work,** not on shopping for models. That’s where the 9% lives.\n\n## What this benchmark doesn’t prove\n\nOne company, one kind of work, one setup. MacInnis says plainly it isn’t a general ranking of intelligence and isn’t neutral: every model but Opus 4.6 ran untuned, and several of them want to be prompted differently than they were.\n\nWhich makes the cheap-model result stronger, not weaker. GLM 5.2 hit 88.7% for $621 with no tuning at all. Rippling’s next step is training those models on their own real work, so if untuned is the floor, that’s the number to watch.\n\nIt also looks a lot more like real production work than any synthetic benchmark will. Which is the reason to take it seriously, and the reason to go run your own.\n\nMacInnis’s own read after the Grok 4.6 re-run is the one to sit with: on everyday business tasks, it feels like the models have hit a frontier, at least for now. If that holds even a few quarters, the next round of gains in your product doesn’t come from the model. It comes from your data, your tools, and your checking.", "url": "https://wpnews.pro/news/rippling-ran-2100-scored-agent-runs-per-model-on-real-payroll-data-the-cheapest", "canonical_source": "https://www.saastr.com/rippling-ran-2100-scored-agent-runs-per-model-on-real-payroll-data-the-cheapest-model-tied-the-most-expensive-one/", "published_at": "2026-08-18 14:10:58+00:00", "updated_at": "2026-08-18 14:42:13.434990+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-research", "ai-products"], "entities": ["Rippling", "Matt MacInnis", "Anthropic", "Opus 4.6", "OpenAI", "GPT-5.5", "GLM 5.2", "Fable 5"], "alternates": {"html": "https://wpnews.pro/news/rippling-ran-2100-scored-agent-runs-per-model-on-real-payroll-data-the-cheapest", "markdown": "https://wpnews.pro/news/rippling-ran-2100-scored-agent-runs-per-model-on-real-payroll-data-the-cheapest.md", "text": "https://wpnews.pro/news/rippling-ran-2100-scored-agent-runs-per-model-on-real-payroll-data-the-cheapest.txt", "jsonld": "https://wpnews.pro/news/rippling-ran-2100-scored-agent-runs-per-model-on-real-payroll-data-the-cheapest.jsonld"}}