AppZen says its finance models beat frontier AI on five of six audit tests AppZen introduced ZenLM Plus on October 1st, a family of finance-specific language models that the company says outperformed seven frontier models on five of six expense-audit control comparisons, scoring 97.4 F1 on targeted policy-category cases versus 88.3 for GPT-5.6 Sol and 92.4 F1 on non-conforming receipt detection, 7.2 points above Opus 5. ZenLM Plus trailed on receipt verification with 93.3 F1 against 94.1 for both Gemini 3.1 Pro and Sonnet 5, and AppZen plans general availability in December, with the company-run benchmark still to be validated in customer deployments. AppZen says its finance models beat frontier AI on five of six audit tests Anant Kale and Kunal Verma are extending AppZen's expense-audit business with models designed for receipts, policies and finance workflows; general availability is planned for December. By RuntimeWire Staff https://runtimewire.com/author/runtimewire-staff ยท Published Primary source: PR Newswire https://www.prnewswire.com/news-releases/appzen-introduces-zenlm-plus-finance-specialized-language-models-that-outperform-frontier-models-on-finance-te-tasks-302889866.html Why it matters AppZen's bet is that finance automation depends on dependable, auditable decisions as much as on model capability. Its benchmark is company-run, so customer deployments will test whether the claimed accuracy and cost advantages hold on real expense policies. Anant Kale https://www.appzen.com/about-us?ref=runtimewire and Kunal Verma https://www.appzen.com/about-us?ref=runtimewire built AppZen https://www.appzen.com/?ref=runtimewire around a particular finance-office problem: software could collect expense reports, but people still had to decide whether the receipts and spending followed policy. On October 1st, AppZen introduced ZenLM Plus, a family of finance-specific models that the company says outperformed seven frontier models on most of the expense-audit controls it tested. The announcement https://www.prnewswire.com/news-releases/appzen-introduces-zenlm-plus-finance-specialized-language-models-that-outperform-frontier-models-on-finance-te-tasks-302889866.html?ref=runtimewire makes a focused case: in this category, task-specific training may count for more than general-purpose model scale. The launch fits the founders' original division of labor. Kale brought experience building enterprise applications, including as a vice president of applications at Fujitsu America; Verma, now AppZen's CTO, was head of AI at Accenture Labs and holds a Ph.D. in computer science from the University of Georgia. Kale recalled that the two neighbors first discussed combining data science with enterprise software while trick-or-treating with their children. Their early product was a mobile assistant for filing expenses. They shifted toward AI auditing after enterprise customers showed stronger interest in the judgment that happened after submission, he told FinTech Global https://fintech.global/2019/09/09/exclusive-appzens-ceo-reveals-how-one-halloween-walk-led-to-the-foundation-of-a-fintech-firm-valued-at-500m/?ref=runtimewire . ZenLM Plus extends that original audit focus. AppZen says its Expense Audit product https://www.appzen.com/ai-for-expense-audit?ref=runtimewire uses the models to assess customer rules, receipt validity, transaction details, itemization and duplicate claims across reports. The system can draw on fields entered by employees, receipt extractions, merchant and card data, and each customer's configured categories and thresholds. That narrow context matters in examples such as telling a payment slip, which can show that a card was charged, from a merchant receipt that substantiates what was bought. What the benchmark says - and what it doesn't AppZen reports in its announcement https://www.prnewswire.com/news-releases/appzen-introduces-zenlm-plus-finance-specialized-language-models-that-outperform-frontier-models-on-finance-te-tasks-302889866.html?ref=runtimewire that ZenLM Plus led five of six individual audit-control comparisons. Across targeted policy-category cases, it scored 97.4 on the F1 scale, against 88.3 for GPT-5.6 Sol https://runtimewire.com/models/native-openai/gpt-5.6-sol-419255f2bb4c6801 , the strongest frontier result in the comparison. AppZen also says its models led all four consolidated policy groups. On non-conforming receipt detection, the reported F1 score was 92.4, 7.2 points above Opus 5. The result was not a clean sweep. On receipt verification, ZenLM Plus scored 93.3, while Gemini 3.1 Pro and Sonnet 5 each scored 94.1. That detail lends the test a more useful shape than a blanket claim of superiority: a customer could use specialized models for controls where they perform well and route another task to a frontier model when it does better. AppZen CEO Kale described that blended approach as part of the product, and Verma said the platform can use a frontier model for a task where it is the better fit. The evaluation remains AppZen's own. Its announcement https://www.prnewswire.com/news-releases/appzen-introduces-zenlm-plus-finance-specialized-language-models-that-outperform-frontier-models-on-finance-te-tasks-302889866.html?ref=runtimewire says each system received the same expense data, supporting documents and customer configuration, then describes precision, recall and the combined F1 score. It does not provide the test-set size, prompts, model settings or an independent evaluation. The reported scores therefore show how ZenLM Plus performed in AppZen's comparison, not how it would perform across every company's expense data or policy rules. AppZen also says in its announcement https://www.prnewswire.com/news-releases/appzen-introduces-zenlm-plus-finance-specialized-language-models-that-outperform-frontier-models-on-finance-te-tasks-302889866.html?ref=runtimewire that ZenLM Plus had the lowest modeled inference cost among the systems tested: about half the cost of GPT-5.6 Luna https://runtimewire.com/models/native-openai/gpt-5.6-luna-e33230db3dd7d2f0 per 1,000 audited expense lines and roughly one-fiftieth the cost of Opus 5. Those estimates could matter in a workflow that checks large volumes of transactions, but the release does not lay out the full cost methodology. Buyers will need to weigh the claimed savings against performance on their own policies and document mix. From models to a finance workflow AppZen is selling ZenLM Plus as part of a broader operating layer, not as a standalone model endpoint. Its Mastermind Platform https://www.appzen.com/mastermind-ai-automation-platform?ref=runtimewire is designed to route work among its specialized models, deterministic rules and frontier models, then apply that output within workflows and controls. The product pitch centers on the surrounding process - including the evidence and audit trail - where finance teams need decisions to be reviewable, rather than competing only on general-purpose model benchmarks. That approach also builds on AppZen's existing focus on expense audits and accounts payable. In September 2025, the company raised $180 million in growth funding https://www.prnewswire.com/news-releases/appzen-raises-180-million-growth-round-led-by-riverwood-capital-to-take-the-next-step-in-autonomous-finance-302559589.html?ref=runtimewire , led by Riverwood Capital https://www.riverwoodcapital.com/?ref=runtimewire , saying the money would support its agent platform and expansion. The financing announcement named Riverwood as lead investor and said its partners joined AppZen's board. ZenLM Plus gives that broader automation pitch a more specific product claim: models trained for the finance decisions inside the workflow. For now, access is limited to select Expense Audit customers, according to the release https://www.prnewswire.com/news-releases/appzen-introduces-zenlm-plus-finance-specialized-language-models-that-outperform-frontier-models-on-finance-te-tasks-302889866.html?ref=runtimewire , with general availability planned for December 2026. AppZen has not named those early customers or provided pricing for ZenLM Plus, according to its product announcement https://www.appzen.com/blog/introducing-zenlm-plus-language-models?ref=runtimewire . The next proof point will come from deployment: whether customers see the benchmark advantage on their own policies, and whether lower modeled inference cost translates into wider audit coverage without sacrificing accuracy or human oversight.