APEX-Accounting Mercor, in partnership with Ramp, introduced APEX-Accounting, a benchmark to assess whether frontier models can perform real accounting tasks such as reconciling accounts and posting transactions. Across nine models, Claude-Fable-5 (Max) leads with 56.4% Mean Criteria@3, while no model scores more than 2.6% Pass^8. The benchmark comprises 160 tasks authored by accounting experts, and leaderboard evals are available on request. arXiv:2607.27189v1 Announce Type: new Abstract: We introduce APEX-Accounting, a benchmark built by Mercor in partnership with Ramp, to assess whether frontier models can do the real work of accountants. Tasks include reconciling accounts, accruing expenses, posting transactions, and producing reports. The private eval set comprises 160 tasks, split across 10 worlds. Each world contains an accounting system, as well as spreadsheets, PDFs, and other files. Every task was authored and solved by experts in accounting and bookkeeping, who also wrote grading rubrics. Across nine frontier models, Claude-Fable-5 Max leads with 56.4% Mean Criteria@3, ahead of Muse-Spark-1.1 xHigh at 52.6%. No model scores more than 2.6% Pass^8 GPT-5.6-Sol Max+Pro and the highest Pass@8 is 21.5% Muse-Spark-1.1 xHigh . We experiment with increasing the token budget from $1 to $50 and observe an instance of Simpson's paradox: scores increase as the token budget increases but within a given budget-constrained harness, scores are lower on tasks where the model spends more tokens. As APEX-Accounting is a closed benchmark, leaderboard evals can be run for any frontier model on request.