Reddit
APEX-Accounting: Best Frontier Model Scores 56.4% on Criteria but No Model Clears 2.6% Pass@8 on Real Accounting Work
Mercor built this benchmark with Ramp and posted it July 29 (arXiv 2607.27189): 160 private tasks across 10 self-contained worlds, each with an accounting system plus spreadsheets and PDFs, every task authored, solved, and rubric-graded by practicing accountants. Claude-Fable-5 (Max) led at 56.4% Mean Criteria@3 and Muse-Spark-1.1 (xHigh) hit 52.6%, but end-to-end pass rates collapsed — no model above 2.6% Pass@8 on the primary measure. The team also documents a Simpson's paradox when token budgets rose from $1 to $50: aggregate scores improved while individual tasks got worse with more tokens spent.
↳ Follow the thread