← The Wire
Entity trail

Simpson

Source-backed findings, relationship evidence, citations, and briefing history from the public MindPattern archive.

Briefing refs
1
Findings
1
Edges
0
Sources
1

Corpus findings

  1. 2026-07-30 / reddit-researcherAPEX-Accounting: Best Frontier Model Scores 56.4% on Criteria but No Model Clears 2.6% Pass@8 on Real Accounting WorkMercor built this benchmark with Ramp and posted it July 29 (arXiv 2607.27189): 160 private tasks across 10 self-contained worlds, each with an accounting system plus spreadsheets and PDFs, every task authored, solved, and rubric-graded by practicing accountants. Claude-Fable-5 (Max) led at 56.4% Mean Criteria@3 and Muse-Spark-1.1 (xHigh) hit 52.6%, but end-to-end pass rates collapsed — no model above 2.6% Pass@8 on the primary measure. The team also documents a Simpson's paradox when token budgets rose from $1 to $50: aggregate scores improved while individual tasks got worse with more tokens spent.

Source trail

Graph sources

entity graphfindings textkg entitiesnewsletter issues