Markets
Rippling Ran 2,100 Scored Agent Runs Per Model on Real Payroll Data and Found an $18 Gap Between the Best and the Cheap One
Rippling's President and CPO Matt MacInnis published a 15-model benchmark run against production payroll data with a brutal grader: every attempt either passes Rippling's production correctness checks or fails, and runs that never finish count as failures. Opus 4.6 scored 91.0% at $1,453 with 154 seconds on the slowest 10%; GPT-5.5 med scored 89.5% at $1,435, an $18 and 1.5-point difference that sits inside the margin of error. The conclusion for builders is not a model recommendation but a method one: several models work, take the cheap one, and no published leaderboard substitutes for your own pass/fail test set on your own data with no partial credit.
Source
↳ Follow the thread