Skills
Agents hold 93% on single API calls and 74% on a 20-step chain, and 77% of the failures got the state right then botched the delivery
APIFlow-Bench tests long-horizon dependent REST workflows with deterministic, provenance-sensitive grading that traces a minted canary through the data flow to the response the answer must come from. Across 19 frontier and open-weight models, success fell from 93% on isolated tasks to 74% on clean 20-subtask chains and 61% including flagged trials; best-case scores spread only 7 points across models while five-of-five reliability spread 44 points. On clean chains 77% of failing runs reached the correct final state and failed only at delivery, which means end-to-end pass/fail hides where your agent is actually breaking.
↳ Follow the thread