DFAH-Bench: Frontier Models Agree on Financial Decisions 95% of the Time but Follow the Same Tool Path Only 77%
Raffi Khatchadourian's replay benchmark measures observable behavioral instability in tool-using financial agents through three channels — tool-call trajectories, evidence contacts, and decision concentration — none requiring access to hidden reasoning text. Across 8,127 replay episodes spanning 10 models and 3 financial tasks, decision agreement of 95% coexists with tool-path agreement of only 77%, an 18-point gap (95% CI [0.14, 0.22]) that outcome-only evaluation cannot see, and over 55% of high-agreement case groups show meaningful trajectory divergence. It separates agents into pattern matchers that collapse to one output regardless of input, stable executors, and trajectory divergers — a distinction that only shows up if you log the path, not the answer.
↳ Follow the thread