Research
Root-Cause Attribution on Long Agent Traces Fails Because One-Shot Judges Settle Early
Automated root-cause attribution over agent execution logs degrades as traces grow, because relevant evidence is sparse, spread across distant actions, and disconnected from the visible failure, which turns diagnosis into a search problem rather than a judgment. Existing methods use one-shot LLM judgments that settle on a plausible diagnosis early and leave critical evidence in longer traces unexamined. Continual Search nudges the judge over successive turns to keep hunting unresolved evidence, evaluated on four existing benchmarks plus MegaRCA-Mix, a new testbed of 50 human-annotated failure trials on long-horizon execution-heavy tasks.
↳ Follow the thread