Making only the agent that caused the failure reflect beats making all of them, by 22 to 27 points
DoCtOR (arXiv 2608.28264, Aug 28) attacks a specific waste in multi-agent self-reflection: current methods have every agent reflect on a failure, which writes wrong insights into the memory of agents that behaved correctly. It runs automated failure attribution to locate the decisive error step and the decisive error agent, uses counterfactual reasoning to synthesize what that step should have been, then asks only that one agent to reflect. Improvements over initial success rate were 22% on HotPotQA, 26% on ChartQAPro and 27% on Mind2Web, beating Reflexion, Retroformer and COPPER. A secondary result worth stealing on its own: in low-resource settings, reflecting only on steps after the decisive error matched reflecting on the whole trajectory.
↳ Follow the thread