Fetching from the wire…
Public story · 2026-08-31 · high
It isolates the exact agent and step behind a failure, then reflects only there, gaining 22 to 27 points over three prior baselines.
Why now: DoCtOR's numbers are new to the record as of August 31, 2026, and they give multi-agent builders a measured comparison against Reflexion, Retroformer and COPPER.
DoCtOR pinpoints the single agent and step that caused a multi-agent run to fail, then rewrites only that agent's next move.
Most multi-agent systems do the opposite. A failure triggers a reflection prompt for every agent in the run, including ones that behaved correctly. That waste is specific. DoCtOR's targeted method gains 22 points on HotPotQA, 26 on ChartQAPro and 27 on Mind2Web over initial success rates, beating Reflexion, Retroformer and COPPER.
The method runs automated failure attribution first, then uses counterfactual reasoning to work out what the failing step should have produced instead. Only the agent responsible for that step gets the reflection prompt. Every other agent's memory stays untouched.
A smaller result in the paper holds up on its own. In low-resource settings, reflecting only on the steps after the decisive error matched reflecting on the whole trajectory. The waste this fixes is one I've caused myself. Broadcast reflection writes a wrong lesson into the memory of an agent that did nothing wrong.
Each link below shares sources, entities, or timing with this story.
DoCtOR competes with Reflexion / Shared entity: Reflexion / Same source domain / Shared topic / Earlier coverage / Tension
Linked by a graph relationship (DoCtOR competes with Reflexion); both cover Reflexion; reported by the same outlet (arxiv.org).
GPT benchmarked against HotpotQA / Shared entity: Gains / Same source domain / Earlier coverage / Downstream implication
Linked by a graph relationship (GPT benchmarked against HotpotQA); both cover Gains; reported by the same outlet (arxiv.org).
Shared entity: Mind2Web / Same source domain / Shared topic / Earlier coverage / Tension
Both cover Mind2Web; reported by the same outlet (arxiv.org); overlapping topics (agent, beat).
GPT benchmarked against HotpotQA / Shared entity: Gains / Same source domain / Earlier coverage
Linked by a graph relationship (GPT benchmarked against HotpotQA); both cover Gains; reported by the same outlet (arxiv.org).
DoCtOR competes with Reflexion / Shared entity: Reflexion / Same source domain / Earlier coverage
Linked by a graph relationship (DoCtOR competes with Reflexion); both cover Reflexion; reported by the same outlet (arxiv.org).
GPT benchmarked against HotpotQA / Same source domain / Shared topic / Tension
Linked by a graph relationship (GPT benchmarked against HotpotQA); reported by the same outlet (arxiv.org); overlapping topics (agent, failure).
Linked by a graph relationship (GPT benchmarked against HotpotQA); reported by the same outlet (arxiv.org); overlapping topics (agent, beat).
Linked by a graph relationship (GPT benchmarked against HotpotQA); reported by the same outlet (arxiv.org); overlapping topics (agent, only).