Fetching from the wire…
Agents2026-09-15 · source-backed
A paired experiment ran five-agent business-intelligence reporting with exactly one variable changed, whether a Manager could reject and request revisions, across 43 paired products and 86 runs judged by a five-model panel plus deterministic spec checks (arXiv 2609.14767). Flat organizations won on Utility (d = 0.42, p = 0.009) and Writing Clarity (d = 0.34, p = 0.030). Hierarchical reports hedged 53% more, each revision loop cost 0.14 points of clarity, and the supervisory tier burned 51.5% more tokens for no quality gain. The rule the authors land on is the one I'd put on a wall: a supervisor pays for itself when it can verify, and becomes a liability when it can only opine. Most orchestrator agents I've seen can only opine.
Each link below shares sources, entities, or timing with this story.
Eight teams per setting formed independently from one base model, each agent keeping a private notebook across ten formation episodes, then role-matched agents were traded between teams (arXiv 2609.05279). Against a placebo reproducing roster-change disruption without changing...
A controlled study ran five Qwen models over eight cases against a DWSIM simulator, 120 slots per arm, with one instruction as the only difference: request a fresh simulation after a substantive modification. No hard gate. Re-verification happened in 94 of 120 guided slots aga...
arXiv 2608.26197 stacked finite-state control, forced tool selection, output validation and bounded retries on two open-weight models, and got mixed results across all four model-task cells. Adding structured planning, where the plan is checked against a fixed schema before an...
arXiv 2608.05144 runs Manager, Planner, Engineer and Reviewer roles over persistent project state with *fixed* model weights, self-evolving through runtime state and control policies rather than training. 76.8% on AARRI-Bench, and mature waves use 21% fewer solve-input tokens...
1. Flip your multi-model pipeline to review-then-generate. Instead of using a reasoning model to plan before code generation, let the specialist generate freely and use reasoning tokens for review. Paper shows 90.2% pass@1 vs 87.2% for the planning pattern. Source 2. Audit you...
This paper traces the Emergence World collapses, where agents committed crimes and enforced unanimous conformity with no external attacker, to an enforcement gap rather than a detection gap (arXiv 2609.15293). Reflexion-style self-critique already flags the dangerous step. The...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.