Fetching from the wire…
Research2026-08-08 · source-backed
arXiv 2608.06171 measures six browser observation modes across eight site-model cells on VisualWebArena and WebArena. Rerunning the same mode on the same tasks flips 12 to 14% of outcomes, so the oracle-routing prize is largely noise, and five routing policies fail to robustly beat fixing one well-chosen mode. What survives: routing only unsolvable tasks to the cheapest mode cuts cost 9.5 to 30.6% in 8 of 8 cells at unchanged success. The obstruction is self-limiting, with correlation 0.95 between label supply and routing opportunity, so weak agents starve the router exactly where it would help.
Each link below shares sources, entities, or timing with this story.
Luan et al. asked whether existing multi-agent repair methods fix causes or just exploit sampling randomness, and built SymTrace, a replay framework that reconstructs execution up to an intervention anchor from recorded logs and regenerates only the downstream trajectory (arXi...
Li, Huo, and Johnson show that one-way message flow between agents produces neither mimicry nor solo behavior but an entirely novel dynamical state, at identical temperature settings. It's conceptual rather than quantitative, but the implication for orchestrator-worker fan-out...
A controlled study ran five Qwen models over eight cases against a DWSIM simulator, 120 slots per arm, with one instruction as the only difference: request a fresh simulation after a substantive modification. No hard gate. Re-verification happened in 94 of 120 guided slots aga...
CAFE (arXiv 2608.24794) makes corrective feedback an in-trajectory intervention the agent chooses to request, using one shared-parameter model alternating between search-agent and critic roles. Online RL shapes request returns from a prompt-level call-versus-skip success gap;...
Agent Lightning v1.0 (arXiv 2608.17528) inverts the standard agentic RL architecture, and the inversion is the whole point. Normally the training engine owns the environment loop. It drives the agent, collects trajectories, computes rewards. Which means your training setup and...
WebMASLab holds task, tools, and browser fixed and varies only architecture. The Telephone Loop attack exploits cross-agent delegation to create cyclical task loops, averaging 80% success with 0% detection against multi-agent versions of Claude Sonnet 4.5, GPT-5.2, and GPT-5.4...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.