Web-Agent Observation Routing Fails Because Router Labels Are Produced at the Agent's Own Success Rate
This paper (2608.06171, submitted 2026-08-06) measures six browser observation modes (text, pixels, hybrids) across eight site-model cells on VisualWebArena and WebArena, and finds the oracle-routing prize is largely noise: rerunning the same mode on the same tasks flips 12-14% of outcomes, so a second run of a mode already in hand gains about as much as adding a new one. Five routing policies — mode picking, spend-on-strong-mode, a zero-cost rule read off task text, a confidence cascade, and pooled cost tiers — none robustly beat simply fixing one well-chosen mode. What does survive is a cost bound: routing only the tasks no mode solves to the cheapest mode cuts cost 9.5-30.6% in 8 of 8 cells at unchanged success. The stated obstruction is self-limiting — routing supervision is generated at the agent's success rate (correlation 0.95 between label supply and routing opportunity), so weak agents starve the router exactly where it would help most.
↳ Follow the thread