Hybrid GUI-MCP agents leave tools on the table: reasoning models call an available tool on only 23.9% of tool-reachable tasks
Under one identical GUI-MCP harness on OSWorld-MCP's 309 tasks, the same MCP tools improved a reasoning model by +4.0pp and degraded a non-reasoning model by -5.9pp across 5 runs each — availability alone does not determine sign. The authors name the "adoption gap": even the reasoning model called a tool on just 55/309 tasks, 23.9% of tool-reachable ones, and a dense tool bonus in multi-turn RL raised spreadsheet adoption from 0.03 to 0.33 without moving held-out accuracy — behavior is steerable, competence is not. Dropping the now-redundant screenshot after a successful tool call and halving image history cut input tokens by about a third; retrained under that rule the compressed agent hit 37.8% vs 33.0% uncompressed at 53% of the input cost.
Source
↳ Follow the thread