"The Illusion of Visual Tool-Use": Causal Audit Finds Crop-and-Zoom Observations Often Have Zero Effect on the Model's Answer
A causal audit of the thinking-with-images paradigm (2608.06270, submitted 2026-08-06) formulates visual tool-use as a causal graph separating observation-mediated paths from action-induced shortcuts, then intervenes at three levels: policy (tool-use vs. direct inference), trajectory (corrupting all observations mid-rollout), and step (counterfactually swapping one observation under a fixed prefix). Across six models and five fine-grained perception benchmarks, two failure modes emerge — Calling Without Looking, where returned observations have no causal effect on the answer, and Looking Without Planning, where observations are informative but the call schedule is incoherent. Aggregate accuracy gains are real but concentrated in a small calibrated minority of rollouts, meaning most of the extra token cost buys nothing causal.
↳ Follow the thread