Research
An Exam for Active Observers: Benchmarking Vision Models That Must Choose Where to Look
Jiarui Zhang, Muzi Tao and Shangshang Wang (arXiv 2607.16165) argue that human vision is a closed loop — gaze is redirected by intermediate hypotheses, not fixed on one snapshot — and build an evaluation where models must actively decide the next fixation rather than consume a single image. Most current VLM benchmarks are passive single-shot, so this measures a capability the field mostly doesn't test. Relevant to anyone building screen-reading or computer-use agents, where 'what do I zoom into next' is the actual bottleneck.
Source
↳ Follow the thread