Research
ComponentBench: Swapping Only the Observation Space Moves the Same Computer-Use Agent by 30 Points
ComponentBench (arXiv 2608.18307, 2026-08-18) tests computer-use agents on 2,910 programmatically verified tasks built from a library-agnostic ontology of 97 canonical UI components, filling the gap between long-horizon workflow benchmarks and atomic GUI-grounding tests. Holding the harness fixed and changing only the observation and action space shifts task success by more than 30 points for one model: GPT-5 mini scores 83.1% with accessibility-tree observations and 48.9% with coordinate-only pixel control. Across seven models (GPT-5.4, Gemini 3 Flash, Qwen3-VL-235B, UI-TARS-1.5-7B and others), even the fastest configuration takes 3.7x as long as the matched human reference trajectory.
↳ Follow the thread