Research
MobileJudgeBench: A Simple Screenshot Baseline Beats Four Purpose-Built LLM Judges for Mobile Agents
Mobile agent benchmarks increasingly grade themselves with LLM judges, but nobody has checked whether those judges are reliable on real trajectories. MobileJudgeBench assembles 931 human-annotated trajectories spanning 6 mobile agent benchmarks, 4 agent models, and 68 apps, then evaluates 6 judge methods — adapted from SPA-Bench, A3 (two modes), AndroidArena, and AgentRewardBench — across multiple LLM backends. The elaborate pipelines don't pay off: a simple baseline judge fed sampled screenshots is competitive with and often exceeds them, and the judge's LLM backbone, not the method, is the decisive factor.
↳ Follow the thread