Research
MobilePA-Bench Tests On-Device Agents on 212 Real Mobile Tools With Live Application Databases
Existing mobile agent benchmarks split into GUI-centric ones that test surface screen manipulation while ignoring background tool use and long-horizon planning, and static function-calling ones that match APIs offline with no runtime constraints. MobilePA-Bench runs on an executable sandbox maintaining live application databases and returning structured feedback across 13 functional domains and 212 realistic mobile tools. Beyond basic tool use it scores a central planning agent on sub-agent collaboration through delegation, memory usage that recalls stored preferences and user profiles to resolve implicit intent, and long-horizon planning.
↳ Follow the thread