Agents
MobileCybench grades agent bug reports by replaying them against 495 executable security probes
MobileCybench (arXiv 2609.23980, 2026-09-21) answers the review-bandwidth problem created by agents filing vulnerability reports faster than maintainers can triage them. Instead of matching against known CVEs, it replays each claimed exploit against 495 author-written probes encoding security properties across 13 Android apps, so a triggered probe proves both that the exploit worked and which property broke. Given only an obfuscated APK, OpenCode with GPT-5.6-Sol triggers probes in 53.8 percent of apps as a malicious local app and 16.7 percent as a remote attacker, and building the benchmark surfaced 23 previously unreported vulnerabilities, most since confirmed by maintainers.
Source
↳ Follow the thread