Research
ERPBench Scores Screenshot-Only Agents Against the ERP Database, Not the Screen
arXiv 2609.17885 (submitted 15 Sep 2026) argues computer-use agent evaluation is still anchored to general desktop and web tasks while ERP systems, which run finance, procurement, inventory and customer operations, present dense interfaces, coordinated multi-step interactions, and errors that alter persistent business records without surfacing on screen. ERPBench evaluates screenshot-only agents on a live, reproducible ERP system and scores each task against ground-truth values in its database rather than against a rendered view. The paper also ships a production-grade harness, and avoids the proprietary-platform and simulated-approximation problems of existing enterprise benchmarks.
↳ Follow the thread