Recommending Usability Improvements with Multimodal Large Language Models
arXiv·medium signal
Traditional usability evaluation requires significant expertise and resources. This paper evaluates whether multimodal LLMs can function as automated usability reviewers by analyzing application screenshots and identifying UX issues with actionable improvement recommendations. Results show MLLMs can identify a meaningful subset of expert-found issues, particularly for layout, contrast, and interaction affordance problems. For builders shipping UI: this is a concrete automation of 'does this screen have obvious UX problems' that could run in CI or design review workflows.