Skills
Vision-language models all fail at driving a browser to find bugs in AI-generated web apps
Code-driven Agentic Testing has the agent write Playwright code to drive the browser, gather feedback, and explore autonomously rather than following a predefined checklist. Evaluated on CATTest, a benchmark of 102 AI-generated web applications with annotated bugs built for complex interactions and subtle defects, all mainstream VLMs performed poorly. The repo at SleepyWithoutCoffee/CATJudge is real (EMNLP'26, pushed August 25) but has 1 star, so treat this as an honest negative result on autonomous end-to-end testing rather than a shipped tool.
↳ Follow the thread