Vibe Coding
1,116 web apps across six models say more verification tools do not buy proportionally better software
arXiv 2608.28795 (28 August 2026) defines an agent's 'verification surface' as its set of self-checking tools (linter, boot probe, shell, screenshot) and makes that the single controlled variable in a minimal coding agent. The authors built 1,116 web applications across six models and eight tool configurations, then had a condition-blind human grade every app against a frozen rubric with automatic probes stress-testing API-observable behavior. This is the first controlled measurement I have seen of the assumption that bolting more checkers onto an agent scales output quality, and it should change how you budget tool slots in a harness.
Source
↳ Follow the thread