In a 100-Person Study, Developers Evaluating AI Code Ran It and Little Else
A remote observational study isolated the evaluation step with 100 participants writing secure and functional code for four C linked-list tasks, cycling through five AI-generated suggestions that varied in security and functionality before selecting and editing one. The design separates the question of whether AI-assisted developers ship secure code from the prior question of whether they can tell which suggestion is insecure, what cues they use, and how trust shapes the choice, with a post-study survey and 23 in-depth interviews. This pairs with a separate 21 Sep paper coding 527 free-text responses from researchers who write code, where over half of accounts described simply running the generated code while automated tests and peer review were rare, and confidence inverted with experience: less experienced programmers trusted the AI more than themselves.
↳ Follow the thread