527 researchers describe how they check AI-written code: over half just run it
arXiv 2609.22049 (18 Sep) codes 527 free-text survey responses from researchers who write code, mostly at U.S. universities, each recounting one real task, how they used an AI tool for it, and how they judged the result. Use concentrated in five tasks: data handling, visualization, debugging, mathematical/scientific computing and statistical analysis. Over half of accounts described running the generated code as the validation step, while automated tests and review by another person were rare. Tasks and evaluation strategies barely varied with programming experience, but confidence inverted: less experienced programmers trusted the AI more than themselves and experienced programmers the reverse, and evaluation confidence correlated with trust in the tool and in oneself rather than with the rigor of the strategy used.
Source
↳ Follow the thread