Full Explanations Raise Trust in AI Code Review but Lower Agreement — 34-Developer Study Finds Moderate Explanation Wins at 89.22%
A within-subjects study with 34 participants (arXiv 2607.24601, July 27) compared three LLM code-review systems: detailed explanation plus review feedback (A), review feedback only (B), and no explanation (C). Full explanations produced the highest perceived trust (M = 3.99/5) but not the highest agreement, while moderate explanations produced the highest agreement at 89.22% — suggesting more explanation prompts developers to question the AI more often. No explanations scored lowest on both. Explanation level did not significantly affect review time, and the most-cited reasons for accept/reject decisions were code readability and correctness. The counterintuitive result matters for tool design: maximizing explanation maximizes trust ratings, not adoption of the recommendation.
↳ Follow the thread