TRACE Framework: Comparing Developer and LLM Biases in Code Evaluation Reveals Systematic Divergences
arXiv·medium signal
Researchers introduce TRACE (Tool for Rubric Analysis in Code Evaluation), a framework for evaluating LLM code judges in realistic interactive settings with partial context and ambiguous intent. The study reveals systematic biases in how both human developers and LLMs evaluate code — biases that diverge in specific, measurable ways. As LLMs are increasingly used as automated code reviewers and judges, understanding these bias patterns is critical for calibrating trust in AI-assisted code quality assessment.