Telling an LLM Judge to Ignore Citation Formatting Kills the Bias and the Measurement Along With It
TraceJudgeBench audits citation-like artifacts in RAG and agent-workflow evaluation across GPT-5.5, Claude Sonnet 4.6, and DeepSeek V4-Flash, with content-equivalent pairs, citation ablations, correctness conflicts, and human-validated quality gaps. Stronger anti-citation prompts cut worse-cited wins from 50.5% to 0%, but some operating points start converting validated moderate-gap decisions into Ties well before the strict stress-test endpoint, even while correctness-conflict accuracy stays at or above 93.0%. TRACE-style decoupling recovers 96.5-100.0% of better-plain resolution, and the authors argue bias suppression, resolution retention, Tie cost, and protocol cost have to be reported together.
↳ Follow the thread