A Prior Score in Judge Metadata Blocks 48% of Error Corrections and Flips 10.18% of Correct Judgments
Kapetanovic et al. ran 192,000 attempted evaluations (185,271 successful) across eight models and found that including a prior score, revision index, or attempt count as context metadata systematically drags the new rating toward the anchor; seven of eight models show 95% task-stratified bootstrap intervals below zero, with Cohen's d reaching 0.71. On categorical industry data with human ground truth, anchored metadata blocks 48% of error corrections and flips 10.18% of correct judgments to an assigned wrong label. Neither chain-of-thought nor an explicit instruction to disregard the metadata removed the effect, so any iterative-refinement pipeline that passes prior scores forward is not running independent judgments.
↳ Follow the thread