Fetching from the wire…
Public story · 2026-08-27 · high
Testing eight models across 192,000 evaluations, researchers found chain-of-thought and direct instructions to ignore the score didn't remove the bias.
Why now: As of August 27, the paper's numbers are the clearest evidence yet that judge-in-the-loop refinement needs isolation between passes.
A prior score biases the re-grade, according to a study spanning eight LLM judges and 192,000 attempted evaluations. Kapetanovic et al.'s paper on prior-score anchoring fed models a prior score, a revision index, or an attempt count as context before asking for a fresh rating. The new score dragged toward the old one in seven of eight models tested. The 95% bootstrap intervals stayed below zero, with an effect size (Cohen's d) reaching 0.71.
The damage shows up in real grading. On categorical industry data checked against human ground truth, judges shown their prior score missed 48% of the corrections they should have caught. They also flipped 10.18% of already-correct judgments to the wrong label.
Chain-of-thought reasoning didn't fix it. Telling the model outright to ignore the metadata didn't fix it either. The anchor held through both.
That matters for multi-pass grading systems, the kind that re-check a prior verdict or score a draft, revise it, then score it again. If the judge model sees what it said last time, it defends the earlier number instead of testing it fresh. It misses the corrections it should make almost half the time.
Each link below shares sources, entities, or timing with this story.
Same source domain / Semantically similar
Reported by the same outlet (arxiv.org); covers closely related ground (similarity 0.78).
Reported by the same outlet (arxiv.org); covers closely related ground (similarity 0.78).
Reported by the same outlet (arxiv.org); covers closely related ground (similarity 0.77).
Reported by the same outlet (arxiv.org); covers closely related ground (similarity 0.77).
Reported by the same outlet (arxiv.org); covers closely related ground (similarity 0.76).
Reported by the same outlet (arxiv.org); covers closely related ground (similarity 0.75).
Reported by the same outlet (arxiv.org); covers closely related ground (similarity 0.75).
Reported by the same outlet (arxiv.org); covers closely related ground (similarity 0.75).