Fetching from the wire…
Research2026-08-27 · source-backed
Across 192,000 attempted evaluations (185,271 successful) on eight models, including a prior score, revision index or attempt count as context metadata systematically drags the new rating toward the anchor. Seven of eight models show 95% task-stratified bootstrap intervals below zero, with Cohen's d reaching 0.71 (arXiv 2608.25869). Neither chain-of-thought nor an explicit instruction to disregard the metadata removed the effect. Any iterative-refinement pipeline that passes prior scores forward is not running independent judgments, it's running one judgment with expensive decoration.
Each link below shares sources, entities, or timing with this story.
Shared entity: Cohen / Same source domain / Shared topic / Earlier coverage
Both cover Cohen; reported by the same outlet (arxiv.org); overlapping topics (context, model).
Both cover Cohen; reported by the same outlet (arxiv.org); overlapping topics (cohen, model).
Shared entity: Cohen / Same source domain / Earlier coverage / Tension
Both cover Cohen; reported by the same outlet (arxiv.org); earlier Cohen coverage from 2026-08-21.
Both cover Cohen; reported by the same outlet (arxiv.org); earlier Cohen coverage from 2026-04-17.
Same source domain / Shared topic / Tension
Reported by the same outlet (arxiv.org); overlapping topics (context, correction, model); pushes against this story (but).
Shared entity: Cohen / Same source domain / Earlier coverage
Both cover Cohen; reported by the same outlet (arxiv.org); earlier Cohen coverage from 2026-08-14.
Both cover Cohen; reported by the same outlet (arxiv.org); earlier Cohen coverage from 2026-07-31.
Shared entity: Cohen / Shared topic / Earlier coverage
Both cover Cohen; overlapping topics (context, model); earlier Cohen coverage from 2026-06-29.