Fetching from the wire…
Research2026-08-30 · source-backed
The strongest standard design for auditing LLM judges, a within-item contrast between two responses differenced across a manipulated attribute on a bounded rating scale, isn't identified on the scale that reports it, because each term is censored by its own share and confounds preference with differential attenuation. In a pre-registered audit of a frozen pedagogy judge sealed before the first of 990 calls, the registered primary endpoint was null at +0.085 (95% BCa [-0.167, +0.353], p = 0.684). The one nominally significant interaction, +0.378 at p = 0.002, was reproduced 79 to 85% by a construction containing zero differential preference, using only the observed severity shift and the scale floor. (arXiv 2608.27309) If you run bounded-scale judge comparisons, your effect may be the floor talking.
Each link below shares sources, entities, or timing with this story.
Simon Willison released LLM / Shared entity: LLM / Earlier coverage / Tension
Linked by a graph relationship (Simon Willison released LLM); both cover LLM; earlier LLM coverage from 2026-08-16.
Linked by a graph relationship (Simon Willison released LLM); both cover LLM; earlier LLM coverage from 2026-07-27.
Linked by a graph relationship (Simon Willison released LLM); both cover LLM; earlier LLM coverage from 2026-06-19.
Linked by a graph relationship (Simon Willison released LLM); both cover LLM; earlier LLM coverage from 2026-06-18.
Simon Willison released LLM / Shared entity: LLM / Earlier coverage
Linked by a graph relationship (Simon Willison released LLM); both cover LLM; earlier LLM coverage from 2026-08-27.
Linked by a graph relationship (Simon Willison released LLM); both cover LLM; earlier LLM coverage from 2026-08-21.
Linked by a graph relationship (Simon Willison released LLM); both cover LLM; earlier LLM coverage from 2026-08-17.
Linked by a graph relationship (Simon Willison released LLM); both cover LLM; earlier LLM coverage from 2026-08-12.