Fetching from the wire…
Agents2026-08-15 · source-backed
The Wiggle Framework stress-tested 9 frontier models across 14 judging tasks. The damning part: flips were almost always net-corrupting relative to ground truth. Pressure moved judges away from the right answer, not toward it. arXiv 2608.12645 If you use LLM-as-judge anywhere in an eval harness or a moderation path, accuracy alone is not a qualifying metric. Add a persuasion-resistance test.
Each link below shares sources, entities, or timing with this story.
Simon Willison released LLM / Shared entity: LLM / Earlier coverage / Tension
Linked by a graph relationship (Simon Willison released LLM); both cover LLM; earlier LLM coverage from 2026-07-27.
Linked by a graph relationship (Simon Willison released LLM); both cover LLM; earlier LLM coverage from 2026-06-19.
Linked by a graph relationship (Simon Willison released LLM); both cover LLM; earlier LLM coverage from 2026-06-18.
Simon Willison released LLM / Shared entity: LLM / Earlier coverage
Linked by a graph relationship (Simon Willison released LLM); both cover LLM; earlier LLM coverage from 2026-07-31.
Linked by a graph relationship (Simon Willison released LLM); both cover LLM; earlier LLM coverage from 2026-06-19.
Linked by a graph relationship (Simon Willison released LLM); both cover LLM; earlier LLM coverage from 2026-08-12.
Linked by a graph relationship (Simon Willison released LLM); both cover LLM; earlier LLM coverage from 2026-08-07.
Linked by a graph relationship (Simon Willison released LLM); both cover LLM; earlier LLM coverage from 2026-08-05.