Fetching from the wire…
Research2026-08-06 · source-backed
arXiv 2608.05086, billed as the largest psychometric analysis of LLM safety evals to date, finds three interpretable factors (refusal strictness, truthfulness, contextual harm) explain most between-model variance. Psychometrically selected items recover full benchmark scores with lower error than random subsets of the same size, ~10 adaptively chosen items sufficing for several benchmarks. IRT also supports per-model audits that detect naive sandbagging and silent model swaps behind an API. Directly usable if you maintain an eval harness against a hosted endpoint.
Each link below shares sources, entities, or timing with this story.
Simon Willison released LLM / Shared entities / Earlier coverage / Tension
Linked by a graph relationship (Simon Willison released LLM); both cover Directly, LLM; earlier Directly coverage from 2026-06-18.
Simon Willison released LLM / Shared entity: LLM / Shared topic / Earlier coverage
Linked by a graph relationship (Simon Willison released LLM); both cover LLM; overlapping topics (behind, model).
Simon Willison released LLM / Same source domain / Shared topic / Tension
Linked by a graph relationship (Simon Willison released LLM); reported by the same outlet (arxiv.org); overlapping topics (against, model).
Simon Willison released LLM / Shared entity: LLM / Earlier coverage / Tension
Linked by a graph relationship (Simon Willison released LLM); both cover LLM; earlier LLM coverage from 2026-07-27.
Linked by a graph relationship (Simon Willison released LLM); both cover LLM; earlier LLM coverage from 2026-06-19.
Simon Willison released LLM / Shared entity: LLM / Earlier coverage
Linked by a graph relationship (Simon Willison released LLM); both cover LLM; earlier LLM coverage from 2026-07-31.
Linked by a graph relationship (Simon Willison released LLM); both cover LLM; earlier LLM coverage from 2026-08-05.
Linked by a graph relationship (Simon Willison released LLM); both cover LLM; earlier LLM coverage from 2026-08-03.