Fetching from the wire…
Research2026-07-26 · source-backed
Izhar Ali compares one model sampled 100 times at τ=1 against an ensemble of 24 LLMs run once each at τ=0 on identical questions, applying a Marchenko-Pastur random-matrix test to separate signal from sampling noise on both sides (arXiv 2607.20464). Within any single model, at most one dimension rises above the noise edge, across five model families and three benchmarks. Across the ensemble, four eigenvalues clear it, against a matched-difficulty Bernoulli null producing at most one in 500 Monte Carlo draws. Temperature sampling gives accurate per-question uncertainty and nothing else. If you want structured knowledge of what a model doesn't know, you need different models, not more samples. This kills a lot of self-consistency ensemble designs.
Each link below shares sources, entities, or timing with this story.
Shared entities / Same source domain / Earlier coverage
Both cover LLMs, Self; reported by the same outlet (arxiv.org); earlier LLMs coverage from 2026-07-13.
Shared entity: LLMs / Same source domain / Shared topic / Earlier coverage
Both cover LLMs; reported by the same outlet (arxiv.org); overlapping topics (benchmark, model).
Both cover LLMs; reported by the same outlet (arxiv.org); overlapping topics (benchmark, model).
Shared entity: Self / Same source domain / Earlier coverage / Downstream implication
Both cover Self; reported by the same outlet (arxiv.org); earlier Self coverage from 2026-07-22.
Shared entity: LLMs / Same source domain / Earlier coverage / Downstream implication
Both cover LLMs; reported by the same outlet (arxiv.org); earlier LLMs coverage from 2026-06-19.
Both cover LLMs; reported by the same outlet (arxiv.org); earlier LLMs coverage from 2026-06-19.
Both cover LLMs; reported by the same outlet (arxiv.org); earlier LLMs coverage from 2026-06-14.
Shared entity: LLMs / Same source domain / Earlier coverage / Tension
Both cover LLMs; reported by the same outlet (arxiv.org); earlier LLMs coverage from 2026-06-09.