Self-Consistency Is Epistemically Shallow: 100 Samples From One Model Surface One Dimension of Signal, 24 Models Surface Four
Izhar Ali runs a clean comparison on identical questions — one model sampled 100 times at τ=1 versus an ensemble of 24 LLMs run once each at τ=0 — and applies a Marchenko-Pastur random-matrix test to separate signal from sampling noise on both sides. Within any single model, at most one dimension rises above the noise edge across five model families and three benchmarks (MMLU, HellaSwag, GSM8K); across the ensemble, four eigenvalues clear it, against a matched-difficulty Bernoulli null that produces at most one in 500 Monte Carlo draws. Temperature sampling gives you accurate per-question uncertainty and nothing else — if you want to know what a model doesn't know as a structure, you need different models, not more samples.
↳ Follow the thread