Persona Panels Beat Single LLM Judges for UI Evaluation, Raising Correlation With Humans From r=0.716 to r=0.922
ESPP (arXiv 2607.28439, July 30) evaluates generative UI by having a panel of psychologically diverse, evidence-grounded personas independently rate a screenshot, exchange opinions under a trait-derived semantically-gated bounded-confidence mechanism, then aggregate via Delphi-inspired social weighting. Correlation with human judgment rises from Pearson r=0.716 for a naive single-pass judge to 0.922, and a prompt-ensemble control recovers only about a third of that gap — isolating persona diversity and evidence grounding, not sampling variance, as the driver. Keeping individual panelist ratings also shows subgroups agree on overall model rankings but diverge sharply on specific dimensions, structure a homogeneous judge erases; code is released.
Source
↳ Follow the thread