Sources
'The Personalization Mirage': 12 Models Fabricate 41.6% of User Attributes on Average, and the Ones Most Confident They Don't Are the Worst Offenders
arXiv 2608.04570 (August 5, Sun, Zhang, Sheng) introduces MirageBench — 150 personas across 6 tasks — and finds every one of 12 tested models over-inferred user attributes on 35–49% of claims (mean 41.6%), ranging 27–59% by task type. The damning result is the self-monitoring inversion: models that rated themselves as over-inferring least were ranked as fabricating most (rho = -0.60, p = 0.044). Inferred attributes also accumulate roughly linearly across multi-turn interactions with almost no revision, which means any memory or personalization layer that trusts the model's own confidence is compounding fabricated profile data — the authors argue for external verification instead of self-report.
↳ Follow the thread