Fetching from the wire…
Research2026-08-06 · source-backed
arXiv 2608.04570 tests 150 personas across 6 tasks. Every one of 12 models over-inferred user attributes on 35-49% of claims, ranging 27-59% by task type. The damning result is the self-monitoring inversion: models rating themselves as over-inferring least ranked as fabricating most (rho = -0.60, p = 0.044). Inferred attributes accumulate roughly linearly across multi-turn interactions with almost no revision. Any memory layer trusting the model's own confidence is compounding fabricated profile data.
Each link below shares sources, entities, or timing with this story.
Same source domain / Shared topic
Reported by the same outlet (arxiv.org); overlapping topics (almost, data, model, task).
Same source domain / Shared topic / Tension
Reported by the same outlet (arxiv.org); overlapping topics (claim, model); pushes against this story (vs).
Reported by the same outlet (arxiv.org); overlapping topics (data, model); pushes against this story (but).
Same source domain / Shared topic / Downstream implication
Reported by the same outlet (arxiv.org); overlapping topics (almost, model); traces where this leads (downstream).
Same source domain / Shared topic / Tension
Reported by the same outlet (arxiv.org); overlapping topics (data, model); pushes against this story (against).
Shared topic / Tension
Overlapping topics (almost, claim, data, model); pushes against this story (but).
Same source domain / Shared topic / Tension
Reported by the same outlet (arxiv.org); overlapping topics (model, task); pushes against this story (against).
Reported by the same outlet (arxiv.org); overlapping topics (data, task); pushes against this story (versus).