Only About Half an LLM's Accuracy Advantage Reaches the Human Consulting It
Available but Unclaimed (arXiv 2609.16793, submitted 15 Sep 2026) ran a between-subjects study where 535 participants solved a 40-item battery of matrix reasoning, mental rotation, syllogisms, and letter-string analogies either unaided or while required to consult GPT-5.6-Luna, Claude Opus 4.8, Gemini 3.6 Flash, or Kimi K3, with each model also answering every item alone 100 times under matched elicitation. In a reference comparison, roughly half the increase in LLM accuracy carried through to assisted accuracy, and how much reached participants differed by model. Post-advice confidence separated correct from incorrect answers less well than unaided confidence, meaning consulting a model degraded people's ability to tell when they were wrong.
↳ Follow the thread