Fetching from the wire…
Public story · 2026-07-20 · high
Participants working alone scored 27 percent correct, while the AI-assisted group scored worse and felt far more certain.
Why now: MIT Tech Review published its hiring-bias findings on July 20, while the study's HN thread was still climbing past 349 points.
AI assistance dropped accuracy from 27 percent to 9 percent while confidence jumped to 76 percent, a new study found. Researchers from University of Milano-Bicocca, École Normale Supérieure, and Sapienza tested people answering questions with and without AI help. The unaided group got 27 percent right. The assisted group got 9 percent, worse by a factor of three, and felt far more certain of it.
The number that matters more is what happened to "I don't know." In the baseline group, 44 percent of participants were willing to say they didn't know an answer. With AI help, that fell to 3 percent. The authors call it cognitive surrender: people adopted the model's output with little scrutiny, and it overrode both intuition and reasoning.
I use Claude Code every day in my personal projects, and I'd like to think I'm not that participant. I have context the model doesn't, and I notice when an answer feels off. But the study's mechanism works by making you feel more certain, not less. That means I can't just ask myself if I'm skeptical enough. Self-assessment is the wrong instrument here.
Two other pieces make the same point from different angles. Blaine Hansen argues LLMs aren't like compilers or power tools. Compilers and power tools are deterministic and verifiable. Models aren't, and treating them the same is the mistake. MIT Tech Review reported that AI résumé screeners are more likely than human screeners to develop hiring bias during evaluation. They don't just inherit it from training data. Different domains, same shape. The model isn't a neutral amplifier of judgment.
The fix isn't a review step that asks whether something looks right, since that mostly produces agreement. It has to be a separate pass, run by someone who hasn't seen your reasoning, whose only job is to find what's wrong. Bring back "I don't know," too. If your process never produces it, that's not confidence. That's a missing capability.
Each link below shares sources, entities, or timing with this story.
Anthropic released Claude Max / Shared entities / Same source domain / Shared topic / Earlier coverage
Linked by a graph relationship (Anthropic released Claude Max); both cover Anthropic, Claude Code, LLM; reported by the same outlet (news.ycombinator.com).
Anthropic released Claude Code / Shared entities / Same source domain / Shared topic / Earlier coverage
Linked by a graph relationship (Anthropic released Claude Code); both cover Claude Code, LLM, LLMs; reported by the same outlet (news.ycombinator.com).
Anthropic released MCP / Shared entities / Shared topic / Earlier coverage / Tension
Linked by a graph relationship (Anthropic released MCP); both cover Anthropic, LLM, LLMs; overlapping topics (agent, context).
DeepSeek competes with Anthropic / Shared entities / Shared topic / Earlier coverage / Tension
Linked by a graph relationship (DeepSeek competes with Anthropic); both cover Claude Code, Different; overlapping topics (agent, context, model, same).
Anthropic released Claude / Shared entities / Earlier coverage
Linked by a graph relationship (Anthropic released Claude); both cover Anthropic, Claude Code, LLM, LLMs; earlier Anthropic coverage from 2026-05-02.
Anthropic released MCP / Shared entities / Shared topic / Earlier coverage
Linked by a graph relationship (Anthropic released MCP); both cover Claude Code, LLM, LLMs; overlapping topics (agent, context).
Anthropic released Claude Code / Shared entities / Shared topic / Earlier coverage
Linked by a graph relationship (Anthropic released Claude Code); both cover Claude Code, LLM, LLMs; overlapping topics (agent, context).
Linked by a graph relationship (Anthropic released Claude Code); both cover Anthropic, Claude Code, LLM; overlapping topics (agent, context).