Fetching from the wire…
Public story · 2026-09-20 · high
A replay of human debate groups found AI stand-ins converging 34 to 44 points more often than the people they were built to model.
Why now: The paper measuring this gap posted to arXiv on September 20.
AI agent panels reached full agreement far more often than the human groups they were built to stand in for. Researchers replayed 100 held-out human debate groups through matched panels of AI agents, seeding each agent with its matching human's pre-discussion answer, then scored both sets with identical code.
The gap matters for anyone using a multi-agent panel as a stand-in for group judgment, including researchers who'd otherwise need to recruit human subjects. If the panel's agreement rate doesn't track how real groups actually behave, it's not a valid substitute.
The task was a set of Wason logic problems, described in the paper replaying human debate groups. Full human consensus ranged from 24.0% to 57.0%, depending on how participation got counted. About a fifth of human group members never posted an answer at all, while their AI counterparts almost always did.
That participation gap got checked twice, not once. Counting only people who submitted an answer, agents beat humans by 34.0 to 43.9 points. Matching participation more strictly between the two groups, the gap held at 34.1 to 44.4 points.
The agents aren't malfunctioning. Copies of a similar reasoning process converge on similar answers, which is exactly what happened here. The paper doesn't test whether the gap holds on tasks less structured than Wason problems, which is the open question for anyone tempted to swap a panel in for a focus group.
Each link below shares sources, entities, or timing with this story.
After 20+ years maintaining Paint.NET, Rick Brewster concluded WINE's Direct2D would never be complete enough for what he needed, so the app now carries its own from-scratch reverse-engineered Direct2D implementation. He puts it at 180,000 lines against 700,000 for the rest of...
Allen Bargi's August 15 post hit 302 points arguing that AI collaboration rewards context-sharing, examples, and feedback over precise instruction (Hacker News). The pushback holds that the piece conflates management with leadership. mikeocool calls it "the most low effort ver...
Riffing on Apple's DRI management concept, he argues accountability requires an entity that can actually be held responsible, and a machine cannot (Simon Willison). It's a sharp, quotable counterweight to the "let the agent own it end-to-end" enthusiasm. I keep this one close...
I've been telling people for months that the agent code I review is *correct and awful*. Correct in the sense that it compiles, passes the tests, does the thing. Awful in the sense that a 400-line function does the work of 80, the same helper exists three times under different...
Simon Willison spent a while taking ChatGPT Work apart and published the map on August 30. Work splits into Work Cloud and Work Local, the latter being the renamed Codex desktop app, at $20/month and up since July 9. He enumerates six capabilities Work has that Chat doesn't, a...
CCP announced the migration covering code that has run on Stackless 2.7 since 2010. The approach is to run futurize across the codebase and then manually review roughly 20,000 places where Python 2 and 3 behavior diverges, including integer division (Simon Willison). No comple...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.