Agents
CentaurBench finds the model that wins at automation loses at augmentation on five of seven tasks
A 46-page benchmark paper (arXiv:2608.18554, submitted 19 Aug 2026) evaluates LLMs on assisting a weaker worker model rather than doing the task alone, across seven real-world tasks scored by blind pairwise LLM judging over ten runs. Rankings between the two regimes are only modestly correlated, the automation winner loses augmentation on five of seven tasks, and on three tasks the unaided worker beats every assisted condition. Only one model's guidance beat no guidance on average, which is a direct warning against picking orchestrator models by solo leaderboard score.
Source
↳ Follow the thread