Agents
CentaurBench: the best automation model loses on augmentation in five of seven tasks
CentaurBench evaluates models on assisting a standardized lower-capacity worker model rather than producing the deliverable directly, across seven economically grounded work tasks scored by a blind LLM judge panel over ten replications. Rankings under automation and augmentation are only modestly correlated, and the automation winner is beaten on augmentation in five of seven tasks. Worse for anyone wiring a strong model as a planner over a cheap executor, the unaided worker outranked every assisted condition on three tasks and only one model's guidance beat no guidance on average.
Source
↳ Follow the thread