Research
τ^τ-bench Makes Agent Construction the Task: Claude Opus 5 Under Claude Code Passes 23.9% Against an 82.2% Expert Ceiling
τ^τ-bench (arXiv 2609.04611) gives a developer agent a real client engagement setup — business records, a requirements-holding client, a production API, an inherited codebase, and serving cost/model limits — and scores it by deploying the customer-service agent it builds against held-out simulated users. Across 53 tasks in four domains, the strongest configuration (Claude Opus 5 under Claude Code) passes 23.9% of evaluation simulations versus an expert-authored reference ceiling of 82.2%. The failure modes are the human ones: shallow queries instead of deep record comprehension, almost no communication back to the client, and too little experimentation with agent architecture and serving setup.
↳ Follow the thread