Skills
Stop benchmarking tool-use on toys: MCP-Atlas and Tool-Decathlon test agents against real MCP servers
Two new 2026 benchmarks evaluate tool-using agents on realistic workloads: MCP-Atlas has 1,000 human-verified tasks across 36 real MCP servers and 220 tools (current leaders: Gemini 3.5 Flash 83.6%, Claude Opus 4.8 82.2%), and Scale moved it to a 100-tool-call budget instead of a 20-turn limit. Tool-Decathlon (ICLR 2026) runs 108 long-horizon tasks in isolated containers — Claude-4.5-Sonnet finishes all 108 in ~70 minutes across 10 parallel processes. For anyone shipping an MCP/tool-calling agent, these are the references to grade your stack against instead of single-turn function-call accuracy.
Source
↳ Follow the thread