Agents
MCP Atlas leaderboard: Muse Spark 1.1 leads tool-calling at 88.1%, Claude Opus 5 at 85.8%
BenchLM's MCP Atlas page, last updated August 21, 2026, ranks 36 closed and open-weight models on interactive tool-calling over Model Context Protocol integrations at an advanced difficulty tier. Meta's Muse Spark 1.1 leads at 88.1%, followed by Claude Opus 5 at 85.8% and Moonshot's Kimi K3 at 84.2%. The benchmark refreshes quarterly and sits in BenchLM's agentic category at 22% weight, though it is currently display-only and excluded from composite rankings, so treat the spread as directional rather than decisive.
Source
↳ Follow the thread