Tools
Kimi K3's Published Benchmarks Put an Open-Weight Model Ahead of Closed Frontier Models on SWE-Marathon, BrowseComp and MCPMark
The K3 repo's evaluation tables report head-to-head scores against Claude Fable 5, GPT-5.6 Sol, Claude Opus 4.8, GPT-5.5 and GLM-5.2. K3 leads on SWE-Marathon (42.0 vs 35.0 for Fable 5 and 39.0 for GPT-5.6 Sol), BrowseComp (91.2 vs 88.0 and 90.4), MCPMark-Verified (94.5 vs 87.4 and 92.9), AutomationBench (30.8) and SpreadsheetBench 2 (34.8), while trailing on HLE-Full (43.5/56.0 vs 53.3/63.0), DeepSWE (67.5 vs 73.0) and GDPval-AA v2 Elo (1686 vs 1747). The pattern is consistent: K3 is competitive on long-horizon agentic and tool-use work and behind on hardest-reasoning benchmarks.
Source
↳ Follow the thread