Fetching from the wire…
Models2026-07-29 · source-backed
The K3 repo's tables put it ahead on SWE-Marathon (42.0 vs 35.0 for Fable 5 and 39.0 for GPT-5.6 Sol), BrowseComp (91.2 vs 88.0 and 90.4), MCPMark-Verified (94.5 vs 87.4 and 92.9), AutomationBench (30.8), and SpreadsheetBench 2 (34.8). It trails on HLE-Full (43.5/56.0 vs 53.3/63.0), DeepSWE (67.5 vs 73.0), and GDPval-AA v2 Elo (1686 vs 1747). The pattern is consistent enough to route on: long-horizon tool use to K3, hardest reasoning elsewhere.
Each link below shares sources, entities, or timing with this story.
Kimi built by Moonshot AI / Shared entities / Shared topic / Earlier coverage
Linked by a graph relationship (Kimi built by Moonshot AI); both cover Fable, Kimi, MCPMark, Verified; overlapping topics (benchmark, frontier).
GPT competes with DeepSeek / Shared entities / Same source domain / Shared topic / Earlier coverage
Linked by a graph relationship (GPT competes with DeepSeek); both cover BrowseComp, GPT, Kimi, SWE; reported by the same outlet (github.com).
GPT competes with Claude / Shared entities / Shared topic / Earlier coverage
Linked by a graph relationship (GPT competes with Claude); both cover DeepSWE, GPT, SWE, Verified; overlapping topics (benchmark, deepswe, gpt-5).
GPT competes with Grok / Shared entities / Shared topic / Earlier coverage / Tension
Linked by a graph relationship (GPT competes with Grok); both cover DeepSWE, Fable, GDPVal, GPT; overlapping topics (benchmark, deepswe, frontier, gpt-5).
Anthropic released Fable / Shared entities / Shared topic / Earlier coverage / Tension
Linked by a graph relationship (Anthropic released Fable); both cover GPT, SWE; overlapping topics (agentic, benchmark, closed, frontier, gpt-5).
Kimi built by Moonshot AI / Shared entities / Shared topic / Earlier coverage
Linked by a graph relationship (Kimi built by Moonshot AI); both cover Fable, Full, GPT; overlapping topics (beat, fable).
Kimi built by Moonshot AI / Shared entities / Shared topic / Earlier coverage / Tension
Linked by a graph relationship (Kimi built by Moonshot AI); both cover GPT, HLE, SWE; overlapping topics (agentic, benchmark, frontier, gpt-5).
Kimi K3 benchmarked against Fable / Shared entities / Shared topic / Earlier coverage
Linked by a graph relationship (Kimi K3 benchmarked against Fable); both cover Fable, GPT, SWE; overlapping topics (benchmark, fable).