Fetching from the wire…
Public story · 2026-08-05 · high
Its published coding scores were run across four different harnesses, so none of them can be compared to each other yet.
Why now: The cluster run surfaced in r/LocalLLaMA discussion covered in the August 5 briefing, while K3's benchmark table draws scrutiny for mixing harnesses.
A LocalLLaMA poster ran Moonshot's full Kimi K3 on 16 Nvidia GB10 chips and clocked 20-plus tokens a second, per a thread with 1,330 upvotes.
That's roughly $64,000 in hardware producing frontier-adjacent coding output at your desk. Dspark speculative decoding pushed bursts to 38 tokens a second, with prefill hitting 750. For anyone weighing a local rig against a Claude Code or Codex subscription, that's a real data point instead of a marketing slide.
Moonshot's scores land in frontier-adjacent territory: 88.3 on Terminal-Bench 2.1, 81.2 on FrontierSWE, 77.8 on ProgramBench raw pass, 67.5 on DeepSWE, 42.0 on SWE Marathon.
But those five numbers weren't run in one harness. Moonshot's own reporting mixes results across Kimi Code, Claude Code, Codex and mini-SWE-agent without saying which score came from which one. K3 is also reportedly sensitive to whether its thinking history gets preserved between turns, another variable the published numbers don't control for.
Anyone comparing K3's scores to Claude Code or Codex output should ask which harness produced them first.
Each link below shares sources, entities, or timing with this story.
Kimi K3 built by Moonshot / Shared entities / Same source domain / Shared topic / Earlier coverage / Tension
Linked by a graph relationship (Kimi K3 built by Moonshot); both cover Bench, LocalLLaMA, MoE, Moonshot; reported by the same outlet (reddit.com).
Kimi K3 built by Moonshot / Shared entities / Shared topic / Earlier coverage
Linked by a graph relationship (Kimi K3 built by Moonshot); both cover Bench, MoE, Moonshot, SWE; overlapping topics (benchmark, code, coding, kimi).
Claude Code uses Kimi K3 / Shared entities / Shared topic / Earlier coverage
Linked by a graph relationship (Claude Code uses Kimi K3); both cover Bench, Claude Code, MoE, SWE; overlapping topics (benchmark, claude, code, coding).
Claude Code uses Kimi K3 / Shared entities / Shared topic / Earlier coverage / Tension
Linked by a graph relationship (Claude Code uses Kimi K3); both cover Bench, Claude Code, Codex, Terminal; overlapping topics (benchmark, claude, code, codex, coding).
Claude Code uses Kimi K3 / Shared entities / Same source domain / Shared topic / Earlier coverage / Tension
Linked by a graph relationship (Claude Code uses Kimi K3); both cover Claude Code, LocalLLaMA, MoE, SWE; reported by the same outlet (reddit.com).
Kimi K3 competes with OpenAI / Shared entities / Shared topic / Earlier coverage
Linked by a graph relationship (Kimi K3 competes with OpenAI); both cover Bench, Claude Code, Codex, SWE; overlapping topics (claude, code, coding).
Kimi K3 benchmarked against Fable / Shared entities / Shared topic / Earlier coverage
Linked by a graph relationship (Kimi K3 benchmarked against Fable); both cover Bench, Claude Code, MoE, SWE; overlapping topics (benchmark, claude, code).
Kimi K3 built by Moonshot / Shared entities / Shared topic / Earlier coverage
Linked by a graph relationship (Kimi K3 built by Moonshot); both cover Codex, Kimi K3, MoE, Moonshot; overlapping topics (coding, kimi).