Kimi K3's 2.8T Open Weights Still Lack Same-Harness Reproduction on Coding Benchmarks
r/LocalLLaMA·medium signal
Moonshot's Kimi K3 (2.8T parameters, 1M context, weights dropped 2026-07-27) reports 88.3 on Terminal-Bench 2.1, 81.2 on FrontierSWE, 77.8 raw pass on ProgramBench, 67.5 DeepSWE, and 42.0 SWE Marathon. The caveat that matters for tool selection: published results mix harnesses — Kimi Code, Claude Code, Codex, and mini-SWE-agent — so no independently reproduced same-harness comparison exists yet, and the model is reportedly sensitive to preserved thinking history. Meanwhile r/LocalLLaMA has it running at 20+ tok/s average (38 peak, 750 prefill) on a 16x GB10 cluster, so self-hosting is real but not cheap.