Sources
Kimi Delta Attention Claims Up to 6x Cheaper Throughput at 1M Context — the Architecture Detail Buried Under K3's Benchmark Headlines
Latent Space's AINews issue (titled, deadpan, 'not much happened today') surfaces the K3 detail the mainstream coverage skipped: Kimi Delta Attention (KDA) delivers up to 6x faster and cheaper throughput at 1M-token context. It also notes K3 placing #3 on DeepSWE — described as the first open-weights model at frontier level on that benchmark — plus 84% on Terminal-Bench v2 and 64% on DeepSWE per Artificial Analysis. For anyone budgeting long-context agent runs, the throughput multiplier matters more than the leaderboard position.
↳ Follow the thread