Fetching from the wire…
Models2026-07-29 · source-backed
SKT put it on Hugging Face July 29, scaling from the 519B K1 on 8.5 trillion tokens with average benchmark performance up 32.2 points over K1 despite fewer training tokens. Scores: 97.1 AIME26, 80.5 KMMLU-Pro, 91.6 CLIcK, 98 on tau²-Bench telecom, all eight problems of the 2026 Korean Math Olympiad round two, 35/42 on IMO 2025, retrieval accuracy holding at roughly 256K. The transferable part is efficiency: proprietary Sparse Gated Attention referencing only needed spans, Gated Norm for stability, and FP8 from the outset rather than post-hoc quantization, which SKT says roughly halves storage and inference cost.
Each link below shares sources, entities, or timing with this story.
43.3% on Frontier-Bench v0.1. Opus 4.8 scored 18.7%. That's not an incremental bump, that's the same benchmark with a different shape of answer. Anthropic released Claude Opus 5 on July 24 at $5/$25 per million input/output tokens, exactly half of Fable 5's $10/$50, while matc...
The alignment team published an aggregate benchmark for argumentation on questions that can't be empirically settled: 60% LMCA (560 position texts, 1,461 expert-rated arguments), 20% ACCoRD (~14,000 constraints checking probability consistency), 20% DTBench (407 expert-written...
The dots.studio lab announced the preview on August 15 with multimodal text, vision, and audio understanding, introducing a TEMPO reinforcement learning method for long-horizon agent training (PANews). A SemiAnalysis chart puts it 4.9 points above the best US open-weight model...
Moonshot AI dropped Kimi K2.6 today and the numbers are hard to ignore. One trillion parameters total, 32 billion active per token across 384 experts, 256K context window, and native multimodal input. It scores 58.6 on SWE-Bench Pro versus GPT-5.4's 57.7 and Claude Opus 4.6's...
Rented GPUs on Modal, trained on exactly 5,000,003,584 decontaminated tokens, keeping K3's actual architecture including Kimi Delta Attention, Gated MLA, attention residuals, LatentMoE with the aux-loss-free balancer, and K3's unmodified 163,840-token tokenizer. 145M active pe...
Finally, a number. Every conversation about "AI can do large-scale migrations now" has been vibes and demo videos. Anthropic's engineering post on AI code migration puts a receipt on the table, and the receipt is detailed enough to model against. Bun's Zig-to-Rust port: over a...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.