Fetching from the wire…
Infra2026-09-17 · source-backed
HOPE points out existing expert-pruning methods decide each expert in isolation and assume contributions are purely additive, when MoE expert usage is cooperative. It derives a second-order objective that provably minimizes an upper bound on pruning error, and shows the state-of-the-art first-order method REAP is the special case where interaction terms are dropped. Across three frontier MoE models up to 122B, two calibration sets, and math, instruction-following, coding and agentic benchmarks, HOPE averages rank 1.58 of 5 methods at 50% pruning against REAP's 2.42, with the widest margins at high pruning rates.
Each link below shares sources, entities, or timing with this story.
The architecture report describes a 125B sparse MoE activating 6B parameters per token, plus 51B of n-gram embedding tables held off the accelerator in host memory with prefetching (arXiv 2608.30320). Against the prior 397B-A17B model it leads on 8 of 14 pre-training benchmark...
A 2.4-trillion-parameter MoE multimodal model with 1M-token context, aimed at long-horizon autonomous software work. Alibaba reports 86.1 against 83.2 for GPT-5.6 Sol Max and 85.0 for Fable 5, the first credible claim that a Chinese lab leads on GUI-driving agentic benchmarks....
Unisound's U2 is a 266B-total / 10B-active MoE tuned for agents, citing 72.2% SWE-bench Verified at $0.15/$0.30 per 1M tokens, while GLM-5.2 is getting named the strongest open-weight coding model across July roundups. Both are roundup-sourced, so verify the benchmarks against...
The assumption that proprietary models own the coding benchmark crown just broke. Moonshot AI's Kimi K2.6 leads on 5 of 8 major agentic coding benchmarks while being the only open-weight model in the top tier. SWE-Bench Pro: 58.6% vs GPT-5.4's 57.7% and Claude Opus 4.6's 53.4%...
Moonshot AI dropped Kimi K2.6 today and the numbers are hard to ignore. One trillion parameters total, 32 billion active per token across 384 experts, 256K context window, and native multimodal input. It scores 58.6 on SWE-Bench Pro versus GPT-5.4's 57.7 and Claude Opus 4.6's...
The approach recasts expert-activation prediction as sequence-to-sequence modeling for multi-step multi-layer forecasts, then treats prefetching as job sequencing with deadlines and adds probabilistic Belady eviction. At 45% residency it averages a 96.97% hit rate. The offload...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.