Fetching from the wire…
Infra2026-09-17 · source-backed
Edge0 attacks why naive SSD offloading fails for MoE: layer N+1's experts must be chosen before layer N's output exists, so reads can't start early enough to hide behind compute. A per-layer prerouter predicts the next layer's routing one token ahead and then uses that prediction as the routing, so the staged expert set equals the routed set and nothing is dropped, with an unmerged recovery LoRA paying back int4 quantization loss. On a single 24GB machine: 20 tok/s inside 3GiB peak active memory, within a few points of its fp16 teacher across five benchmarks. Framework, checkpoints and adapters open sourced.
Each link below shares sources, entities, or timing with this story.
PortLLM claimed training-free, data-free transfer of LoRA patches onto updated base models, but only over short horizons and without theoretical grounding. This study runs 10 continual-pretraining steps on Mistral, Gemma, and Qwen and finds portability persists long-run, meani...
Xiaomi released MiMo-V2.5-Pro, a 1.02 trillion parameter mixture-of-experts model (42B active) with 1M token context, fully MIT licensed. In benchmarks, it achieves 63.8% success on agentic tasks using 40-60% fewer tokens than Claude Opus 4.6 or GPT-5.4 for comparable results....
The September 11 fine-tuning expansion adds 18 open-weight models (GLM 5.3/5.2/5.1, DeepSeek-V4-Flash variants, Kimi K2.7-Code, Qwen 3.8-27B, Gemma 4) plus Expert LoRA, which puts adapters on the experts themselves. On invented-fact recall, expert-inclusive adapters reached 89...
Edge0-AI/Edge0, created September 8, went from 269 to 583 stars in two days. It packages SSD expert offload, Recover-LoRA and prerouter routing prediction into an MLX-backed framework: edge0-35b is a 4-bit 40-layer 256-expert model built on Qwen3.5-MoE 35B-A3B needing about 2....
Google DeepMind released Gemma 4 on April 2 with four model sizes (E2B, E4B, 26B MoE, 31B Dense) under Apache 2.0. Multimodal (text, vision, audio). 256K context. Native thinking and tool-calling optimized for agentic workflows. Day-zero ecosystem support across vLLM, llama.cp...
The models are good. The license is the real story. Google released Gemma 4 on April 2 with four variants: E2B, E4B, 26B MoE, and 31B Dense. All built on the Gemini 3 architecture. The 31B Dense variant claimed #3 on Arena AI's text leaderboard, beating models 20x its size. Th...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.