Fetching from the wire…
Infra2026-09-14 · source-backed
The approach recasts expert-activation prediction as sequence-to-sequence modeling for multi-step multi-layer forecasts, then treats prefetching as job sequencing with deadlines and adds probabilistic Belady eviction. At 45% residency it averages a 96.97% hit rate. The offloading runtime keeps end-to-end graph capture working, which is the part most offloading schemes break.
Each link below shares sources, entities, or timing with this story.
arXiv 2607.24667 recasts eviction as estimation of a hidden reuse signal along a commit-lag axis, with StreamingLLM/H2O/SnapKV at lag 0 and Belady's optimum at full future knowledge, then fills the middle: wait a bounded number of steps, observe what a correct near-future pred...
Power availability is now a primary limit on AI infrastructure growth, but making training power-flexible requires knowing how throughput responds to reduction, which nobody had characterized. The index is a normalized metric for the performance cost of a power cut that double...
Most looped-transformer results compare at fixed model size, conflating architecture with extra compute. SMELT matches per-token FLOPs, non-embedding parameters and KV cache against an unlooped baseline, scaling to 54B non-embedding parameters with a separate Chinchilla-style...
The architecture report describes a 125B sparse MoE activating 6B parameters per token, plus 51B of n-gram embedding tables held off the accelerator in host memory with prefetching (arXiv 2608.30320). Against the prior 397B-A17B model it leads on 8 of 14 pre-training benchmark...
Load Hijack modifies nothing but router weights in a checkpoint. When a private trigger appears, token-to-expert assignment concentrates on experts co-located on a single GPU, making it a straggler while peers idle (arXiv 2608.10614). Across three MoE families and four corpora...
arXiv 2607.14530 gets Hyper-Connections past the N=4 wall by sparsely updating only k=4 streams plus temporal feature augmentation, scoring 4.0 points higher on average downstream than prior mHC on an 18B MoE. Vanilla and mHC need 1.50x and 1.19x xHC's compute to hit the same...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.