Fetching from the wire…
OSS2026-09-15 · source-backed
Seongeun So maintains a free online textbook, last updated September 15, covering symbolism versus connectionism, RNNs, Transformers, MoE, scaling laws, RLHF, Flash Attention, quantization, speculative decoding, retrieval, multimodal learning, state space models and mechanistic interpretability, with PyTorch examples, quizzes and interactive visualizers (sungeuns.github.io). Explicitly a living document.
Each link below shares sources, entities, or timing with this story.
Hy4-Preview runs 256 routed experts plus one always-active shared expert with top-8 routing, combining Multi-head Latent Attention, DeepSeek Sparse Attention with shared indexer layers, gated MLA with learnable attention sinks, and Independent Hyper-Connections replacing the p...
Gemma 4 12B dropped June 3, and the spec sheet is the kind of thing I read twice to make sure I wasn't misreading it. 11.95 billion params, Apache 2.0, reads text, image, audio, and video. No separate vision encoder. No separate audio encoder. The model handles all of it nativ...
A 15-author Huawei team argues kernel-generation benchmarks are almost entirely CUDA and Triton, leaving less-documented hardware with no shared yardstick (arXiv 2607.20518). CANN Bench covers 53 operators and 1,060 test cases in four difficulty tiers, from elementwise primiti...
google-research/timesfm gained 326 stars, second on the all-language board. google/timesfm-3.0-pytorch was created August 24 and shows 257 likes against 0 downloads (Hugging Face). Nine days of likes with no downloads means people are bookmarking, not running. Useful calibrati...
The Hugging Face repo contains a tokenizer, prompt-encoding reference and a minimal PyTorch implementation covering the vision encoder and aligner, DFlash attention, MoE, Hyper-Connections and the DSpark forward path. The model is a 284B-parameter MoE of twenty 13B experts, AP...
Strix Halo and Strix Point default to Vulkan instead of ROCm for up to 23% faster prompt processing and 8% faster generation, and AMD iGPUs without ROCm move to Vulkan instead of CPU on Linux. On Apple Silicon, gated-delta models train up to 25% faster and quantized MLX KV cac...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.