Fetching from the wire…
Models2026-09-01 · source-backed
Adapted from ICLR 2026 and MLSS 2026 workshops, covering masked diffusion, block diffusion for variable-length generation, encoder-decoder architectures, remasking-based error correction, sampling distillation, guidance and RL post-training (Kuleshov Group). It catalogs what you can actually build on: LLaDA at 8B with open weights and LLaMA compatibility, Google's Gemma Diffusion, NVIDIA's Nemotron Diffusion family up to 35B, ESM3 at 100B for proteins, Nucleotide Transformer v3 on about a trillion DNA tokens. The commercial datapoint is Mercury 2 at 1,000+ tokens/second on standard GPUs, which the authors put at 5-10x comparable-quality models.
Each link below shares sources, entities, or timing with this story.
Amid a week of pricing and commerce stories, here's hard tech you can actually download. Google released DiffusionGemma on June 10, a 26B-parameter Mixture-of-Experts model (3.8B active) that generates text by diffusion instead of left-to-right decoding. The architecture is th...
PR #19378 landed in llama.cpp this week, and I think most people are underselling what it means. Backend-agnostic tensor parallelism via --split-mode tensor makes multi-GPU inference work across AMD, Intel, and Apple Silicon. Not just CUDA. Everything. For context: llama.cpp h...
It's a vertical-integration play that mirrors Google's TPU and Meta's Iris silicon, aimed at cutting Nvidia dependence while positioning for public markets (BuildFastWithAI). If it's real, it says the model makers now think the chip is part of the moat, not a commodity you ren...
Xiaomi released MiMo-V2.5-Pro, a 1.02 trillion parameter mixture-of-experts model (42B active) with 1M token context, fully MIT licensed. In benchmarks, it achieves 63.8% success on agentic tasks using 40-60% fewer tokens than Claude Opus 4.6 or GPT-5.4 for comparable results....
Harvard and Google released the first TPU-native benchmark for AI-generated kernel optimization: 50 JAX workloads, 17 production operators from MaxText architectures (Llama-3.1, DeepSeek-V3, Mixtral, Mamba-2, AlphaFold2) and 33 translated from KernelBench at sizes tuned for hi...
TechCrunch's August 29 piece frames Nvidia's durable advantage as system-level, built around Vera Rubin pairing the Rubin GPU with the Vera CPU, a Groq 3 LPX inference accelerator, and storage and networking racks. VP of storage technology Jason Hardy is quoted claiming "upwar...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.