Fetching from the wire…
Infra2026-09-19 · source-backed
arXiv 2609.20734 shows a pretrained model's decoding states already carry information predictive of whether a global read will help, before the read happens. Training only a small recall head to invoke global attention selectively, with pretrained weights untouched and the full historical KV cache available for later recall, plus GPU-side conditional execution implemented in vLLM, turns reduced global reads into real decoding speedups at long context. Across Qwen and Gemma including hybrid-attention backbones, selective recall recovered most of the accuracy lost under pure local attention. arXiv
Each link below shares sources, entities, or timing with this story.
Amid a week of pricing and commerce stories, here's hard tech you can actually download. Google released DiffusionGemma on June 10, a 26B-parameter Mixture-of-Experts model (3.8B active) that generates text by diffusion instead of left-to-right decoding. The architecture is th...
Google DeepMind released Gemma 4 on April 2 with four model sizes (E2B, E4B, 26B MoE, 31B Dense) under Apache 2.0. Multimodal (text, vision, audio). 256K context. Native thinking and tool-calling optimized for agentic workflows. Day-zero ecosystem support across vLLM, llama.cp...
Google's HF org lists diffusiongemma-26B-A4B-it (~4B active), an image-text-to-text Gemma member that's diffusion-style rather than purely autoregressive (Hugging Face). No detailed announcement yet, which is why I'm flagging it low. But a diffusion approach inside the Gemma o...
Google released Gemma 4 12B June 3 under clean Apache 2.0, native multimodality, up to 256K context on larger variants, with the 31B reportedly at 85.2% MMLU Pro. Two days later came QAT versions optimized for mobile and laptop hardware. The 12B size targets the single-GPU swe...
Gemma 4 12B dropped June 3, and the spec sheet is the kind of thing I read twice to make sure I wasn't misreading it. 11.95 billion params, Apache 2.0, reads text, image, audio, and video. No separate vision encoder. No separate audio encoder. The model handles all of it nativ...
Google DeepMind shipped Quantization-Aware Training checkpoints for every Gemma 4 size, and the headline number is genuinely useful: the smallest model goes from 11.4GB to 1.1GB. That's 0.84GB if you go text-only. Up to ~72% lower VRAM and 2x faster inference on mobile NPUs, w...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.