Fetching from the wire…
Research2026-09-01 · source-backed
Translating key-value state from one model into a form another can consume works across scale, architecture, attention configuration, tokenizer and family (arXiv 2608.30963). Llama3.1-70B to Qwen2.5-7B reaches 44.0% accuracy against 45.7% native while dropping latency to 138ms; Qwen2.5-1.5B to Gemma-2-2B cuts target-side prefill up to 67.05% at 4K context. Within-family, Qwen2.5-7B to Qwen2.5-1.5B raises LongBench2 accuracy from 27.59% to 34.48%, so the small model inherits capability from the large one's prefill. I want to see this reproduced before I believe the cross-family numbers.
Each link below shares sources, entities, or timing with this story.
Google DeepMind released Gemma 4 on April 2 with four model sizes (E2B, E4B, 26B MoE, 31B Dense) under Apache 2.0. Multimodal (text, vision, audio). 256K context. Native thinking and tool-calling optimized for agentic workflows. Day-zero ecosystem support across vLLM, llama.cp...
Hugging Face published its Summer 2026 State of Open Models report on August 14, and one statistic in it went almost entirely unremarked in the coverage. By July 2026, agents rather than humans became the Hub's primary users. Claude Code alone accounted for 44.4% of all agent...
Google's HF org lists diffusiongemma-26B-A4B-it (~4B active), an image-text-to-text Gemma member that's diffusion-style rather than purely autoregressive (Hugging Face). No detailed announcement yet, which is why I'm flagging it low. But a diffusion approach inside the Gemma o...
Huang used his inaugural X post on July 24 to publish "Open Weights and American AI Leadership," a three-page letter on Nvidia's own servers signed by 25 companies including Meta, Microsoft, IBM, Mistral, Mozilla, Hugging Face, a16z, Palantir and the Linux Foundation. Within a...
Amid a week of pricing and commerce stories, here's hard tech you can actually download. Google released DiffusionGemma on June 10, a 26B-parameter Mixture-of-Experts model (3.8B active) that generates text by diffusion instead of left-to-right decoding. The architecture is th...
Google DeepMind shipped Quantization-Aware Training checkpoints for every Gemma 4 size, and the headline number is genuinely useful: the smallest model goes from 11.4GB to 1.1GB. That's 0.84GB if you go text-only. Up to ~72% lower VRAM and 2x faster inference on mobile NPUs, w...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.