Fetching from the wire…
Research2026-07-28 · source-backed
arXiv 2607.24667 recasts eviction as estimation of a hidden reuse signal along a commit-lag axis, with StreamingLLM/H2O/SnapKV at lag 0 and Belady's optimum at full future knowledge, then fills the middle: wait a bounded number of steps, observe what a correct near-future prediction actually attended to, then evict. In controlled settings it identifies used memory far better than accumulated attention. Run inside NVIDIA's KVPress harness against KVPress's own implementations, the advantage largely vanishes. The authors say plainly the contribution is the framework and "an honest map of when measuring beats accumulating, not a new state of the art." More papers should end like this. (arXiv 2607.24667)
Each link below shares sources, entities, or timing with this story.
This one annoyed me, in the good way. Researchers took 206 real developer-agent sessions from 13 developers, extracted each developer's preferences from their actual interaction traces via rule-based bootstrapping plus evidence-grounded refinement, then replayed everything aga...
Designed with Broadcom, built by TSMC, going into manufacturing after a six-week test run found no major issues (Reuters). Iris supplements rather than replaces Meta's Nvidia/AMD GPUs and supports a buildout targeting 7GW by end-2026 and 14GW in 2027 against up to $145B in 202...
Mira Murati's startup secured a multi-year deal with NVIDIA for at least 1 gigawatt of next-generation Vera Rubin chip-powered servers, plus an undisclosed equity investment. The company has raised over $2 billion since its February 2025 founding. A gigawatt of compute — rough...
Nemotron-Terminal (arXiv:2602.21193, 66 HF upvotes) — First systematic study of data engineering for terminal/CLI agents. Terminal-Task-Gen pipeline with Dockerized environment interaction. Qwen3-initialized 8B model goes from 2.5% to 13.0% on Terminal-Bench 2.0. All checkpoin...
Adapted from ICLR 2026 and MLSS 2026 workshops, covering masked diffusion, block diffusion for variable-length generation, encoder-decoder architectures, remasking-based error correction, sampling distillation, guidance and RL post-training (Kuleshov Group). It catalogs what y...
TechCrunch's August 29 piece frames Nvidia's durable advantage as system-level, built around Vera Rubin pairing the Rubin GPU with the Vera CPU, a Groq 3 LPX inference accelerator, and storage and networking racks. VP of storage technology Jason Hardy is quoted claiming "upwar...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.