Fetching from the wire…
Infra2026-09-17 · source-backed
Fathom targets the scan that ranks all n keys for a top-k step when million-token agent sessions keep KV caches and indexes in host memory. It stores the 4-bit K cache channel-major as bit planes, so a prefix of t planes is exactly that channel's t-bit quantizer, and each query spends its bit budget by reverse water-filling over variance-weighted channel importance. At one million tokens on Qwen3-8B a decode step is 1.67x faster in GPU time than the 136-bit scans of Double Sparsity, Loki and SparQ, and on real coding-agent sessions it matches the most accurate 136-bit scan's step agreement at 92 bits.
Each link below shares sources, entities, or timing with this story.
Simulating HBM, DRAM and SSD tiers against a random-forest execution-time predictor across chat, agent and document QA workloads, tiering supported 73.02x more concurrent sessions per GPU at 62.04x lower cost per session. The authors attribute the gains to tier capacities of 1...
arXiv 2608.04074 reframes KV quantization as transform coding where distortion is measured on the attention product, deriving closed-form optimal transforms from calibration statistics that satisfy a generalized Parseval relation. At two bits per element it recovers most of th...
Activation probes are usually evaluated against agents who don't know they're monitored, which is a generous assumption. This study held models, probes and thresholds fixed and varied only the disclosure: nothing, monitor present, or monitor present plus last round's score. Ac...
Multi-turn serving hurts because the KV cache grows linearly with conversation length, choking GPU memory and bandwidth. Tangram spends memory unevenly across the cache instead of treating all tokens equally, cutting the footprint of long sessions. Source: arXiv Directly usefu...
Paritok-4B (arXiv 2608.24188) is a LoRA on Qwen3-4B distilled from a gpt-4.1-mini teacher over 67,074 real OpenHands trajectories. It's extractive rather than paraphrasing, with 96.0% of emitted identifiers, paths and numbers already present in its input, and intent-conditione...
Measuring expert time on two datacenter GPU generations shows it is linear in neither token count (EPLB, LPLB, UltraEP) nor activated-expert count (METRO). Below roughly 156 to 168 tokens, HBM weight streaming dominates so cost attaches to activated replicas. Above it, grouped...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.