Fetching from the wire…
Research2026-09-19 · source-backed
arXiv 2609.19499 fixes N=8 on 500 GSM8K prompts and compares four generation schedules (1x8, 2x4, 4x2, 8x1) on A100s. Eight serial calls consume 4.64 to 4.86 times the gross GPU-device energy and show 5.77 to 6.12 times the P95 latency of one batched call producing the same eight candidates, replicated across three independently scheduled A100 nodes and short-output SciQ/V100 runs. Raising N from 1 to 8 gained 8.4 accuracy points on Phi-3-mini and 18.4 on Qwen2.5-1.5B. arXiv The quality gain is real and the reported budget N hides a 5x cost swing.
Each link below shares sources, entities, or timing with this story.
arXiv 2609.17863 measured 54 configurations of Qwen2.5-7B-Instruct on vLLM 0.12 across L4, A100 and H100, then calibrated a simulator reproducing them with cross-campaign drift under 1.5%. On the calibrated grid, 18 of 36 configurations reach the cost/quality/latency frontier,...
Counterfactual regret minimization has been one of the few large numerical workloads that ran faster on CPUs, because each iteration issues millions of tiny interdependent gather and scatter steps where kernel launch and framework dispatch dominate. GPU-CFR uses the fact that...
Foundry-Local, alongside community runtimes like LocalAI and cua's sandboxes, points to a first-party push to run agents locally with GPU acceleration and OpenAI-compatible APIs. For cost- and privacy-sensitive loops that fan out lots of cheap subagent calls, a local runtime b...
Translating key-value state from one model into a form another can consume works across scale, architecture, attention configuration, tokenizer and family (arXiv 2608.30963). Llama3.1-70B to Qwen2.5-7B reaches 44.0% accuracy against 45.7% native while dropping latency to 138ms...
Replicated David Ng's RYS method — duplicating layers 12-14 routes hidden states through the reasoning circuit twice, boosting BBH Logical Deduction from 0.22 to 0.76 and GSM8K from 0.48 to 0.64. Zero training, no weight changes. Same technique on Qwen2.5-Coder-32B improved re...
arXiv 2608.08097 exploits decode-time attention sparsity: keep a 2,048-token budget resident and speculatively prefetch the rest via lookahead prediction. Reported results are 1.69x speedup on reasoning workloads, up to 2.1x on multi-GPU long-context serving, roughly 2x throug...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.