Fetching from the wire…
Infra2026-09-17 · source-backed
arXiv 2609.17863 measured 54 configurations of Qwen2.5-7B-Instruct on vLLM 0.12 across L4, A100 and H100, then calibrated a simulator reproducing them with cross-campaign drift under 1.5%. On the calibrated grid, 18 of 36 configurations reach the cost/quality/latency frontier, and combined optimizations reach it more often than single ones. Quality testing on 200 GSM8K questions reorders the winners: AWQ 4-bit cuts per-token latency to 0.34x baseline on L4 but loses 5.9% strict accuracy, narrowly missing a 95% quality floor, while flexible answer extraction recovers FP16 parity. The loss is formatting, not arithmetic, which is a different fix entirely.
Each link below shares sources, entities, or timing with this story.
Translating key-value state from one model into a form another can consume works across scale, architecture, attention configuration, tokenizer and family (arXiv 2608.30963). Llama3.1-70B to Qwen2.5-7B reaches 44.0% accuracy against 45.7% native while dropping latency to 138ms...
Auditing Qwen2.5-7B-Instruct on RGB and HotpotQA with a hallucination detector, NLI entailment and an LLM judge, INT8 is near-lossless on accuracy and faithfulness (arXiv 2608.30996). INT4 lowers accuracy, and among answers that stay factually correct, over 90% of faithfulness...
Counterfactual regret minimization has been one of the few large numerical workloads that ran faster on CPUs, because each iteration issues millions of tiny interdependent gather and scatter steps where kernel launch and framework dispatch dominate. GPU-CFR uses the fact that...
NCP-ArchPreview (arXiv 2609.10715) trains an 8.9B model on 5.73T Dolma-3 tokens. Alongside next-token prediction, it predicts the next concept from a product-quantized vocabulary built from the model's own hidden states. After full pretraining it beats OLMo-3-7B by 2.45 points...
arXiv 2608.24641 partially replicated Khojah et al. across Zero-Shot, Few-Shot, Chain-of-Thought, Contrastive CoT and an adapted Program-of-Thought on three version pairs (GPT-3.5-Turbo/GPT-4o, Qwen2 7B/Qwen2.5 7B, Mistral-7B-Instruct/Mistral-Large) over 218 context-rich Pytho...
English, Spanish, French, German, Italian, Vietnamese, Mandarin, Hindi, Japanese, Modern Standard Arabic, Korean, Brazilian Portuguese, male and female voices each (HF). Latency by hardware: 32ms/239ms at 64 concurrent on B200, 47ms/275ms on H100, 79ms/395ms on A100. Character...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.