Fetching from the wire…
Infra2026-09-16 · source-backed
JustFit combines compressed KV execution, component residency swapping and state-preserving serving transitions. On an M4 Pro running Qwen3.8-27B MXFP4, three independent runs completed 196,608 input and 16,384 output tokens, lifting single-request context from 30,720 positions to 212,992. A 32K-input probe reached 19.11 tokens/s with a median peak footprint of 16,374 MiB, and the runtime answered 29 of 30 AIME 2026 problems correctly.
Each link below shares sources, entities, or timing with this story.
Accuracy drops 30–50% well before you hit the documented context limit. Not at the limit. Before it. Cross-model testing across GPT-4.1, the Claude 4 family, Gemini 2.5, and Qwen3 quantified what everyone shipping long-context features has felt and couldn't measure (Glasp). Th...
Community 4-bit quantizations of Qwen3.8-27B kept the 48 GDN layers' decay and write-strength gates at 8 or 16 bits, on the intuition that recurrence errors accumulate. Minima pushes NVFP4 W4A4 through all 496 linear layers including GDN and matches BF16 within seed noise acro...
DeepGrove published it to Hugging Face under MIT: 20B total and 1B active, 24 layers, 256 experts with 8 active, 3:1 SWA-512:GA attention, 131,072-token context. Reports 5–16x faster than Gemma 4, Qwen3.5 and gpt-oss at comparable quality, with scores on LCBv6, AIME 2026, HMMT...
Paritok-4B (arXiv 2608.24188) is a LoRA on Qwen3-4B distilled from a gpt-4.1-mini teacher over 67,074 real OpenHands trajectories. It's extractive rather than paraphrasing, with 96.0% of emitted identifiers, paths and numbers already present in its input, and intent-conditione...
A developer published v100-skinny with hand-written NVFP4 W4A16 CUDA kernels plus chain-MTP speculative serving: four V100s at 219.1 ± 5.9 tok/s decode against a 5090 running NInfer at 214.7 ± 9.2, both 5/5 correct on AIME 2026 problem 1 across five seeds. (r/LocalLLaMA) The m...
The leaderboard says first place. The methodology says you should check your own bill. Qwen3.8 Max now ranks first on Artificial Analysis' agentic index, scoring 86.1 on OSWorld-Verified ahead of GPT-5.6 Sol Max at 83.2 and Fable 5 at 85.0, priced at $2.00/M input and $6.00/M...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.