Fetching from the wire…
Infra2026-09-22 · source-backed
Federico Viticci's review measures 1.2 TB/s memory bandwidth (50% over the M3 Ultra's 819 GB/s), an 80-core GPU with Neural Accelerators, and 256 GB unified memory on the tested unit with a 512 GB model due in October. On Qwen3.8-Flash-Next it does about 2,733 tok/s prompt processing and 54-108 tok/s generation, roughly 2.5x the M3 Ultra on prefill. An RTX 5090 still edges it on raw prompt processing at ~3,000 tok/s but can't hold three concurrent Flash-Next sessions the way 256 GB of unified memory can. (MacStories)
Each link below shares sources, entities, or timing with this story.
On a warmed M3 Ultra A/B, Qwen3.8-27B went from 458 tok/s baseline at 4K prefill to 588 with the ANE/GPU split and 625 with ANE/CPU/GPU, costing about 7 GB peak memory. GitHub The compile-memory drop is arguably the bigger deal, since a 35.8 GB spike locked out most Macs from...
The published retirement date for gemini-2.5-pro and gemini-2.5-flash is October 16, 2026, pushed back from an original June date, with Gemini 3.1 Pro and 3.6 Flash as the recorded upgrade paths at higher list prices. The specific loss named in the thread is 2.5 Pro's document...
Hardcoded model IDs now return API failures, and teams moving to 2.5-flash face a second mandatory cutover by October 16 (Google). Three major versions in under a year. The lesson isn't which model to pick, it's to stop hardcoding model IDs entirely. Abstract the selection or...
Part 4 of a running 2x3090 series: prefill was 80+ seconds to first token on an 8k prompt and 24 minutes on a 119k one, and releasing the 150-slot expert cache off the GPU during prompt processing bought the speedup. The comments are why it's here. A reader calculated the 2.8s...
The day-0 SGLang post dated August 26 gives the architecture the release megathread didn't: 125B main parameters plus a separate 51.2B Per-Layer Embedding table at about 95.4 GiB in BF16, 6B activated per token, 48 layers split 36 GDN linear attention and 12 QSA sparse attenti...
An r/LocalLLaMA builder patched vLLM to offload most of the KV cache to host RAM and reports 1M context on 3x RTX 3090: about 80 tok/s at short context, dropping to roughly 60 once QSA hits its 2,048-token budget and then staying flat as context grows, ~150 tok/s at four concu...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.