Fetching from the wire…
Models2026-09-02 · source-backed
A two-month follow-up on a dual R9700 setup: MXFP4 first reached parity with FP8, then passed it using W4A8 kernels, which the builder believes is the hardware limit for these cards (r/LocalLLaMA). BetterBench decode with DFlash2 runs 280.0 t/s on JSON down to 116.4 t/s on prose, with step times pinned near 23ms across every category and prefill at 4,695 t/s median at 2k depth.
Each link below shares sources, entities, or timing with this story.
A developer published v100-skinny with hand-written NVFP4 W4A16 CUDA kernels plus chain-MTP speculative serving: four V100s at 219.1 ± 5.9 tok/s decode against a 5090 running NInfer at 214.7 ± 9.2, both 5/5 correct on AIME 2026 problem 1 across five seeds. (r/LocalLLaMA) The m...
FP8 at xhigh beat BF16 at medium and equalled BF16 at xhigh while running faster. r/LocalLLaMA On problem 7 both BF16 and FP8 at xhigh exhausted the generation budget without emitting an answer, so that miss is a context-limit failure rather than a wrong answer. Practical read...
The progression: 82, then ~114, then ~138 with DFlash2 drafting and lookup-augmented drafting, and now ~133 tok/s on real chat prompts with 382 tok/s when the model reproduces its own context. Stack is fp8 KV cache, int8 lm_head and embed_tokens, fp16 recurrent state, int8 act...
An r/LocalLLaMA thread asking where the promised MoE went turned up a hard artifact: modelscope/ms-swift commit a45f1d4, titled "fix wrong model-ids," removes the Qwen/Qwen3.8-35B-A3B and -FP8 entries from swift/model/models/qwen.py and substitutes the dense 27B. Why it matter...
Alibaba published a fine-grained MoE with 2.4T total / 95B active, 512 experts, and a 92-layer hybrid full/linear attention backbone. vLLM shipped day-0 support verified on NVIDIA and AMD with ready 4-bit checkpoints (NVFP4 at 1.32 TiB for an 8xB300 node, MXFP4 at 1.45 TiB for...
The file is 1.56TB. That's the first thing you notice about the moonshotai/Kimi-K3 Hugging Face repo that went live today. 2.8 trillion total parameters, 104B activated, 896 routed experts with 16 selected plus 2 shared per token, 93 layers split 69 Kimi Delta Attention and 24...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.