Fetching from the wire…
Public story · 2026-09-10 · high
On an M3 Ultra, decode ran at 59% of memory bandwidth because small kernels between weight-streaming steps ate dispatch latency.
Why now: The maintainer posted the benchmark and fix on September 10.
GLM 5.3 Flash was decoding at just 59% of an M3 Ultra's memory bandwidth, according to the maintainer's GLM 5.3 Flash benchmark post.
Idle bandwidth on hardware that's already paid for. The fix took decode speed from 29 to 40 tokens per second.
The big weight-streaming kernels weren't the issue. Dozens of smaller kernels ran between them, and each one paid its own dispatch latency.
Fusing that in-between work into larger dispatches raised decode speed at every context length tested. At 62k context, speed measured 24 to 38 tokens per second. A Claude Code harness run at 200k depth averaged over 38 t/s under the same fix.
That's 81% of the chip's bandwidth ceiling, up from 59%. On Apple Silicon, once weights stream fast enough, dispatch count becomes the limit.
The maintainer's post doesn't say whether the fusion technique generalizes to other models or depends on the q4 quantization path used here.
Each link below shares sources, entities, or timing with this story.
$3,054 against $38,370. Same benchmark, better score. Praxist (arXiv 2608.25955, submitted August 26) replaces per-attempt agent memory with a typed evidence graph of findings, plus lane-structured frontiers and agendas, so later attempts inherit validated mechanisms rather th...
Alibaba released Qwen3.6-27B on April 22. Dense architecture. Open weights. 77.2% on SWE-bench Verified, within 3.7 points of Claude Opus 4.6. On SkillsBench, it scores 48.2% versus its own 397B MoE predecessor's 30.0%. That's a 77% improvement with 14.8x fewer parameters. Let...
The defining number of developer tooling in 2026 isn't adoption. It's the gap between adoption and trust. Stack Overflow's latest analysis puts developer AI tool adoption at 84%, up from 76% in 2024. Usage keeps climbing. But trust in AI accuracy has cratered to 29%, down from...
A user running the dsh developer preview reports it worked for two hours where Claude Code stalls, then decided it needed more context, left the correctly configured project directory, and started reading elsewhere on disk (r/LocalLLaMA). The top comment argues Anthropic's har...
An r/LocalLLaMA post at 1,330 upvotes reports the first run of full K3, Moonshot's 2.8T open-weight MoE, on a 16x NVIDIA GB10 cluster with dspark speculative decoding: 20+ tok/s average, 38 peak, 750 prefill. That's roughly $64K of hardware for frontier-adjacent tokens at your...
diegosouzapw/OmniRoute added 1,343 stars on July 20, a single MIT-licensed gateway across 268+ providers (50+ free) and 500+ models including Claude, GPT, Gemini, Kimi K3, GLM and DeepSeek, wired for Claude Code, Codex, Cursor, Cline and Copilot. Quota-aware automatic fallback...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.