Fetching from the wire…
Infra2026-09-21 · source-backed
PR #57885 removes four validity copies and one unused division per decode step from sparse-attention metadata prep. On DeepSeek-V4.1-Flash across 4x GB200 at TP4, mean TPOT moved 4.467ms to 4.457ms and median ITL 4.383 to 4.365. The author states plainly that mean TPOT is effectively flat and claims no material end-to-end speedup. The change is justified by removing work, not by a number, which is a standard I wish more perf PRs held.
Each link below shares sources, entities, or timing with this story.
PR #56962, merged September 15, reopens a closed PR rebased onto the DeepGEMM fork pin and retargets the model path to deepseek_v41 (GitHub). CUPTI spans on a single GB200 at H=5120/7168 with hc_mult=4 show Mega-mHC ahead by 1.14x at one decode token and 1.51x at 64, widening...
PR #57204 removes the MegaMoE intermediate-width padding that had widened the 2304 checkpoint width to 2560. The bundled DeepGEMM uses layout::Data(..., false) for activation-scale rows, so the old 16-byte TMA row alignment rationale no longer applies, and removing it strips 1...
The September 5 release adds beam search via a beam_width request parameter returning the n best sequences, though it doesn't yet combine with speculative decoding, disaggregation, DP attention or HiCache. DeepEP v2's fixed-size ElasticBuffer engine as --moe-a2a-backend deepep...
"Gemini is Cooked but GCP is Cooking" argues Google quietly shelved 3.5 Pro, which industry chatter placed at roughly Opus 4.5 level, shipping Gemini 3.6 Flash as a bridge the authors call worse than Muse Spark 1.2, Grok 4.5, and tier-1 Chinese open-source models. The hard num...
Task executor migration to Go, a fifth batch of agent PRs ported, Go rewrites of the Tavily, Google search, Google Scholar, DuckDuckGo and Wikipedia canvas components, plus GCS support and an admin server in Go. GitHub GitHub now labels the 88,966-star repo as Go, which confir...
Willison's August 13 release adds Gemini 3.7 Flash, 3.6 Flash and 3.5 Flash-Lite plus gemini-embedding-2 and -001, rebuilding on LLM 0.32's structured message and streaming APIs so reasoning, tool calls and results emit as typed stream events while preserving Gemini thought si...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.