Fetching from the wire…
Public story · 2026-07-23 · high
A specialized model from the same training run beat GPT-5.4-Mini by 3.98 points on Operations Research tasks.
Why now: As of July 23, most Western AI reporting still assumes frontier-scale training in China depends on Nvidia hardware.
SLAI post-trained DeepSeek-V4 to 34.22% MFU on an Ascend NPU SuperPOD, a 2.93 times improvement over the open source baseline recipe, per the arXiv preprint. Hitting that number on non-Nvidia hardware weakens the assumption, common in Western coverage and baked into export-control policy, that frontier-scale training still requires Nvidia GPUs.
MFU measures how much of a chip's peak compute a training run captures, so 34.22% means SLAI used about a third of the SuperPOD's ceiling. The gains came from optimizing three layers together: model parallelism, computation-communication orchestration, and low-level kernel code tuned to the SuperPOD's architecture. The paper reports training stability held throughout the run for a trillion-parameter MoE model.
A specialized version, DeepSeek-V4-Flash, scored 71.81% zero-shot Pass@1 on Operations Research tasks, 3.98 points ahead of GPT-5.4-Mini on the same benchmark.
SLAI publishing full non-Nvidia training internals, something Western labs rarely do even on GPUs, matters more than whether 34.22% MFU survives outside scrutiny. Watch whether more full training-stack writeups like this one come out of Ascend-based labs than Nvidia-based ones over the next year. That's the sign the training-stack gap is closing faster than the export-control conversation assumes.
As of July 23, most Western AI reporting still assumes frontier-scale training in China depends on Nvidia hardware.
Each link below shares sources, entities, or timing with this story.
DeepSeek released V4 Preview on April 24 with two open-weight variants: V4-Pro (1.6T total parameters, 49B activated via MoE) and V4-Flash (284B parameters, 13B activated). Both support 1M-token context windows. Both are Apache 2.0 licensed. Both are live right now on Hugging...
$3,054 against $38,370. Same benchmark, better score. Praxist (arXiv 2608.25955, submitted August 26) replaces per-attempt agent memory with a typed evidence graph of findings, plus lane-structured frontiers and agendas, so later attempts inherit validated mechanisms rather th...
For about a year, "run your agent locally" meant accepting a model that couldn't reliably call a tool twice in a row. That excuse is gone. Meta Superintelligence Labs published Muse Glimmer today: a 29.6B dense causal transformer, 52 layers, 6,656 hidden dim, with a ~1.8B ViT-...
DeepSeek-V4-Flash-0731 landed July 31 under MIT with a DSpark speculative-decoding module attached. Terminal Bench 2.1: 82.7. Toolathlon-Verified: 70.3. DSBench-FullStack: 68.7. DeepSWE: 54.4. NL2Repo: 54.2. The model card claims it beats DeepSeek-V4-Pro (Preview) "despite its...
DeepSeek released V4 on April 24 and the numbers demand attention. V4-Pro is 1.6 trillion parameters total with 49 billion active, MIT-licensed, native 1M-token context. It scores 80.6% on SWE-bench Verified, putting it within 0.2 points of Claude Opus 4.6. On Terminal-Bench 2...
GPT-5.6 Luna went to $0.20 input / $1.20 output per million tokens on July 30. That's an 80% cut. Terra dropped 20%. Luna's input now undercuts Gemini 3.1 Flash-Lite ($0.25/$1.50) and sits at one-fifth of Claude Haiku 4.5's $1 input. Simon Willison covered the announcement and...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.