Fetching from the wire…
Public story · 2026-07-23 · high
A specialized model from the same training run beat GPT-5.4-Mini by 3.98 points on Operations Research tasks.
Why now: As of July 23, most Western AI reporting still assumes frontier-scale training in China depends on Nvidia hardware.
SLAI post-trained DeepSeek-V4 to 34.22% MFU on an Ascend NPU SuperPOD, a 2.93 times improvement over the open source baseline recipe, per the arXiv preprint. Hitting that number on non-Nvidia hardware weakens the assumption, common in Western coverage and baked into export-control policy, that frontier-scale training still requires Nvidia GPUs.
MFU measures how much of a chip's peak compute a training run captures, so 34.22% means SLAI used about a third of the SuperPOD's ceiling. The gains came from optimizing three layers together: model parallelism, computation-communication orchestration, and low-level kernel code tuned to the SuperPOD's architecture. The paper reports training stability held throughout the run for a trillion-parameter MoE model.
A specialized version, DeepSeek-V4-Flash, scored 71.81% zero-shot Pass@1 on Operations Research tasks, 3.98 points ahead of GPT-5.4-Mini on the same benchmark.
SLAI publishing full non-Nvidia training internals, something Western labs rarely do even on GPUs, matters more than whether 34.22% MFU survives outside scrutiny. Watch whether more full training-stack writeups like this one come out of Ascend-based labs than Nvidia-based ones over the next year. That's the sign the training-stack gap is closing faster than the export-control conversation assumes.
As of July 23, most Western AI reporting still assumes frontier-scale training in China depends on Nvidia hardware.
Each link below shares sources, entities, or timing with this story.
NVIDIA invested in OpenAI / Shared entities / Earlier coverage
Linked by a graph relationship (NVIDIA invested in OpenAI); both cover DeepSeek, Flash, GPT, MoE; earlier DeepSeek coverage from 2026-04-24.
DeepSeek partners with Huawei / Shared entities / Earlier coverage
Linked by a graph relationship (DeepSeek partners with Huawei); both cover DeepSeek, Flash, MoE, NVIDIA; earlier DeepSeek coverage from 2026-04-26.
NVIDIA partners with Google / Shared entities / Earlier coverage / Tension
Linked by a graph relationship (NVIDIA partners with Google); both cover DeepSeek, GPT, MoE; earlier DeepSeek coverage from 2026-04-20.
NVIDIA partners with Google / Shared entities / Earlier coverage
Linked by a graph relationship (NVIDIA partners with Google); both cover DeepSeek, GPT, MoE; earlier DeepSeek coverage from 2026-05-01.
Anthropic uses NVIDIA / Shared entities / Earlier coverage
Linked by a graph relationship (Anthropic uses NVIDIA); both cover GPT, GPU, MoE; earlier GPT coverage from 2026-04-09.
NVIDIA released Blackwell / Shared entities / Earlier coverage
Linked by a graph relationship (NVIDIA released Blackwell); both cover DeepSeek, GPU, NVIDIA; earlier DeepSeek coverage from 2026-04-03.
Anthropic uses NVIDIA / Shared entities / Same source domain / Tension
Linked by a graph relationship (Anthropic uses NVIDIA); both cover Flash, GPT; reported by the same outlet (arxiv.org).
Anthropic uses NVIDIA / Shared entities / Shared topic / Earlier coverage
Linked by a graph relationship (Anthropic uses NVIDIA); both cover DeepSeek, GPT; overlapping topics (baseline, model).