Research
SLAI T-Rex Reports 34.22% MFU Post-Training DeepSeek-V4 on Ascend NPUs — 2.93× Over the Open-Source Baseline Recipe
A rare end-to-end systems report for trillion-parameter MoE post-training outside the GPU world: hierarchical optimization across model parallelism, computation-communication orchestration, and low-level kernels on an Ascend NPU SuperPOD reaches 34.22% Model FLOPs Utilization, a 2.93× improvement over the open-source baseline recipe, with training stability maintained. On top of it they build CPT/SFT pipelines for Operations Research using solver-verified synthetic optimization documents (10K SFT samples, four task categories), and the specialized DeepSeek-V4-Flash model posts the highest zero-shot Pass@1 at 71.81% — 3.98 points over GPT-5.4-Mini and 11.27 over its own base.
↳ Follow the thread