CANN Bench Opens Agent Kernel Generation Beyond CUDA: 53 Operators, 1,060 Test Cases on Huawei Ascend NPU
A 15-author Huawei-affiliated team argues that kernel-generation benchmarks are almost entirely CUDA and Triton, leaving less-documented hardware ecosystems with no shared yardstick. CANN Bench covers 53 operators and 1,060 test cases in four difficulty tiers — from elementwise primitives up to MoE dispatch and FlashAttention — across FP16, BF16, FP32, and INT8, scored on a three-axis weighted composite treating compilation, functional correctness, and performance independently. Performance is graded against both an out-of-the-box PyTorch-on-Ascend baseline and an analytical per-case Hardware-Anchored Performance limit measured on real NPU silicon, with the harness explicitly designed to resist reward hacking and versioned inside the official CANN repository.
↳ Follow the thread