SK Telecom Releases A.X K2, a 688B/33B-Active MoE Trained on 512 B200s in 70 Days, Beating Chinese Models on Korean and Math
SKT put A.X K2 on Hugging Face on July 29, scaling from the 519B A.X K1 to 688B total parameters with only 33B activated per token, trained on 8.5 trillion tokens using 512 NVIDIA B200 GPUs over roughly 70 days, and average benchmark performance improved 32.2 percentage points over K1 despite fewer training tokens. It scores 97.1 on AIME26, 80.5 on KMMLU-Pro, 91.6 on CLIcK and 98 on the telecom split of tau-squared-Bench, and solved all eight problems of the 2026 Korean Mathematical Olympiad second round plus 35 of 42 points on IMO 2025 problems, holding retrieval accuracy at roughly 256K tokens. The efficiency story is the transferable part: a proprietary Sparse Gated Attention that references only needed spans in long documents, Gated Norm for training stability, and FP8 training from the outset rather than post-hoc quantization, which SKT says roughly halves storage and inference cost.
↳ Follow the thread