Sources
Upstage Ships Solar Open 2: a 250B/15B-Active Agent-First Open-Weight MoE That Runs on Two H200s Quantized
Korea's Upstage released Solar Open 2 (250B total, 15B active, 320 routed + 1 shared experts, top-8) under a commercially usable license, trained on ~12T tokens over 2M B200 GPU-hours. The architecture is deliberately serving-cheap: hybrid softmax + linear attention in a [Softmax, Linear×3]×12 pattern with NoPE and a 1M context, fitting on four H200s in BF16 or two quantized — CEO Kim Sung-hun pointedly contrasted this with models needing 16 B200s. Reported scores include 70.4 SWE-Bench Verified, 92.4 LiveCodeBench, 86.2 MMLU-Pro and 58.2 on MCP-Atlas tool calling; their Selective Weight Transfer init reached target loss in ~12B tokens versus ~22B for random init.
↳ Follow the thread