Reddit
Spark-X2.5 arrives as a genuinely new 4B/1.7B architecture on Apache 2.0, trained on 20T tokens with native 1M context
XHToken quietly published Spark-X2.5-4B and Spark-X2.5-1.7B on Hugging Face. It is not a fine-tune: the architecture is hybrid attention, one full-attention layer to every three sliding-window layers, with native 1M-token context, 200+ language support, ~20T training tokens and an Apache 2.0 license. Claimed scores for the 4B include 65.1 on BFCL-V4, 75.1 on tau-squared-bench, 44.4 on SWE-Bench Pro and 90.7 on AIME 2026, benchmarked against Qwen3.5 and Gemma4 at 2B through 12B. It does not run on upstream llama.cpp yet — PR ggml-org/llama.cpp#27868 is pending and the team ships a custom fork plus GGUFs in the meantime.
↳ Follow the thread