Ant Ling's 124B Ling-3.0-flash Claims Parity With Its Own 1-Trillion-Parameter Flagship at 5.1B Active Params
Announced July 23, Ling-3.0-flash is a hybrid-reasoning MoE with 124B total parameters activating only 5.1B per token — roughly one-eighth the total and one-twelfth the active parameters of Ant Ling's 1T flagship, which the founder's thread claims it matches or beats on most benchmarks. The architecture interleaves K-directed attention (KDA) and multi-layer aggregation (MLA) layers at a 5:1 ratio with 1/64 expert activation, shipping a 256K context window natively and claiming scalability to 1M with additional engineering. It is live on OpenRouter and free through August 3, 2026, with post-free-period pricing undisclosed; note the benchmark claims trace to the founder's X thread rather than an independent evaluation.
Source
↳ Follow the thread