ByteDance Is Pre-Training a 10-Trillion-Parameter Model Aimed Squarely at Anthropic's Mythos, FT Reports
Slashdot / Financial Times·high signal
The Financial Times reported on August 7, citing three people familiar with the project, that ByteDance is training a model with as many as 10 trillion parameters — over three times the size of Moonshot's Kimi K3 at ~2.8T, currently among the largest Chinese models. The model is still in pre-training, a phase that typically runs three to six months before fine-tuning and any release. Treat the number skeptically: without active-parameter count, training budget, data mix or any published evals, 10T says almost nothing about capability, and the FT noted it could not independently verify the details.