By comparison, Moonshot AI’s Kimi K3 has 2.8 trillion parameters, while Meituan’s LongCat-2.0 and DeepSeek’s V4-Pro are estimated at around 1.6 trillion parameters each. The ByteDance model is currently in the pre-training stage. ByteDance Expands Its AI DevelopmentThe project forms part of ByteDance’s wider push into advanced AI. Its Seed team focuses on areas including model pre-training, post-training, inference, memory, learning and interpretability. Model distillation generally involves training a smaller model to reproduce knowledge or outputs from a larger model.