The model carries 124 billion total parameters but activates only around 5.1 billion of them per token during inference. Ling-3.0-Flash matches or beats Ling-2.6-1T, Ant’s previous flagship model, across core reasoning and instruction-following tasks. The native context window sits at 262,000 tokens, with Ant targeting eventual expansion to one million tokens. Built for agents, not just chatThe design brief for Ling-3.0-Flash is unusually specific: production-grade AI agents running at high frequency. Ant made the model available on Hugging Face under the MIT license shortly after the initial release, listing it under the inclusionAI organization.