None
EN
Kimi K3 and DeepSeek V4 expose widening divide over native multimodality
['Cheng Zi']
KrASIA
Alibaba Group and ByteDance are also pursuing native multimodality, treating vision as a foundational capability of general-purpose models.
It can then pass that information to a text model for reasoning.
In a natively multimodal model, he said, the communication channel between visual input and the language backbone has greater bandwidth.
A natively multimodal model can interpret changes on the page directly and decide what to do next.
It has 2.4 trillion total parameters and activates 95 billion parameters for each token while supporting native vision, coding, and Cowork capabilities.