SHANGHAI, July 28, 2026 /PRNewswire/ — AGIBOT today announced that WITA-Omni Preview, its multimodal foundation model for embodied interaction, has ranked first on the Daily-Omni audio-visual reasoning benchmark. WITA-Omni Preview ranks first on Daily-Omni with an average accuracy of 85.21%Understanding Audio and Visual Information TogetherDaily-Omni is a third-party benchmark designed to evaluate audio-visual reasoning and temporal alignment in everyday scenarios. The Thinker-Talker-Actor architecture consists of three core components:Thinker serves as the multimodal reasoning core. As the foundation of AGIBOT’s Interaction Intelligence, WITA-Omni will continue to evolve alongside the company’s Manipulation Intelligence and Locomotion Intelligence under its “Three Intelligences in One” architecture. AGIBOT’s “Three Intelligences in One” architecture integrates Locomotion Intelligence, Interaction Intelligence, and Manipulation Intelligence into a unified embodied system.