Zhiyuan's WITA-Omni Preview Tops DailyOmni Multimodal Understanding Benchmark
Zhiyuan AGIBOT's WITA-Omni Preview has achieved first place on the DailyOmni benchmark with a score of 85.21, outperforming models from Qwen, Gemini, Doubao, and Nvidia across six of eight metrics.
Zhiyuan's WITA-Omni Preview Tops DailyOmni Multimodal Understanding Benchmark
July 28, 2026 — According to an announcement from Zhiyuan AGIBOT, their self-developed embodied native multimodal large model, WITA-Omni Preview, has achieved the top position on the authoritative DailyOmni benchmark for embodied full-modal understanding. The model scored 85.21 overall, surpassing well-known models such as Qwen, Gemini, Doubao, and Nvidia, and secured first place in six out of eight sub-indicators.
DailyOmni is widely recognized as a third-party benchmark for testing audio-video temporal alignment and cross-modal joint reasoning. Unlike typical tasks designed for images and short videos, DailyOmni extensively uses real-world audio-visual materials, evaluating the model's ability to fuse visual scenes with environmental audio, accurately identify correspondences between sound and image, understand event sequences, and perform cross-modal semantic analysis. These capabilities are directly relevant to the core perception skills required for humanoid robots to interact in real physical spaces.
According to the company, the key to WITA-Omni's success lies in its architecture. Rather than simply embedding a chat model into a robot, WITA-Omni elevates physical actions and facial expressions to primary outputs alongside speech. It uses a native end-to-end model that uniformly drives perception, decision-making, and expression, expanding from the Thinker-Talker paradigm to a Thinker-Talker-Actor architecture tailored for embodied outputs.
The model's performance is not solely due to scale but to a layered data system and targeted training methods built around embodied full-modal interaction. During training, WITA-Omni utilized millions of hours of open-source and proprietary multimodal data to upgrade its strong text and vision base into a full-modal model capable of "hearing, seeing, and jointly reasoning about audio and visuals."
Public audio-video data often lacks turn-taking, response timing, action, and expression information essential for real interaction. To address this, Zhiyuan constructed a large-scale, high-quality full-modal dataset centered on humans in real interaction scenarios, preserving the natural temporal relationships among sound, visuals, language, actions, and expressions.