Zhiyuan's full-modal large model has reached number one globally! Surpassing Google and Nvidia, it outperforms major companies.

The integration of intelligence, intelligence, and empathy is constantly evolving, ushering in a new future of embodied intelligence.

Zhiyuan's full-modal large model has reached number one globally! Surpassing Google and Nvidia, it outperforms major companies.

When a robot stands in an exhibition hall where many people are talking, can it tell from the noise who is talking to it and whether a sentence needs a response? Can it pick up the conversation when the other person pauses and continue naturally after being suddenly interrupted? Can it answer questions while making appropriate gestures and expressions to match its tone?

These seemingly ordinary details of communication are the core capabilities of embodied intelligence, and also the most difficult link to fill when robots cross the threshold of natural interaction.

For a long time, the robotics industry has been trapped in an awkward paradox: continuous breakthroughs in individual capabilities have led to the development of "..."See, hear, speak, move"It was kneaded into a coherent, human-like whole, but it was always lacking a breath."

The essence of this chasm isInteractionThe capabilities are not keeping up—perception, reasoning, speech, and action each operate independently and are never coordinated on the same timeline.

Recently, WITA-Omni Preview, a full-modal large model developed by Logic Interactive, has delivered a significant result.

On DailyOmni, the industry-recognized authoritative ranking list for embodied multimodal understanding, it topped the list with an overall score of 85.21, surpassing leading domestic and international models such as Qianwen, Gemini, Doubao, and NVIDIA, and taking first place in six of the eight sub-indicators.

7d17b22aff32e818f7eee74c9d2c10a7.jpg
7d17b22aff32e818f7eee74c9d2c10a7.jpg

On the core interactive capability line of "understanding the environment, understanding the sound, thinking while listening, and acting while speaking," a company focused on embodied intelligence has taken the lead over general large-scale model manufacturers.

Previously, Zhiyuan's self-developed world model, Genie Envisioner-Sim 2.0, had already won first place in the WorldArena "World Model Perception and Motion Response" track.This time, WITA-Omni once again topped the all-modal understanding leaderboard, with two achievements corresponding to the key capabilities of robots in understanding the physical world, simulating the physical world, and interacting with the physical world.

This has led outsiders to observe that, in addition to manufacturing the robot itself, Logic is simultaneously building a self-developed AI landscape integrating "interaction, operation, and movement."

01.

A major test of "sound and image matching":

What criteria does DailyOmni use to verify true full-modal capability?


The simple addition of "being able to see images + being able to hear sounds" does not equate to full-modality computing. The real challenge lies in whether the model can distinguish who is speaking, what happened first, and which detail to focus on when sound and images arrive simultaneously. Its core is the synchronous and collaborative ability to hear, see, and think simultaneously.

DailyOmni's evaluation logic precisely addresses this challenge.

It extensively uses audio and video from real open environments to examine the correspondence between sound and image, the judgment of event sequence, and cross-modal joint reasoning. It tests whether the model can "know who is speaking in the picture when it hears the sound" and "know the sequence of events when it sees the action", and understand what is happening in a real physical world in a continuous time stream.

Unlike image-based evaluations that can be completed using only a single visual or linguistic prior, many questions in DailyOmni require the model to simultaneously utilize both auditory and visual information; neither can be omitted.

WITA-Omni's performance report: topped the list with an overall score of 85.21, 86.13 for audio-video alignment, 84.64 for event timing, and 85.06 for integrated reasoning, and also achieved a leading position in long-term tasks such as comparative understanding and 30-second/60-second video understanding.Six out of eight indicators ranked first, establishing a systematic lead across the entire main line of "audio-visual alignment, temporal judgment, and cross-modal reasoning".

Let's look at the benchmarks for competition. On the DailyOmni list are commercial or flagship models from general-purpose large-scale model manufacturers such as Qianwen, Gemini, Doubao, and NVIDIA. Zhiyuan is one of the few "intelligent-integrated track players" among them.

A company that started with robotics has outperformed major model manufacturers in full-modal understanding. Beyond the score difference, this ranking points to the choice of technological approach: building full-modal capabilities from native robot interaction scenarios is proving its competitiveness.

WITA-Omni's top ranking proves that Calcium Robotics has taken the lead in enabling robots to "naturally and fluently understand and respond to the world." For the industry, this is also a cognitive calibration: in the past, the market focused more on Calcium Robotics' robot assembly and mass production capabilities; but this top ranking shows that Calcium Robotics' AI large-scale model capabilities have entered the first tier of competition with major companies.


02.

Thinker–Talker–Actor:

Let the robot think, speak and move at the same time


Traditional robot interaction is like an assembly line: first, recognize speech; then, understand semantics; then, generate a response; then, synthesize speech; and finally, output an action. Each step has a delay, and each change of link in the chain results in loss. As a result, the robot "thinks before it speaks, and speaks before it acts," failing to achieve the natural feeling of thinking and expressing oneself as humans do.

This cascaded architecture has three inherent flaws..

  • Firstly, the serial calls to multiple modules cause a continuous accumulation of response delays, making the sluggishness increasingly noticeable over time.

  • Secondly, the conversion of information across modules results in semantic loss; the understanding from one stage is already compromised when it is passed to the next.

  • Third, voice, actions, and facial expressions are generated independently, making it impossible to maintain rhythm and emotional synchronization: the mouth is smiling but the hands are not following, or the words are finished but the facial expression arrives late.

Ultimately, a robot may be able to function on every module, but it will never truly resemble a coordinated and complete intelligent agent.

WITA-Omni's native architecture breaks this serial paradigm at its core.

It extends the Thinker-Talker paradigm with the Actor module, elevating robot actions and expressions to a first-level output alongside speech..

Thinker is the core of multimodal reasoning.Text, images, audio, and video are mapped to a unified representation space through their respective encoders and lightweight projection layers, enabling cross-modal joint reasoning and high-level interactive decision-making;Talker is conditional on Thinker's hidden state.Streaming to generate natural speech;An actor consists of an action head and an expression head, which drive the actions and expressions..

The three modules proceed synchronously on a unified timeline, replacing the traditional serial mode of "processing one module before handing it over to the next".

The cross-stream synchronization mechanism further enhances this synchronization. The Actor not only receives the Thinker's semantic state, but also takes into account the Talker's audio latent variables, enabling the action to perceive speech rate, pauses, stress, and emotional changes.

56de6996fef91aeb9b66bd8d5e618b94.jpg
56de6996fef91aeb9b66bd8d5e618b94.jpg

The robot's gestures and facial expressions respond in real time during the speech generation process:Speaking and moving simultaneously, expressing oneself while moving, with voice, actions, and facial expressions all on the same timeline, allowing for simultaneous understanding and response.

The core breakthrough of the WITA-Omni architecture lies in the word "native," which transforms "understanding" and "expression" from two separate steps into different dimensions of the same generative process.

Most general-purpose full-modal models from large-scale model vendors prioritize adapting to online digital content interaction; however, Logic AI, starting from the native interaction needs of robots, has redefined the structure of its full-modal model. This is not only an innovation in engineering architecture but also reflects the differentiated competitiveness brought about by the unique perspective of an embodied intelligence company.


03.

Millions of hours of training + thousands of hours of high-quality post-training

The "steelmaking" of embodied data


The competition among base models is moving beyond the dimension of "who has the most parameters," while the data challenge of embodied intelligence is even more difficult: publicly available audio and video data lacks interaction rounds, response timings, and temporal annotations of actions and expressions. Models trained on this data will never learn "when to speak and when to move."

WITA-Omni's leading position stems from a hierarchical data system and targeted training methods built around embodied multimodal interaction.

During the training phase, WITA-Omni usedTens of millions of hours of multimodal dataThe strong text and visual foundation will be upgraded to a full-modal model that can "hear, see, and perform audio-visual joint reasoning".

The data ratios were carefully designed.Audio and video combined dataApproximately 40% of the training focuses on the correspondence between sound, visuals, and events.Pure audio dataApproximately 30% is used to enhance speech, acoustic events, and environmental sound understanding;Visual and text dataApproximately 20% and 10% of these are used to maintain existing visual and language abilities.

This phase focuses on three questions: which subject the sound originates from and which visual event it corresponds to; the chronological order and temporal relationship of different events; and cross-modal reasoning that requires the simultaneous combination of sound and visuals. The logic of data configuration always revolves around the core capability requirements of embodied interaction, rather than simply increasing the quantity.

Publicly available data is far from sufficient. Therefore,Zhiyuan has built its own large-scale, high-quality, full-modal dataset centered on humans and oriented towards real-world interaction scenarios.It fully preserves the natural temporal relationship between sound, image, language, action and expression.

The data mainly comes from three types of scenarios:

  • Human-to-human and human-to-computer interaction videos with synchronized voice, actions, and facial expressions, with annotations of turn-by-turn switching and response timing;

  • The robot collects first-person perception and motion data from its own sensors during teleoperation interaction;

  • A semi-automatic data pipeline is used to transform raw videos into supervised samples with annotations for speakers, actions, events, time intervals, instructions, and questions and answers.

This data allows the model to learn "when to speak, to whom to speak, and how to coordinate actions" directly from real interactions, skipping the step of manually piecing together multiple modules of rules.

On thousands of hours of high-quality dataWITA-Omni further adopts a three-stage post-training strategy: SFT–OPD–RL..

  • SFT supervised fine-tuning establishes foundational capabilities for full-modal understanding, interactive decision-making, and multi-stream generation.

  • OPD same-policy distillation allows the same model to act as both "teacher" and "student" simultaneously. The teacher model generates high-quality answers after acquiring privileged information such as additional time intervals and target objects, and then distills the capabilities back to the student model, which does not rely on privileged information. This enhances the reasoning ability in audio and video contexts while maintaining the model's own expressive style.

  • RL reinforcement learning focuses on optimizing Thinker's interactive decision-making ability, with reward signals covering verifiable answers, time interval matching, target person selection, waiting or responding decisions, and active trigger accuracy, and optimizing model strategies through GRPO.

Thus, the model not only learns to "answer correctly," but also learns to judge whether to respond, when to respond, and to whom in real-world scenarios.

Around the core scenario of embodied interaction, WITA-Omni has built a complete system of capabilities at the two levels of data collection and training strategies. Ultimately, WITA-Omni integrates audio and video joint understanding, interactive decision-making, and synchronous expression, enabling the robot not only to "understand what it sees and hears" but also to respond naturally. This is its core competitiveness in embodied interaction scenarios.

50d6b0f7d41966fa08f6dc994bd5f33e.jpg
50d6b0f7d41966fa08f6dc994bd5f33e.jpg


04.

Not only do we need to "think accurately"

More importantly, you should "chat like a human being."


Accurate understanding is merely the minimum requirement; the higher standards of interactive intelligence lie in the details: whether it interrupts you when you speak, whether it can quietly think when you pause, and whether it can naturally return to the point when you interrupt. These details determine whether humans and robots can truly "get along."

WITA-OmniVoice interaction sideIt also withstood the test.

On two publicly available benchmarks used by the internationally authoritative third-party large model evaluation platform Artificial Analysis Speech Arena, Zhiyuan conducted evaluations using the same protocol: Big Bench Audio, which measures speech reasoning, and a subset of Full Duplex Bench, which measures full-duplex dialogue capabilities (covering pause handling, turn-switching, interruption handling, and echo handling).

26f1b13e7e291b39335870dae3b59c91.png
26f1b13e7e291b39335870dae3b59c91.png

▲Big Bench Audio, a measure of speech reasoning

8f6d3d32e3d310c3ce42a443b4d9af0c.png
8f6d3d32e3d310c3ce42a443b4d9af0c.png

▲ A subset of the Full Duplex Bench that measures full-duplex dialogue capability

WITA-Omni Preview achieved first place in both categories, surpassing closed-source commercial models such as GPT-Realtime.

Full-duplex dialogueSimultaneous listening and speaking, and real-time judgment of speech turn switching, is an extremely difficult aspect of voice interaction. WITA-Omni's leading capability in this area means that it can react naturally in the rhythm of real conversations: speak when it should speak, stop when it should stop, resume naturally when interrupted, and maintain its rhythm when hearing agreement.

Many models can answer questions correctly in offline evaluations, but their shortcomings are exposed when they enter real-world dialogue scenarios: they don't stop when they should, hesitate when they should speak, and lose focus when interrupted. WITA-Omni's leading performance in dynamic dialogue metrics shows that "correct interaction rhythm" and "strong comprehension ability" are equally important, and perhaps even more difficult to achieve.

This sense of rhythm is the direct experience users have when robots enter homes, factories, and service scenarios. Users don't care about the model's parameters; they only directly feel whether "this robot speaks like a human."

WITA-Omni's capabilities cover "Perception—Reasoning—Expression—Interaction"A complete link. This end-to-end capability marks a key difference between it and the general large model, which is characterized by strong single-modality and weak interaction."

General-purpose large-scale model vendors' multimodal models are more geared towards content consumption scenarios; while WITA-Omni was designed from the outset with realistic physical interaction in mind, giving it a natural advantage in the dimension of interaction rhythm. Understanding what you see and hear determines the lower limit, while smooth conversation and natural movement determine the upper limit.WITA-Omni's leading position in voice interaction completes a crucial link in embodied intelligence in the dimension of "natural expression".


05.

Conclusion: Integrated Intelligence, Leading the Industry


For a long time, Zhiyuan has been building its differentiated competitiveness around the "three intelligences integration" of operational intelligence, interactive intelligence, and motion intelligence.

In Zhiyuan's "One Body, Three Intelligences" architecture, motion intelligence is the basic intelligence, serving as the actuator of the physical carrier; operational intelligence provides labor productivity, enabling robots to work autonomously; and interactive intelligence provides service productivity, allowing robots to communicate naturally.

Previously, Logic had already achieved the top score in the WorldArena embodied intelligence comprehensive evaluation, and its GO series basic models for operational intelligence and its motion control base model for motion intelligence have established industry-leading advantages. This global achievement in interactive intelligence demonstrates Logic's complete capability portfolio integrating three intelligences.

The significance of the integration of the three intelligences (motor intelligence, operational intelligence, and interactive intelligence) will be particularly clear in 2026. This year, the embodied intelligence industry will see a crucial leap from development to deployment—robots will enter production lines, stores, and homes, creating real productivity. The core threshold for deployment is the synergy of the three intelligences: motion intelligence determines whether a robot can enter a scenario, operational intelligence determines whether it can complete tasks, and interactive intelligence determines whether it can provide emotional value and generate service productivity. Among the seven productivity solutions pioneered by Logic, the complete capability stack integrating the three intelligences is its underlying support.

As robots gradually enter factories, commercial services, and public spaces, the standards for measuring intelligence levels will transcend individual indicators and move towards overall collaboration. WITA-Omni's achievement brings robots one step closer to natural interaction in the real world and also solidifies Zhiyuan's positioning of "having both a physical body and AI" into a verifiable capability loop.